MCPcopy Create free account

hub / github.com/NVIDIA/cutlass / functions

Functions19,553 in github.com/NVIDIA/cutlass

↓ 5 callersMethodepilogue
Epilogue implementation for non-TMA store version. :param epi_tidx: Thread index :type epi_tidx: cutlass.Int32 :para
examples/python/CuTeDSL/cute/blackwell/kernel/dense_gemm/dense_gemm.py:1052
↓ 5 callersFunctionexponent_biased
Returns the biased exponent
include/cutlass/exmy_base.h:1070
↓ 5 callersFunctionf
(t)
python/docs/_static/clipboard.min.js:7
↓ 5 callersMethodfence_tensormap_initialization
( self, *, loc: Optional[ir.Location] = None, ip: Optional[ir.InsertionPoint]
python/CuTeDSL/cutlass/utils/tensormap_manager.py:106
↓ 5 callersFunctionfixup
include/cutlass/gemm/kernel/sm90_tile_scheduler_stream_k.hpp:406
↓ 5 callersMethodforeach_read_tensor
Execute the given function for each supplemental read tensor.
examples/python/CuTeDSL/cute/blackwell/efc/common_efc.py:1483
↓ 5 callersMethodgK_strideM
examples/41_fused_multi_head_attention/kernel_backward.h:713
↓ 5 callersMethodgV_strideM
examples/41_fused_multi_head_attention/kernel_backward.h:716
↓ 5 callersFunctiongcd
include/cutlass/fast_math.h:167
↓ 5 callersMethodgen_code
(self)
examples/44_multi_gemm_ir_and_codegen/ir_gen/gen_verify.py:53
↓ 5 callersMethodget
Returns a pointer
include/cutlass/transform/threadblock/predicated_tile_access_iterator.h:1001
↓ 5 callersMethodget
include/cutlass/pipeline/sm90_pipeline.hpp:127
↓ 5 callersMethodgetOnDiag
Return if the address in on the diagonal
include/cutlass/transform/threadblock/predicated_tile_access_iterator_triangular_matrix.h:874
↓ 5 callersFunctionget_c_pointers
Given the `obj`, recursively go through it to extract all contained C pointers
python/CuTeDSL/cutlass/base_dsl/typing.py:222
↓ 5 callersMethodget_desc_ptr
(self, tensor_name: str, executor_idx: Int32)
examples/python/CuTeDSL/cute/blackwell/kernel/moe/moe_utils.py:710
↓ 5 callersMethodget_gmem_tensor
( self, tensor_name: str, gmem_tensor_in_moe_view: cute.Tensor, offs: Tuple[cu
examples/python/CuTeDSL/cute/blackwell/kernel/moe/moe_sched_extension.py:315
↓ 5 callersMethodget_k_tile_count
Get the current k_index, k_tile_count, and local split_kv value for the MLA kernel. :param split_kv: Split_kv value :type split_kv: c
examples/python/CuTeDSL/cute/blackwell/kernel/attention/mla/mla_decode_fp16.py:1397
↓ 5 callersMethodget_k_tile_count
Get the current k_index, k_tile_count, and local split_kv value for the MLA kernel. :param split_kv: Split_kv value :type split_kv: c
examples/python/CuTeDSL/cute/blackwell/kernel/attention/mla/mla_decode_fp8.py:1463
↓ 5 callersFunctionget_linearized_problem_shape_MNKL
include/cutlass/conv/detail.hpp:118
↓ 5 callersMethodget_masked_trip_count
examples/77_blackwell_fmha/collective/fmha_fusion.hpp:53
↓ 5 callersMethodget_registered_adapter
Get the registered JIT argument adapter for the given argument.
python/CuTeDSL/cutlass/base_dsl/runtime/jit_arg_adapters.py:174
↓ 5 callersMethodget_shape_A
Get A extents. fprop: A extents array contains [N,D,H,W,C]. Turn that into ((W,H,D,N), (C)) dgrad: A extents array contains [N,Z,P,Q,K]. Turn that int
include/cutlass/conv/convnd_problem_shape.hpp:402
↓ 5 callersFunctionget_sm_version
Get the SM (compute capability) version of a CUDA device.
examples/python/CuTeDSL/cute/blackwell/kernel/rmsnorm/rmsnorm.py:106
↓ 5 callersFunctionget_smem_layout_atom_ab
Simple heuristics to select the optimal SMEM layout atom based on the majorness, the data type, and the major mode size. :param major_mode: T
python/CuTeDSL/cutlass/utils/blackwell_helpers.py:612
↓ 5 callersFunctionget_swizzle_portion
include/cute/swizzle_layout.hpp:133
↓ 5 callersMethodget_unmasked_trip_count
Calculate the number of unmasked trips for the current block. This represents the number of trips that don't require special
examples/python/CuTeDSL/utils/fmha_helpers.py:836
↓ 5 callersFunctionget_wgmma_level_from_global_level
(global_level: int)
python/cutlass_library/sm90_utils.py:90
↓ 5 callersMethodinitialize
Initializes the workspace
tools/library/src/gemm_operation.h:277
↓ 5 callersFunctioninitialize_tensor
test/unit/transform/device/testbed_sparse_gemm_compressor.hpp:99
↓ 5 callersMethodinverse
include/cutlass/layout/tensor_op_multiplicand_sm70.h:376
↓ 5 callersMethodinverse
include/cutlass/layout/tensor_op_multiplicand_sm75.h:303
↓ 5 callersMethodir_value
Return the underlying MLIR vector value.
python/CuTeDSL/cutlass/base_dsl/_mlir_helpers/arith.py:1191
↓ 5 callersFunctionis_constexpr_field
Check if a field is a constexpr field.
python/CuTeDSL/cutlass/base_dsl/utils/tree_utils.py:170
↓ 5 callersFunctionis_cupy_tensor
(inp)
python/cutlass_cppgen/utils/datatypes.py:124
↓ 5 callersMethodis_valid
examples/88_hopper_fmha/kernel/fmha_tile_scheduler.hpp:67
↓ 5 callersFunctionitofp
( src: ir.Value, signed: Union[bool, None], res_elem_type: ir.Type, *, loc: Optional[ir.Lo
python/CuTeDSL/cutlass/base_dsl/_mlir_helpers/arith.py:185
↓ 5 callersFunctionleft_inverse
include/cute/layout.hpp:1322
↓ 5 callersMethodload
Loads a fragment from memory at the location pointed to by the iterator.
include/cutlass/gemm/warp/mma_tensor_op_tile_iterator_sm80.h:2359
↓ 5 callersMethodload_ab_init
include/cutlass/gemm/collective/sm100_mma_warpspecialized_blockwise_scaling.hpp:641
↓ 5 callersMethodload_tail
Perform a Producer Epilogue to prevent early exit of blocks in a Cluster
examples/63_hopper_gemm_with_weight_prefetch/collective/sm90_mma_tma_gmma_ss_warpspecialized_with_prefetch.hpp:629
↓ 5 callersFunctionload_with_byte_offset
Loads a fragment from memory
examples/41_fused_multi_head_attention/iterators/epilogue_predicated_tile_iterator.h:338
↓ 5 callersFunctionload_with_pointer_offset
Loads a fragment from memory with additional logical offset
include/cutlass/gemm/warp/mma_tensor_op_tile_iterator_sm80.h:301
↓ 5 callersMethodload_with_pointer_offset
Loads a fragment from memory
examples/41_fused_multi_head_attention/iterators/predicated_tile_iterator_residual_last.h:874
↓ 5 callersMethodload_with_pointer_offset
Loads a fragment from memory with additional logical offset
include/cutlass/gemm/warp/mma_tensor_op_tile_iterator_sm80.h:2366
↓ 5 callersMethodload_with_pointer_offset
Loads a fragment from memory
include/cutlass/transform/threadblock/regular_tile_iterator_tensor_op_sm70.h:507
↓ 5 callersMethodload_with_pointer_offset
Loads a fragment from memory
include/cutlass/transform/threadblock/predicated_tile_iterator.h:1459
↓ 5 callersMethodload_with_pointer_offset
Loads a fragment
include/cutlass/transform/threadblock/regular_tile_iterator_pitch_linear.h:339
↓ 5 callersFunctionmax
include/cute/numeric/math.hpp:48
↓ 5 callersFunctionnormalize_field_to_ir_name
Normalize a field specifier to its IR logical field name. Accepted inputs: - Enum value present in admissible_fields (must expose _to_i
python/CuTeDSL/cutlass/cute/nvgpu/common.py:111
↓ 5 callersFunctionoffs_to_group_sizes
Convert cumulative end offsets to per-group sizes.
examples/python/CuTeDSL/cute/blackwell/kernel/moe/torch_scaled_grouped_mm.py:2233
↓ 5 callersMethodoperations
Returns whether the provided operation class supports the provided data type combination :param op_class: operation class to conside
python/cutlass_cppgen/library_defaults.py:505
↓ 5 callersMethodpack_values_to_alloca
Pack values to an alloca that lays out in the order of the values. Parameters ---------- current_block : ir.Block
python/CuTeDSL/cutlass/base_dsl/tvm_ffi_builder/mlir_builder.py:435
↓ 5 callersFunctionprint_tensor
include/cute/util/print_tensor.hpp:102
↓ 5 callersMethodprocedural_name
The full procedural name indicates architecture, extended name, tile size, and layout.
python/cutlass_library/conv3d_operation.py:145
↓ 5 callersFunctionproducer_get_barrier
include/cutlass/pipeline/sm100_pipeline.hpp:1072
↓ 5 callersFunctionproduct
( a: Union[IntTuple, Shape], *, loc: Optional[ir.Location] = None, ip: Optional[ir.InsertionPo
python/CuTeDSL/cutlass/cute/tuple.py:116
↓ 5 callersFunctionproduct_each
( a: IntTuple, *, loc: Optional[ir.Location] = None, ip: Optional[ir.InsertionPoint] = None, )
python/CuTeDSL/cutlass/cute/tuple.py:190
↓ 5 callersMethodqualified_name
Returns a string name for debugging
tools/profiler/include/cutlass/profiler/problem_space.h:215
↓ 5 callersMethodraise_error_and_return
Raise an error and return -1. Instead of concatenating parts at compile time, we define each part as a global string and call a helpe
python/CuTeDSL/cutlass/base_dsl/tvm_ffi_builder/tvm_ffi_builder.py:723
↓ 5 callersMethodraw
Obtains raw bits
include/cutlass/tfloat32.h:161
↓ 5 callersFunctionrelatively_equal_float
include/cutlass/relatively_equal.h:57
↓ 5 callersMethodremove_edge
Remove edge src -> dst
python/cutlass_cppgen/backend/evt/ir/dag_ir.py:108
↓ 5 callersMethodreturn_
Create a return statement.
python/CuTeDSL/cutlass/base_dsl/tvm_ffi_builder/mlir_builder.py:300
↓ 5 callersMethodround
tools/profiler/include/cutlass/profiler/problem_space.h:360
↓ 5 callersMethodrun
Runs the kernel
tools/library/src/gemm_operation.h:299
↓ 5 callersMethodrun
Executes one test
test/unit/gemm/device/gemm_testbed_3x_ptr_array.hpp:2185
↓ 5 callersMethodrun
Executes one test
test/unit/conv/device/conv3d_with_broadcast_testbed.h:294
↓ 5 callersMethodrun
Executes one test
test/unit/conv/device/conv3d_testbed.h:212
↓ 5 callersMethodrun
(self)
examples/python/CuTeDSL/cute/blackwell/kernel/moe/torch_grouped_mm.py:1907
↓ 5 callersMethodset_kgroup_index
Notify the iterator which k-group it is currently pointing to. This does not advance the iterator. Rather, it overrides its internal tracking with co
include/cutlass/gemm/warp/mma_tensor_op_tile_iterator_sm80.h:2427
↓ 5 callersMethodset_residual_tile
examples/41_fused_multi_head_attention/iterators/predicated_tile_access_iterator_residual_last.h:802
↓ 5 callersMethodset_sign_bit
include/cutlass/exmy_base.h:523
↓ 5 callersMethodsigmoid
Sigmoid activation function: f(x) = 1 / (1 + exp(-x))
examples/python/CuTeDSL/cute/blackwell/efc/common_efc.py:1265
↓ 5 callersMethodsignificand_hidden_bits
include/cutlass/exmy_base.h:629
↓ 5 callersFunctionsin
Compute element-wise sine of the input tensor. :param a: Input tensor (in radians) :type a: Union[TensorSSA, Numeric] :param fastmath: En
python/CuTeDSL/cutlass/cute/math.py:554
↓ 5 callersFunctionsize
include/cute/int_tuple.hpp:265
↓ 5 callersMethodsoftmax
Softmax for one k-tile. Updates the related pipeline states and returns the computed results. :param common_params: The common parameters
examples/python/CuTeDSL/cute/blackwell/kernel/attention/mla/mla_decode_fp8.py:2414
↓ 5 callersMethodsoftmax
Compute softmax on attention scores from QK matrix multiplication. This method handles the softmax computation for either the first or second
examples/python/CuTeDSL/cute/blackwell/kernel/attention/fmha/fmha.py:1797
↓ 5 callersMethodstore_init
include/cutlass/epilogue/collective/detail.hpp:406
↓ 5 callersFunctionstore_with_byte_offset
Stores a fragment to memory
examples/41_fused_multi_head_attention/iterators/epilogue_predicated_tile_iterator.h:411
↓ 5 callersFunctionstore_with_pointer_offset
Store
include/cutlass/epilogue/warp/tile_iterator_simt.h:195
↓ 5 callersFunctionstore_with_pointer_offset
Store
include/cutlass/epilogue/warp/tile_iterator_tensor_op_mixed.h:203
↓ 5 callersMethodstore_with_pointer_offset
Store a fragment to memory
examples/41_fused_multi_head_attention/iterators/predicated_tile_iterator_residual_last.h:892
↓ 5 callersMethodstore_with_pointer_offset
Stores a fragment to memory at the location pointed to by the iterator
include/cutlass/gemm/warp/mma_tensor_op_tile_iterator_wmma.h:761
↓ 5 callersMethodstore_with_pointer_offset
Store a fragment to memory
include/cutlass/transform/threadblock/regular_tile_iterator_tensor_op_sm70.h:519
↓ 5 callersMethodstore_with_pointer_offset
Store a fragment to memory
include/cutlass/transform/threadblock/predicated_tile_iterator.h:1477
↓ 5 callersMethodstore_with_pointer_offset
Stores a fragment
include/cutlass/transform/threadblock/regular_tile_iterator_pitch_linear.h:357
↓ 5 callersFunctionstrided_dgrad_tile_m_per_filter
/////////////////////////////////////////////////////////////////////////////////////////////// Strided dgrad helper functions
include/cutlass/conv/conv2d_problem_size.h:617
↓ 5 callersMethodstructure_sparse_zero_mask_fill
include/cutlass/transform/kernel/sparse_gemm_compressor.hpp:191
↓ 5 callersMethodtensor
Return the output tensor (concept: cutlass_cppgen.backend.evt.ir.tensor)
python/cutlass_cppgen/backend/evt/ir/node.py:184
↓ 5 callersMethodtensormaps_cp_fence_release
include/cutlass/gemm/collective/sm90_mma_array_tma_gmma_ss_warpspecialized.hpp:735
↓ 5 callersFunctiontest_distribute
test/unit/cute/core/domain_distribute.cpp:45
↓ 5 callersMethodtest_wait
include/cutlass/arch/barrier.h:360
↓ 5 callersMethodtile_descriptions
Returns a list of valid tile descriptions for the operations :returns: list of valid tile descriptions for the operations :r
python/cutlass_cppgen/op/gemm.py:405
↓ 5 callersFunctiontrait_ratio
include/cute/numeric/integral_ratio.hpp:292
↓ 5 callersFunctionumma_arrive
include/cutlass/arch/barrier.h:796
↓ 5 callersFunctionumma_arrive_multicast_2x1SM
UMMA arrive for MMA_2x1SM + TMA_LOAD_MULTICAST combination
include/cutlass/arch/barrier.h:847
↓ 5 callersMethodupdate_arguments_base
Constructs the arguments structure given the configuration and arguments
tools/library/src/grouped_gemm_operation_3x.hpp:153
↓ 5 callersMethodvisit
include/cutlass/epilogue/fusion/sm90_visitor_tma_warpspecialized.hpp:597
↓ 5 callersFunctionyield_out
Generate a yield operation. It it used to return values from a loop, if-else, or while region.
python/CuTeDSL/cutlass/cutlass_dsl/cutlass.py:2236
← previousnext →1,301–1,400 of 19,553, ranked by callers