MCPcopy Create free account

hub / github.com/NVIDIA/cutlass / functions

Functions19,553 in github.com/NVIDIA/cutlass

↓ 6 callersMethodget_grid_shape
Computes the grid shape
include/cutlass/conv/device/conv_universal_adapter.hpp:174
↓ 6 callersMethodget_ir_values
Get all ir.Values from this leaf (may be multiple for DynamicExpression).
python/CuTeDSL/cutlass/base_dsl/leaf_utils.py:174
↓ 6 callersMethodgood
tools/profiler/include/cutlass/profiler/performance_report.h:91
↓ 6 callersFunctionh
(j)
docs/jquery.js:53
↓ 6 callersFunctionhas_underscore
(a: XTuple)
python/CuTeDSL/cutlass/cute/core.py:1980
↓ 6 callersFunctionis_block_scaled
(gemm_kind)
python/cutlass_library/library.py:324
↓ 6 callersFunctionis_int
(x)
python/pycute/int_tuple.py:43
↓ 6 callersFunctionis_tma_epilogue
(epilogue_schedule_type)
python/cutlass_library/library.py:984
↓ 6 callersMethodis_valid
examples/112_blackwell_ssd/kernel/sm100_ssd_tile_scheduler.hpp:95
↓ 6 callersMethodis_valid
examples/111_hopper_ssd/kernel/sm90_ssd_tile_scheduler.hpp:95
↓ 6 callersMethodisinstance
Check if an ``ir.Value`` matches this parameterized vector type.
python/CuTeDSL/cutlass/base_dsl/_mlir_helpers/arith.py:1098
↓ 6 callersMethodload_ffi_any_array_item_v_ptr
Get the v_ptr from the index-th field of tvm_ffi_any_type struct. Semantics as follows: .. code-block:: c void* get_v_p
python/CuTeDSL/cutlass/base_dsl/tvm_ffi_builder/tvm_ffi_builder.py:366
↓ 6 callersMethodload_sf
include/cutlass/gemm/collective/sm100_mma_warpspecialized_blockwise_scaling.hpp:900
↓ 6 callersMethodload_with_ell_index
include/cutlass/transform/threadblock/ell_predicated_tile_iterator.h:888
↓ 6 callersMethodload_with_ell_index_fast
include/cutlass/transform/threadblock/ell_predicated_tile_iterator.h:893
↓ 6 callersFunctionload_with_pointer_offset
Loads a fragment from memory
include/cutlass/conv/threadblock/predicated_scale_bias_vector_iterator.h:185
↓ 6 callersMethodload_with_pointer_offset
Loads a fragment
include/cutlass/transform/threadblock/regular_tile_iterator_pitch_linear_2dthreadtile.h:148
↓ 6 callersFunctionlocal_partition
include/cute/tensor_impl.hpp:1077
↓ 6 callersFunctionmake_im2col_tma_copy
include/cute/atom/copy_traits_sm90_im2col.hpp:822
↓ 6 callersFunctionmake_iterator
tools/library/src/reference/block_scaled_gemm_reference_operation.h:61
↓ 6 callersFunctionmake_layout
include/cute/swizzle_layout.hpp:68
↓ 6 callersFunctionmake_rmem_tensor
Creates a tensor in register memory with the specified layout/shape and data type. This function allocates a tensor in register memory (rmem) usu
python/CuTeDSL/cutlass/cute/tensor.py:835
↓ 6 callersFunctionmark_values
(*args_list)
test/utils/test_sharding.py:201
↓ 6 callersMethodmax
(cls)
python/CuTeDSL/cutlass/base_dsl/typing.py:582
↓ 6 callersFunctionmerge_desc_sorted_arrays
include/cutlass/epilogue/fusion/sm90_visitor_topk_softmax.hpp:228
↓ 6 callersMethodmma_partition_c
(self, tiled_mma, tile_shape_mnk, tmem_acc_ptr, acc_stages)
examples/python/CuTeDSL/cute/blackwell/kernel/attention/mamba2_ssd/mamba2_ssd.py:2884
↓ 6 callersMethodnvcc
(self)
python/cutlass_cppgen/backend/compiler.py:169
↓ 6 callersMethodoperand_B_ref
Returns a TensorRef to the B operand
examples/41_fused_multi_head_attention/gemm/mma_from_smem.h:225
↓ 6 callersFunctionoperator*
include/cutlass/half.h:802
↓ 6 callersFunctionoperator+
include/cutlass/half.h:775
↓ 6 callersFunctionoperator-
include/cutlass/half.h:784
↓ 6 callersFunctionoperator/
include/cutlass/half.h:811
↓ 6 callersMethodpartition
include/cute/atom/partitioner.hpp:83
↓ 6 callersMethodpartition_S
( self, src: Tensor, *, loc: Optional[ir.Location] = None, ip: Optiona
python/CuTeDSL/cutlass/cute/atom.py:913
↓ 6 callersMethodpartition_accumulator_shape
Construct A Single Stage's Accumulator Shape
include/cutlass/gemm/collective/sm100_mma_warpspecialized.hpp:453
↓ 6 callersFunctionpermute
(x, indices: tuple)
python/cutlass_cppgen/epilogue/evt_ops.py:87
↓ 6 callersMethodpostreduce
After reduce call, before smem async fence. Smem stores usually performed here. Upon exit, all smem stores for TMA must have been issued
include/cutlass/epilogue/fusion/sm90_visitor_tma_warpspecialized.hpp:242
↓ 6 callersMethodprefetch_tma_descriptors
Issue Tma Descriptor Prefetch -- ideally from a single thread for best performance
include/cutlass/gemm/collective/sm100_mma_warpspecialized.hpp:446
↓ 6 callersMethodprevisit
Before visit callback. Smem broadcasts usually performed here. Upon entry, all producer loads for this subtile are completed and visible.
include/cutlass/epilogue/fusion/sm90_visitor_tma_warpspecialized.hpp:206
↓ 6 callersMethodprocedural_name
The full procedural name indicates architecture, extended name, tile size, and layout.
python/cutlass_library/conv2d_operation.py:171
↓ 6 callersFunctionprologue
examples/41_fused_multi_head_attention/gemm/mma_from_smem.h:528
↓ 6 callersMethodpromote_if_needed
include/cutlass/gemm/collective/fp8_accumulation.hpp:170
↓ 6 callersFunctionraw
Accesses raw internal state
include/cutlass/exmy_base.h:1058
↓ 6 callersMethodreadExpect
(self, what: str)
examples/41_fused_multi_head_attention/piped_subprocess.py:124
↓ 6 callersFunctionrecast_ptr
( ptr: Pointer, swizzle_: Optional[Swizzle] = None, dtype: Optional[Type[Numeric]] = None, loc
python/CuTeDSL/cutlass/cute/core.py:4287
↓ 6 callersMethodreduce
Perform reduce on selected modes with given predefined reduction op. :param op: The reduction operator to use (operator.add or opera
python/CuTeDSL/cutlass/cute/tensor.py:2516
↓ 6 callersMethodrequires_separate_reduction
include/cutlass/gemm/kernel/sm100_tile_scheduler.hpp:648
↓ 6 callersMethodrun
Executes one test
test/unit/conv/device/conv2d_with_broadcast_testbed.h:281
↓ 6 callersFunctionsetClassAttr
(elem,attr)
docs/search/search.js:716
↓ 6 callersFunctionslice_
(crd: Union[None, tuple, int], trg: Union[tuple, int])
python/pycute/int_tuple.py:204
↓ 6 callersMethodstore_with_pointer_offset
Stores a fragment
include/cutlass/transform/threadblock/regular_tile_iterator_pitch_linear_2dthreadtile.h:189
↓ 6 callersMethodthrfrg_A
include/cute/atom/mma_atom.hpp:289
↓ 6 callersMethodthrfrg_B
include/cute/atom/mma_atom.hpp:328
↓ 6 callersFunctiontile_atom_to_shape_SF
A helper function to get dynamic SFA/SFB layout by filling dynamic A/B shape to the scale factor atom layout. :param Shape: The shape of the
python/CuTeDSL/cutlass/utils/blockscaled_layout.py:66
↓ 6 callersFunctiontma_store_arrive
Indicate arrival of warp issuing TMA_STORE
include/cute/arch/copy_sm90_tma.hpp:1224
↓ 6 callersMethodto_rmem_tensor
Pack work tile info fields into an rmem tensor of shape (4,) for vectorized smem copy.
examples/python/CuTeDSL/cute/blackwell/kernel/moe/moe_persistent_scheduler.py:119
↓ 6 callersFunctionto_underlying_arguments
include/cutlass/conv/collective/sm100_implicit_gemm_umma_warpspecialized.hpp:382
↓ 6 callersMethodtype
(self)
python/CuTeDSL/cutlass/cute/atom.py:196
↓ 6 callersMethodvalue
Get the scale value. :return: The scale value
python/CuTeDSL/cutlass/cute/core.py:807
↓ 6 callersFunctionwrap
Wraps the input into a tuple if not a tuple.
python/CuTeDSL/cutlass/cute/tuple.py:31
↓ 6 callersFunctionzipped_divide
(layoutA, layoutB)
python/pycute/layout.py:343
↓ 5 callersFunctionCreateConvOperator3x
Create zero or more CUTLASS 3 two-dimensional convolution operators. Create a CUTLASS 3 two-dimensional convolution operator for all feasible
python/cutlass_library/generator.py:1071
↓ 5 callersFunctionCudaToolkitVersionSatisfies
(semantic_ver_string, major, minor, patch = 0)
python/cutlass_library/sm90_utils.py:57
↓ 5 callersFunctionCutlassUnitTestProblemCount
test/unit/common/filter_architecture.cpp:146
↓ 5 callersFunctionEncodeOperator
test/unit/conv/cache_testbed_output.h:256
↓ 5 callersMethodGemmCoord
Default ctor
include/cutlass/gemm_coord.h:108
↓ 5 callersMethod__extract_mlir_attributes__
(self)
python/CuTeDSL/cutlass/base_dsl/_mlir_helpers/arith.py:1145
↓ 5 callersMethod__init__
( self, x: Union[bool, int, float, ir.Value, "Integer", "Float"], *, loc: Opti
python/CuTeDSL/cutlass/base_dsl/typing.py:1878
↓ 5 callersMethod_build_select_and_assign
( self, *, name: str, test: ast.expr, body: ast.expr, orelse:
python/CuTeDSL/cutlass/base_dsl/ast_preprocessor.py:1473
↓ 5 callersMethod_get_problem_for_group
Load gemm problem (m,n,k,l) for the specified group from global memory to register. :param problem_shape_mnkl: Tensor in global memo
python/CuTeDSL/cutlass/utils/grouped_gemm_persistent_tile_scheduler.py:417
↓ 5 callersFunction_normalize_ptr
Helper function to normalize pointer types to MLIR ir.Value. Supports: - ir.Value (LLVM pointer): returned as-is - cute.ptr (_Pointe
python/CuTeDSL/cutlass/cute/arch/nvvm_wrappers.py:2282
↓ 5 callersMethod_reset_epilogue_functor_alignment
Reset the alignment of the current epilogue functor based on alignment C
python/cutlass_cppgen/op/op.py:321
↓ 5 callersMethod_reset_options
Resets the kernel options based on cc :param cc: compute capability to reset to :type cc: int
python/cutlass_cppgen/op/op.py:124
↓ 5 callersMethodaccum_init
include/cutlass/gemm/collective/sm100_mma_warpspecialized_blockwise_scaling.hpp:807
↓ 5 callersFunctionadd_free_func_and_tensor
(free_func, tensor)
examples/python/CuTeDSL/cute/blackwell/kernel/distributed/distributed_gemm_reduce_scatter_blackwell.py:2391
↓ 5 callersMethodadd_node
(self, node)
python/cutlass_cppgen/backend/evt/frontend/frontend_base.py:132
↓ 5 callersFunctionadd_pointer_offset
include/cutlass/subbyte_reference.h:1330
↓ 5 callersMethodadd_pointer_offset
Adds a pointer offset in units of Element
examples/41_fused_multi_head_attention/iterators/predicated_tile_access_iterator_residual_last.h:808
↓ 5 callersMethodadd_pointer_offset
Adds a pointer offset in units of Element
include/cutlass/transform/threadblock/regular_tile_iterator_tensor_op_sm70.h:477
↓ 5 callersMethodadd_pointer_offset
Adds a pointer offset in units of Element
include/cutlass/transform/threadblock/predicated_tile_access_iterator.h:988
↓ 5 callersMethodadd_tile_offset
Adds a tile offset
include/cutlass/transform/threadblock/regular_tile_iterator_tensor_op_sm70.h:483
↓ 5 callersMethodadd_tile_offset
Advances an iterator along logical dimensions of matrix in units of whole tiles
include/cutlass/transform/threadblock/predicated_tile_access_iterator.h:995
↓ 5 callersMethodalign_offset
Return the round-up offset up to the next multiple of align.
python/CuTeDSL/cutlass/cute/core.py:5953
↓ 5 callersMethodand_
Create a logical AND operation between two values.
python/CuTeDSL/cutlass/base_dsl/tvm_ffi_builder/mlir_builder.py:190
↓ 5 callersFunctionany_
Logical OR operation for any element in an iterable. Returns True if any element in the iterable is truthy, otherwise False. This is the DSL
python/CuTeDSL/cutlass/cutlass_dsl/cutlass.py:2160
↓ 5 callersFunctionappend
include/cute/algorithm/tuple_algorithms.hpp:779
↓ 5 callersFunctionapply_variable_length_offset
examples/77_blackwell_fmha/collective/fmha_fusion.hpp:361
↓ 5 callersFunctionc
(e,n,o)
python/docs/_static/scripts/furo.js:2
↓ 5 callersMethodcan_implement
Determines whether the GEMM can execute the given problem.
examples/13_two_tensor_op_fusion/device/b2b_gemm.h:201
↓ 5 callersFunctionceil_div
include/cute/numeric/arithmetic_tuple.hpp:349
↓ 5 callersMethodcheck_valid_work_for_seqlen_q
Check if the current work index is valid for the given query sequence length. This method verifies that the current work tile index
examples/python/CuTeDSL/utils/fmha_helpers.py:180
↓ 5 callersMethodclear
Clear the cache.
python/CuTeDSL/cutlass/base_dsl/cache_helpers.py:493
↓ 5 callersMethodclear
< Efficiently disables all accesses guarded by mask
include/cutlass/epilogue/threadblock/predicated_tile_iterator.h:171
↓ 5 callersFunctioncongruent
include/cute/int_tuple.hpp:439
↓ 5 callersMethodconstruct
Constructs a ``cutlass_cppgen.backend.Conv2dOperation`` based on the input parameters and current kernel specification of the ``Conv2
python/cutlass_cppgen/op/conv.py:545
↓ 5 callersMethodconsumer_release_from_umma
examples/112_blackwell_ssd/utils/pipeline.h:172
↓ 5 callersFunctioncontinue_current_work
Returns whether the current work_tile_info passed in should continue to be used. This occurs only in the stream-K decomposition with stream-K work uni
include/cutlass/gemm/kernel/sm90_tile_scheduler_stream_k.hpp:299
↓ 5 callersFunctioncvtf
( src: ir.Value, res_elem_type: ir.Type, *, loc: Optional[ir.Location] = None, ip: Optiona
python/CuTeDSL/cutlass/base_dsl/_mlir_helpers/arith.py:114
↓ 5 callersMethodemit
(self)
python/cutlass_cppgen/backend/gemm_operation.py:903
↓ 5 callersMethodenter_control_flow_scope
Context manager for entering a new dynamic control-flow scope. This scope rule diverge from Python's local scope, variables defined
python/CuTeDSL/cutlass/base_dsl/ast_preprocessor.py:198
← previousnext →1,201–1,300 of 19,553, ranked by callers