MCPcopy Create free account

hub / github.com/SHI-Labs/NATTEN / functions

Functions2,059 in github.com/SHI-Labs/NATTEN

↓ 1 callersFunctioniterable_to_static_cute_tuple
(shape_in)
scripts/autogen_hopper_fmha_bwd.py:184
↓ 1 callersFunctioniterable_to_static_cute_tuple
(shape_in)
scripts/autogen_hopper_fmha.py:193
↓ 1 callersFunctionkernel_type_to_tag
(kernel_type: KernelType)
scripts/autogen_hopper_fmha.py:210
↓ 1 callersFunctionkernel_type_to_tag
(kernel_type: KernelType)
scripts/autogen_hopper_fna.py:174
↓ 1 callersFunctionlayout_op_mk_v
csrc/include/natten/cuda/fna_hopper/collective/fna_common.hpp:184
↓ 1 callersFunctionlayout_op_mk_v
csrc/include/natten/cuda/fmha_hopper/collective/fmha_common.hpp:141
↓ 1 callersFunctionload
Loads a fragment from memory at the location pointed to by the iterator.
csrc/include/natten/cuda/fmha/iterators/warp_iterator_from_smem.h:259
↓ 1 callersFunctionload
Loads a fragment from memory at the location pointed to by the iterator.
csrc/include/natten/cuda/fna/iterators/warp_iterator_from_smem.h:259
↓ 1 callersMethodload
csrc/include/natten/cuda/fmha_blackwell/kernel/sm100_fmha_bwd_kernel_tma_warpspecialized.hpp:555
↓ 1 callersMethodload
csrc/include/natten/cuda/fna_blackwell/kernel/sm100_fna_bwd_kernel_tma_warpspecialized.hpp:587
↓ 1 callersMethodload_maybe_q
csrc/include/natten/cuda/fmha_hopper/collective/fmha_collective_tma_warpspecialized.hpp:444
↓ 1 callersMethodload_with_byte_offset
Loads a fragment from memory
csrc/include/natten/cuda/fna/iterators/predicated_tile_iterator.h:795
↓ 1 callersMethodload_with_pointer_offset
Loads a fragment from memory
csrc/include/natten/cuda/fna/iterators/predicated_tile_iterator.h:789
↓ 1 callersFunctionmain
()
scripts/testing/monitor.py:246
↓ 1 callersMethodmake_additional_kv_tensors
( self, device, requires_grad: bool, heads_last: bool )
src/natten/profiling_utils/problem.py:188
↓ 1 callersMethodmake_qkvo_tensors
( self, device, requires_grad: bool, heads_last: bool, flatten: bool = False )
src/natten/profiling_utils/problem.py:168
↓ 1 callersFunctionmeasure_natten_runtime
( problem: Problem, warmup_steps: int, backend: Optional[str] = None, fmha_backend: Optional[s
src/natten/profiling_utils/profiling.py:68
↓ 1 callersMethodmma
csrc/include/natten/cuda/fmha_blackwell/kernel/sm100_fmha_bwd_kernel_tma_warpspecialized.hpp:839
↓ 1 callersMethodmma
csrc/include/natten/cuda/fna_blackwell/kernel/sm100_fna_bwd_kernel_tma_warpspecialized.hpp:872
↓ 1 callersFunctionmod_tuple
csrc/include/natten/cuda/reference/mask.hpp:201
↓ 1 callersMethodo_strideM
csrc/include/natten/cuda/fmha/kernel_backward.h:636
↓ 1 callersFunctionopt_progress_bar
(fn, total)
src/natten/profiling_utils/pretty_printer.py:27
↓ 1 callersFunctionparse_env_int
(env_var: str, default: int)
src/natten/utils/environment.py:43
↓ 1 callersFunctionprefetch
csrc/include/natten/cuda/fmha/iterators/epilogue_predicated_tile_iterator.h:298
↓ 1 callersFunctionprofile
( batch_size: int, heads: int, input_size, dim: int, dim_value: int, window_size: Dime
src/natten/profiler.py:513
↓ 1 callersFunctionread_progress
Read progress JSON for a test. Returns dict or None.
scripts/testing/monitor.py:111
↓ 1 callersMethodreduce
csrc/include/natten/cuda/fmha_hopper/collective/fmha_collective_tma_warpspecialized.hpp:480
↓ 1 callersMethodrelease
csrc/include/natten/cuda/fna/kernel_backward.h:194
↓ 1 callersFunctionsdpa
(q: Tensor, k: Tensor, v: Tensor, backend: str)
src/natten/profiling_utils/ops.py:30
↓ 1 callersFunctionsdpa_ref
( q: Tensor, k_list: list, v_list: list, do: Tensor, backend: str, )
tests/test_attn_merge.py:92
↓ 1 callersFunctionsdpa_split
( q: Tensor, k_list: list, v_list: list, do: Tensor, backend: str, torch_compile: bool
tests/test_attn_merge.py:51
↓ 1 callersFunctionset_iteration_index
Overrides the internal iteration index
csrc/include/natten/cuda/fmha/iterators/predicated_tile_access_iterator_residual_last.h:286
↓ 1 callersMethodset_mask
Sets the predicate mask, overriding value stored in predicate iterator
csrc/include/natten/cuda/fna/iterators/predicated_tile_iterator_residual_last.h:565
↓ 1 callersFunctionshould_include
(filename)
setup.py:246
↓ 1 callersMethodsoftmax
csrc/include/natten/cuda/fmha_blackwell/collective/sm100_fmha_fwd_mainloop_tma_warpspecialized.hpp:795
↓ 1 callersMethodsoftmax
csrc/include/natten/cuda/fna_blackwell/collective/sm100_fna_fwd_mainloop_tma_warpspecialized.hpp:864
↓ 1 callersMethodstep_interleave_begin
csrc/include/natten/cuda/fna_hopper/collective/fna_collective_softmax.hpp:201
↓ 1 callersMethodstep_interleave_begin
csrc/include/natten/cuda/fmha_hopper/collective/fmha_collective_softmax.hpp:165
↓ 1 callersMethodstore
csrc/include/natten/cuda/fna_blackwell/kernel/sm100_fna_bwd_kernel_tma_warpspecialized.hpp:1216
↓ 1 callersMethodstore_with_byte_offset
Store a fragment to memory
csrc/include/natten/cuda/fna/iterators/predicated_tile_iterator.h:813
↓ 1 callersMethodstore_with_pointer_offset
Store a fragment to memory
csrc/include/natten/cuda/fna/iterators/predicated_tile_iterator.h:807
↓ 1 callersFunctionstrip_str_name
(name: str)
src/natten/profiling_utils/formatting.py:223
↓ 1 callersFunctionsub_tuple
(X: tuple, Y: tuple)
src/natten/utils/tuples.py:41
↓ 1 callersMethodto_str_name
(self)
src/natten/profiling_utils/formatting.py:151
↓ 1 callersFunctionto_string
csrc/include/natten/cuda/hopper_fmha_fna.h:40
↓ 1 callersFunctiontoken_permute_cutlass
( tensor: Tensor, tile_shape: DimensionType, dilation: DimensionType, flip_tiled_dims, )
src/natten/token_permute/cutlass_impl.py:219
↓ 1 callersFunctiontoken_permute_torch
( tensor: Tensor, tile_shape: DimensionType, dilation: DimensionType, flip_tiled_dims: bool, )
src/natten/token_permute/torch_impl.py:309
↓ 1 callersFunctiontoken_unpermute_cutlass
( tensor: Tensor, token_layout_shape: DimensionType, tile_shape: DimensionType, dilation: Dime
src/natten/token_permute/cutlass_impl.py:255
↓ 1 callersFunctiontoken_unpermute_torch
( tensor: Tensor, token_layout: DimensionType, tile_shape: DimensionType, dilation: DimensionT
src/natten/token_permute/torch_impl.py:332
↓ 1 callersFunctionvalid
Returns whether access is valid or not
csrc/include/natten/cuda/fmha/iterators/predicated_tile_access_iterator_residual_last.h:436
↓ 1 callersFunctionvalid
Returns whether access is valid or not
csrc/include/natten/cuda/fna/iterators/predicated_tile_access_iterator_residual_last.h:434
↓ 1 callersFunctionvalid
Returns whether access is valid or not
csrc/include/natten/cuda/fna/iterators/predicated_tile_access_iterator.h:881
↓ 1 callersFunctionwriteFragsToGmem
csrc/include/natten/cuda/fna/kernel_backward.h:2556
↓ 1 callersFunctionwrite_combined_source_file
(path, filename, headers, kernels)
scripts/autogen_hopper_fna_bwd.py:218
↓ 1 callersFunctionwrite_combined_source_file
(path, filename, headers, sources)
scripts/autogen_fna.py:264
↓ 1 callersFunctionwrite_combined_source_file
(path, filename, headers, kernels)
scripts/autogen_blackwell_fmha_bwd.py:240
↓ 1 callersFunctionwrite_combined_source_file
(path, filename, headers, kernels)
scripts/autogen_blackwell_fna.py:211
↓ 1 callersFunctionwrite_combined_source_file
(path, filename, headers, kernels)
scripts/autogen_blackwell_fna_bwd.py:215
↓ 1 callersFunctionwrite_combined_source_file
(path, filename, headers, kernels)
scripts/autogen_reference_fna.py:242
↓ 1 callersFunctionwrite_combined_source_file
(path, filename, headers, kernels)
scripts/autogen_blackwell_fmha.py:212
↓ 1 callersFunctionwrite_combined_source_file
(path, filename, headers, kernels)
scripts/autogen_hopper_fmha_bwd.py:222
↓ 1 callersFunctionwrite_combined_source_file
(path, filename, headers, kernels)
scripts/autogen_hopper_fmha.py:258
↓ 1 callersFunctionwrite_combined_source_file
(path, filename, headers, kernels)
scripts/autogen_hopper_fna.py:253
↓ 1 callersFunctionwrite_combined_source_file
(path, filename, headers, sources)
scripts/autogen_fmha.py:237
MethodApplyExp
csrc/include/natten/cuda/fmha/epilogue/epilogue_thread_apply_logsumexp.h:136
MethodApplyExp
csrc/include/natten/cuda/fna/epilogue/epilogue_thread_apply_logsumexp.h:136
FunctionAssertDimsAre128BitAligned
csrc/include/natten/helpers.h:404
FunctionAssertOddKernelSize
csrc/include/natten/helpers.h:36
MethodCUTE_DEVICE Pow2
csrc/include/natten/cuda/fmha_blackwell/common/pow_2.hpp:46
MethodCUTE_DEVICE Pow2
csrc/include/natten/cuda/fna_blackwell/common/pow_2.hpp:46
FunctionCUTLASS_DEVICE accumApplyExpToSmem
NOTE(alih): we only apply the exp operator here; LSE is pre-fetched into shared memory, and subtracted when we check the NA mask and set invalid weigh
csrc/include/natten/cuda/fna/gemm/mma_from_smem.h:1773
FunctionCUTLASS_DEVICE accumToSmem
csrc/include/natten/cuda/fna/gemm/mma_from_smem.h:1687
FunctionCUTLASS_HOST_DEVICE CustomPredicatedTileAccessIterator
csrc/include/natten/cuda/fna/iterators/predicated_tile_access_iterator.h:521
FunctionCUTLASS_HOST_DEVICE PredicatedTileAccessIteratorResidualLast
csrc/include/natten/cuda/fmha/iterators/predicated_tile_access_iterator_residual_last.h:100
FunctionCUTLASS_HOST_DEVICE PredicatedTileAccessIteratorResidualLast
csrc/include/natten/cuda/fna/iterators/predicated_tile_access_iterator_residual_last.h:102
MethodCausalIndividualTileScheduler
csrc/include/natten/cuda/fmha_blackwell/kernel/fmha_causal_tile_scheduler.hpp:61
FunctionCheckArgs
csrc/include/natten/helpers.h:93
FunctionCheckArgsAgainstDim
csrc/include/natten/helpers.h:171
FunctionCheckIfBatchHeadsHeadDimMatch
csrc/include/natten/helpers.h:304
FunctionCheckIfPropertiesMatch
csrc/include/natten/helpers.h:217
FunctionCheckIfTensorShapesMatch
csrc/include/natten/helpers.h:239
FunctionCheckIfTensorShapesMatchExceptHeadDim
csrc/include/natten/helpers.h:257
FunctionCheckLogSumExp
csrc/include/natten/helpers.h:314
FunctionCheckLogSumExpHeadsFirst
csrc/include/natten/helpers.h:354
MethodCollectiveSoftmax
csrc/include/natten/cuda/fna_hopper/collective/fna_collective_softmax.hpp:45
MethodCollectiveSoftmax
csrc/include/natten/cuda/fmha_hopper/collective/fmha_collective_softmax.hpp:45
MethodCustomPredicatedTileAccessIterator
Default constructor
csrc/include/natten/cuda/fna/iterators/predicated_tile_access_iterator.h:631
MethodCustomPredicatedTileAccessIterator
Default constructor
csrc/include/natten/cuda/fna/iterators/predicated_tile_access_iterator.h:1006
MethodCustomPredicatedTileAccessIterator
Default constructor
csrc/include/natten/cuda/fna/iterators/predicated_tile_access_iterator.h:1308
MethodCustomPredicatedTileAccessIterator
Default constructor
csrc/include/natten/cuda/fna/iterators/predicated_tile_access_iterator.h:1531
FunctionCustomPredicatedTileAccessIterator operator++
Increment and return an instance to self.
csrc/include/natten/cuda/fna/iterators/predicated_tile_access_iterator.h:849
MethodCustomPredicatedTileAccessIteratorParams
csrc/include/natten/cuda/fna/iterators/predicated_tile_access_iterator.h:174
MethodCustomPredicatedTileAccessIteratorPredicates
Default constructor
csrc/include/natten/cuda/fna/iterators/predicated_tile_access_iterator.h:414
MethodCustomPredicatedTileIterator
Default constructor
csrc/include/natten/cuda/fna/iterators/predicated_tile_iterator.h:177
MethodCustomPredicatedTileIterator
Default constructor
csrc/include/natten/cuda/fna/iterators/predicated_tile_iterator.h:462
MethodCustomPredicatedTileIterator
Default constructor
csrc/include/natten/cuda/fna/iterators/predicated_tile_iterator.h:693
MethodCustomPredicatedTileIterator
Constructor
csrc/include/natten/cuda/fna/epilogue/predicated_tile_iterator.h:239
MethodCustomPredicatedTileIteratorParams
csrc/include/natten/cuda/fna/epilogue/predicated_tile_iterator_params.h:138
FunctionFnaRPBChecks
csrc/include/natten/helpers.h:61
MethodIndividualTileScheduler
csrc/include/natten/cuda/fmha_blackwell/kernel/fmha_tile_scheduler.hpp:50
← previousnext →701–800 of 2,059, ranked by callers