MCPcopy Create free account

hub / github.com/SHI-Labs/NATTEN / functions

Functions2,059 in github.com/SHI-Labs/NATTEN

↓ 94 callersMethodappend
(self, na_dim: int)
scripts/autogen_fna.py:311
↓ 70 callersMethoddebug
(self, *args, **kwargs)
src/natten/utils/log.py:124
↓ 60 callersFunctionprint
csrc/include/natten/cuda/fna_blackwell/common/pow_2.hpp:81
↓ 51 callersFunctionceil_div
csrc/include/natten/cuda/fna/gemm_kernel_utils.h:63
↓ 46 callersMethodget
Returns a pointer
csrc/include/natten/cuda/fna/iterators/predicated_tile_access_iterator.h:1597
↓ 46 callersFunctionmaybe_contiguous
(x)
src/natten/_libnatten/torch_wrappers.py:73
↓ 43 callersFunctionidx2crd
(index, shape)
src/natten/backends/flex.py:270
↓ 41 callersMethodadd_tile_offset
Advances an iterator along logical dimensions of matrix in units of whole tiles
csrc/include/natten/cuda/fmha/iterators/predicated_tile_access_iterator_residual_last.h:823
↓ 37 callersMethodadd_tile_offset
Advances an iterator along logical dimensions of matrix in units of whole tiles
csrc/include/natten/cuda/fna/iterators/predicated_tile_access_iterator.h:1591
↓ 37 callersMethodget
Returns a pointer
csrc/include/natten/cuda/fmha/iterators/predicated_tile_access_iterator_residual_last.h:829
↓ 34 callersMethodload
csrc/include/natten/cuda/fmha/gemm/mma_from_smem.h:300
↓ 33 callersFunctionwarp_uniform
csrc/include/natten/cuda/fmha/gemm_kernel_utils.h:215
↓ 31 callersFunctionceil_div
csrc/include/natten/cuda/fmha/gemm_kernel_utils.h:118
↓ 30 callersFunctionget_device_cc
(device: Optional[torch.device] = None)
src/natten/utils/device.py:41
↓ 27 callersMethodset_iteration_index
Overrides the internal iteration index
csrc/include/natten/cuda/fmha/iterators/predicated_tile_access_iterator_residual_last.h:805
↓ 26 callersMethodload
csrc/include/natten/cuda/fna/gemm/mma_from_smem.h:301
↓ 24 callersMethodset_iteration_index
csrc/include/natten/cuda/fna/iterators/predicated_tile_access_iterator.h:427
↓ 24 callersFunctiontoken_permute_operation
( tensor: Tensor, tile_shape: DimensionType, dilation: Optional[DimensionType] = None, flip_ti
src/natten/token_permute/frontend.py:42
↓ 23 callersMethodclear_mask
Clears the predicate set efficiently
csrc/include/natten/cuda/fmha/iterators/predicated_tile_iterator_residual_last.h:853
↓ 23 callersFunctionwarp_uniform
csrc/include/natten/cuda/fna/gemm_kernel_utils.h:160
↓ 20 callersMethodvalid
Returns whether access is valid or not
csrc/include/natten/cuda/fmha/iterators/predicated_tile_access_iterator_residual_last.h:884
↓ 19 callersMethodbackward
(ctx, grad_out: Tensor, grad_lse: Tensor)
src/natten/backends/fmha.py:121
↓ 18 callersMethodclear
< Efficiently disables all accesses guarded by mask
csrc/include/natten/cuda/fmha/iterators/epilogue_predicated_tile_iterator.h:162
↓ 17 callersMethodvalid
Returns whether access is valid or not
csrc/include/natten/cuda/fna/iterators/predicated_tile_access_iterator.h:480
↓ 16 callersMethodclear
< Efficiently disables all accesses guarded by mask
csrc/include/natten/cuda/fna/epilogue/predicated_tile_iterator.h:168
↓ 16 callersMethodclear_mask
< Efficiently disables all accesses guarded by mask
csrc/include/natten/cuda/fna/epilogue/predicated_tile_iterator.h:604
↓ 15 callersMethodadd_pointer_offset
Adds a pointer offset in units of Element
csrc/include/natten/cuda/fmha/iterators/predicated_tile_iterator_residual_last.h:822
↓ 15 callersMethodclear
csrc/include/natten/cuda/fna/kernel_backward.h:1198
↓ 15 callersMethodinit_state
csrc/include/natten/cuda/fna_hopper/collective/fna_collective_load.hpp:134
↓ 15 callersMethodinit_state
csrc/include/natten/cuda/fmha_hopper/collective/fmha_collective_load.hpp:167
↓ 15 callersFunctiontoken_unpermute_operation
( tensor: Tensor, token_layout_shape: DimensionType, tile_shape: DimensionType, dilation: Opti
src/natten/token_permute/frontend.py:94
↓ 14 callersFunction_make_logger
(name: str)
tests/test_log.py:26
↓ 14 callersFunctioncheck_all_args
( na_dim: int, kernel_size: Any, stride: Any, dilation: Any, is_causal: Any )
src/natten/utils/checks.py:473
↓ 14 callersFunctioncheck_tile_shape
( tile_shape: Any, )
src/natten/utils/checks.py:529
↓ 14 callersMethodset_residual_tile
csrc/include/natten/cuda/fmha/iterators/predicated_tile_iterator_residual_last.h:858
↓ 13 callersFunction_reset_everything
(random_seed: int = 42, torch_seed: int = 42)
tests/test_flex.py:58
↓ 13 callersMethodappend
(self, na_dim: int)
scripts/autogen_hopper_fna_bwd.py:275
↓ 13 callersMethodappend
(self, na_dim: int)
scripts/autogen_blackwell_fna.py:266
↓ 13 callersMethodappend
(self, na_dim: int)
scripts/autogen_blackwell_fna_bwd.py:270
↓ 13 callersMethodappend
(self, na_dim: int)
scripts/autogen_hopper_fna.py:310
↓ 13 callersFunctionis_cuda
(device: torch.device)
src/natten/utils/device.py:29
↓ 13 callersFunctionis_problem_shape_variable_length
csrc/include/natten/cuda/fmha_hopper/collective/fmha_varlen.hpp:86
↓ 12 callersFunction_run_empty_batch_kernel_check
(backend, use_compile, test_causal=False)
tests/test_varlen_recompile.py:184
↓ 12 callersFunctionis_fp8
(dtype: torch.dtype)
src/natten/utils/dtype.py:35
↓ 12 callersFunctiontuple_add
csrc/include/natten/cuda/fna_hopper/collective/fna_common.hpp:74
↓ 11 callersFunction_reset_everything
(random_seed: int = 42, torch_seed: int = 42)
tests/test_fna.py:47
↓ 11 callersMethod_test_all_dtypes_against_cutlass_2x_fna
( self, batch, heads, head_dim, input_shape, kernel_size,
tests/test_flex.py:88
↓ 11 callersMethodaccum_ref
csrc/include/natten/cuda/fmha/gemm/mma_from_smem.h:125
↓ 11 callersMethodaccum_ref
csrc/include/natten/cuda/fna/gemm/mma_from_smem.h:126
↓ 11 callersFunctionceil_div_tuple
(X: tuple, Y: tuple)
src/natten/utils/tuples.py:31
↓ 11 callersMethodclear
csrc/include/natten/cuda/fmha/kernel_backward.h:1086
↓ 11 callersFunctiongemm_zero_acc
csrc/include/natten/cuda/fna_hopper/collective/fna_common.hpp:111
↓ 11 callersFunctiongemm_zero_acc
csrc/include/natten/cuda/fmha_hopper/collective/fmha_common.hpp:68
↓ 11 callersMethodget_block_coord
csrc/include/natten/cuda/fmha_hopper/kernel/fmha_tile_scheduler.hpp:78
↓ 11 callersFunctionload_with_byte_offset
Loads a fragment from memory
csrc/include/natten/cuda/fmha/iterators/epilogue_predicated_tile_iterator.h:347
↓ 11 callersMethodstore
Store a fragment to memory
csrc/include/natten/cuda/fmha/iterators/predicated_tile_iterator_residual_last.h:912
↓ 11 callersFunctionstore_with_byte_offset
Stores a fragment to memory
csrc/include/natten/cuda/fmha/iterators/epilogue_predicated_tile_iterator.h:420
↓ 10 callersFunction_reset_everything
(random_seed: int = 42, torch_seed: int = 42)
tests/test_fmha.py:60
↓ 10 callersMethodappend
(self, na_dim: int)
scripts/autogen_reference_fna.py:298
↓ 10 callersMethodappend
(self, sm: int)
scripts/autogen_fmha.py:280
↓ 10 callersMethodapply
csrc/include/natten/cuda/fna/epilogue/epilogue_pipelined.h:85
↓ 10 callersFunctionattention
Runs standard dot product attention. This operation is used to implement neighborhood cross attention, in which we allow every token to inter
src/natten/functional.py:74
↓ 10 callersFunctioncan_run_flex_attention
( query: Tensor, key: Tensor, value: Tensor, torch_compile: bool, is_causal: bool = False,
src/natten/backends/configs/checks.py:598
↓ 10 callersMethodis_valid
csrc/include/natten/cuda/fmha_hopper/kernel/fmha_tile_scheduler.hpp:73
↓ 10 callersFunctionlayout_acc_mn
csrc/include/natten/cuda/fmha_hopper/collective/fmha_common.hpp:127
↓ 10 callersFunctionna_tensor_checks
( query: Tensor, key: Tensor, value: Tensor, must_match_head_dims: bool = False, supports_
src/natten/utils/checks.py:85
↓ 10 callersMethodnum_splits_key_device
csrc/include/natten/cuda/fmha/kernel_backward.h:673
↓ 10 callersMethodstore
Stores a fragment to memory
csrc/include/natten/cuda/fna/epilogue/predicated_tile_iterator.h:484
↓ 10 callersFunctiontuple_add
csrc/include/natten/cuda/fna_blackwell/collective/fna_fusion.hpp:63
↓ 9 callersMethodappend
(self, dtype: DataType)
scripts/autogen_blackwell_fmha_bwd.py:295
↓ 9 callersMethodappend
(self, dtype: DataType)
scripts/autogen_blackwell_fmha.py:267
↓ 9 callersMethodappend
(self, dtype: DataType)
scripts/autogen_hopper_fmha_bwd.py:277
↓ 9 callersMethodappend
(self, dtype: DataType)
scripts/autogen_hopper_fmha.py:313
↓ 9 callersFunctionfmha_tensor_checks
( query: Tensor, key: Tensor, value: Tensor, must_match_head_dims: bool = False, supports_
src/natten/utils/checks.py:185
↓ 9 callersFunctionis_torch_compiling
()
src/natten/utils/environment.py:74
↓ 9 callersFunctionmul_tuple
(X: tuple, Y: tuple)
src/natten/utils/tuples.py:36
↓ 9 callersFunctionneighborhood_attention_generic
( query: Tensor, key: Tensor, value: Tensor, kernel_size: DimensionTypeOrDed, stride: Dime
src/natten/functional.py:354
↓ 9 callersMethodset_mask
Sets the predicate mask, overriding value stored in predicate iterator
csrc/include/natten/cuda/fmha/iterators/predicated_tile_access_iterator_residual_last.h:872
↓ 9 callersMethodwarning
(self, *args, **kwargs)
src/natten/utils/log.py:128
↓ 8 callersFunction_get_log_pipe
()
src/natten/utils/log.py:73
↓ 8 callersFunction_reset_everything
(random_seed: int = 42, torch_seed: int = 42)
tests/test_fmha_varlen.py:56
↓ 8 callersFunction_reset_everything
(random_seed: int = 42, torch_seed: int = 42)
tests/test_hopper_fna.py:44
↓ 8 callersFunction_reset_everything
(random_seed: int = 42, torch_seed: int = 42)
tests/test_blackwell_fna.py:45
↓ 8 callersMethodadd_pointer_offset
Adds a pointer offset in units of Element
csrc/include/natten/cuda/fna/epilogue/predicated_tile_iterator.h:299
↓ 8 callersFunctiongetQueryStart
Iteration order logic
csrc/include/natten/cuda/fna/kernel_backward.h:2337
↓ 8 callersMethodget_mask
Gets the mask
csrc/include/natten/cuda/fmha/iterators/predicated_tile_access_iterator_residual_last.h:878
↓ 8 callersMethodget_trip_count
csrc/include/natten/cuda/fmha_hopper/collective/fmha_fusion.hpp:121
↓ 8 callersFunctionlayout_acc_mn
csrc/include/natten/cuda/fna_hopper/collective/fna_common.hpp:170
↓ 8 callersMethodload
Loads a fragment from memory
csrc/include/natten/cuda/fna/epilogue/predicated_tile_iterator.h:379
↓ 8 callersMethodreset
()
src/natten/context.py:46
↓ 8 callersFunctionreset_torch_compile
(cache_size_limit, recompile_limit: int | None = None)
tests/utils.py:37
↓ 8 callersMethodset_residual_tile
csrc/include/natten/cuda/fna/iterators/predicated_tile_iterator_residual_last.h:553
↓ 8 callersMethodstep
csrc/include/natten/cuda/fna_hopper/collective/fna_collective_load.hpp:198
↓ 8 callersMethodstep
csrc/include/natten/cuda/fmha_hopper/collective/fmha_collective_load.hpp:189
↓ 8 callersFunctiontuple_mul
csrc/include/natten/cuda/fna_hopper/collective/fna_common.hpp:80
↓ 8 callersFunctiontuple_mul
csrc/include/natten/cuda/fna_blackwell/collective/fna_fusion.hpp:69
↓ 7 callersMethod_test_backend_against_natten
( self, batch, heads, head_dim, seqlen_q, seqlen_kv, i
tests/test_fmha.py:956
↓ 7 callersFunctionapply_variable_length
csrc/include/natten/cuda/fmha_hopper/collective/fmha_varlen.hpp:61
↓ 7 callersMethodclear_mask
Clears the predicate set efficiently
csrc/include/natten/cuda/fmha/iterators/predicated_tile_access_iterator_residual_last.h:860
↓ 7 callersMethodenable_mask
Clears the predicate set efficiently
csrc/include/natten/cuda/fmha/iterators/predicated_tile_access_iterator_residual_last.h:866
next →1–100 of 2,059, ranked by callers