MCPcopy Create free account

hub / github.com/SHI-Labs/NATTEN / functions

Functions2,059 in github.com/SHI-Labs/NATTEN

↓ 3 callersMethodstore
csrc/include/natten/cuda/fna/kernel_backward.h:164
↓ 3 callersMethodstore_with_byte_offset
Store a fragment to memory
csrc/include/natten/cuda/fmha/iterators/predicated_tile_iterator_residual_last.h:906
↓ 3 callersFunctionunstageSmemLayout
csrc/include/natten/cuda/fmha_blackwell/collective/fmha_common.hpp:71
↓ 3 callersFunctionunstageSmemLayout
csrc/include/natten/cuda/fna_hopper/collective/fna_common.hpp:284
↓ 3 callersFunctionunstageSmemLayout
csrc/include/natten/cuda/fna_blackwell/collective/fna_common.hpp:71
↓ 3 callersFunctionunstageSmemLayout
csrc/include/natten/cuda/fmha_hopper/collective/fmha_common.hpp:241
↓ 3 callersFunctionwrite_header_file
(content, path, namespaces, extra_includes=None)
scripts/autogen_fmha.py:364
↓ 2 callersMethod__init__
( self, embed_dim: int, num_heads: int, mlp_ratio: int, qkv_bias: bool
tests/test_torch_compile.py:162
↓ 2 callersFunction_build_cases
()
tests/test_varlen_recompile.py:224
↓ 2 callersFunction_check_cuda_arch
(arch: Any)
setup.py:146
↓ 2 callersFunction_find_attention_kernels
Return (all_names, fwd_dict, bwd_dict) of attention kernels. Each dict maps kernel name to its call count from the profiler. Forward kernels
tests/test_varlen_recompile.py:145
↓ 2 callersFunction_merge_attentions_fn
( outputs: List[Tensor], lse_tensors: List[Tensor] )
src/natten/attn_merge.py:48
↓ 2 callersFunction_merge_attentions_op
( outputs: List[Tensor], lse_tensors: List[Tensor], torch_compile: bool = True )
src/natten/attn_merge.py:115
↓ 2 callersFunction_profile_case
(label, fn, q, k, v, cumseq_q, cumseq_kv, max_q, max_kv)
tests/test_varlen_recompile.py:127
↓ 2 callersFunction_prologue
csrc/include/natten/cuda/fmha/gemm/mma_from_smem.h:977
↓ 2 callersFunction_prologue
csrc/include/natten/cuda/fna/gemm/mma_from_smem.h:978
↓ 2 callersFunction_reset_everything
(random_seed: int = 42, torch_seed: int = 42)
tests/test_torch_compile.py:38
↓ 2 callersFunction_run_flex_attn
( q: Tensor, k: Tensor, v: Tensor, block_mask: BlockMask, scale: float, torch_compile:
src/natten/backends/flex.py:88
↓ 2 callersMethod_test_against_reference_inputs
( self, inputs, reference, batch: int, heads: int, head_dim: i
tests/test_fmha.py:228
↓ 2 callersMethod_test_permute_unpermute_torch
( self, B, H, S, D, tile_shape, dilation, flip_dims, eps, dtype, device="cuda" )
tests/test_token_permute.py:48
↓ 2 callersFunctionadd_pointer_offset
csrc/include/natten/cuda/fna/iterators/predicated_tile_access_iterator.h:690
↓ 2 callersMethodadd_pointer_offset
Adds a pointer offset in units of Element
csrc/include/natten/cuda/fna/iterators/predicated_tile_iterator.h:734
↓ 2 callersFunctionadd_tile_offset
Advances an iterator along logical dimensions of matrix in units of whole tiles
csrc/include/natten/cuda/fmha/iterators/predicated_tile_access_iterator_residual_last.h:306
↓ 2 callersFunctionadd_tile_offset
Advances an iterator along logical dimensions of matrix in units of whole tiles
csrc/include/natten/cuda/fna/iterators/predicated_tile_access_iterator.h:712
↓ 2 callersFunctionadditional_kv_tensor_checks
( query: Tensor, key: Tensor, value: Tensor, add_key: Optional[Tensor] = None, add_value:
src/natten/utils/checks.py:278
↓ 2 callersMethodappend
(self, sm: int)
scripts/autogen_fna.py:365
↓ 2 callersMethodappend
(self, dtype: DataType)
scripts/autogen_fna.py:419
↓ 2 callersMethodappend
(self, cm)
scripts/autogen_fna.py:482
↓ 2 callersMethodappend
(self, dtype: DataType)
scripts/autogen_fmha.py:330
↓ 2 callersMethodapply_mask
csrc/include/natten/cuda/fna_blackwell/collective/fna_fusion.hpp:384
↓ 2 callersMethodapply_padded_mask
csrc/include/natten/cuda/fna_blackwell/collective/fna_fusion.hpp:442
↓ 2 callersMethodcan_implement
Determines whether the GEMM can execute the given problem.
csrc/include/natten/cuda/fna_hopper/device/fna_sm90.hpp:87
↓ 2 callersFunctionceil_div_int
(x: int, y: int)
src/natten/utils/tuples.py:27
↓ 2 callersFunctionceil_tuple
csrc/include/natten/cuda/fna_hopper/collective/fna_common.hpp:62
↓ 2 callersFunctionceil_tuple
csrc/include/natten/cuda/fna_blackwell/collective/fna_fusion.hpp:51
↓ 2 callersFunctioncheck_input_size_arg
(na_dim: int, input_size: Any)
src/natten/utils/checks.py:373
↓ 2 callersFunctioncheck_kernel_schedule
(kernel_schedule: Any)
src/natten/utils/checks.py:545
↓ 2 callersFunctionchoose_backend
( query: Tensor, key: Tensor, value: Tensor, torch_compile: bool )
src/natten/backends/__init__.py:88
↓ 2 callersFunctionchoose_fmha_backend
( query: Tensor, key: Tensor, value: Tensor, is_causal: bool, is_varlen: bool, torch_c
src/natten/backends/__init__.py:113
↓ 2 callersFunctionclassify
Returns (state, start, end, rc, gpu, worker).
scripts/testing/monitor.py:69
↓ 2 callersMethodclear
< Efficiently disables all accesses guarded by mask
csrc/include/natten/cuda/fna/iterators/epilogue_predicated_tile_iterator.h:171
↓ 2 callersMethodclear_mask
Clears the predicate set efficiently
csrc/include/natten/cuda/fna/iterators/predicated_tile_access_iterator_residual_last.h:942
↓ 2 callersMethodclear_mask
Clears the predicate set efficiently
csrc/include/natten/cuda/fna/iterators/predicated_tile_iterator.h:765
↓ 2 callersMethodcompute
csrc/include/natten/cuda/fmha_hopper/collective/fmha_collective_tma.hpp:283
↓ 2 callersFunctionconvert_to_natten_profiler_ops
( profiler: torch_profile, )
src/natten/profiling_utils/formatting.py:319
↓ 2 callersFunctioncopy_tiles_and_advance
csrc/include/natten/cuda/fmha/gemm/custom_mma_multistage.h:296
↓ 2 callersFunctioncopy_tiles_and_advance
csrc/include/natten/cuda/fna/gemm/custom_mma_multistage.h:296
↓ 2 callersFunctioncopy_tiles_and_advance_1
csrc/include/natten/cuda/fmha/gemm/mma_from_smem.h:943
↓ 2 callersFunctioncopy_tiles_and_advance_1
csrc/include/natten/cuda/fna/gemm/mma_from_smem.h:944
↓ 2 callersFunctioncorrect_qkv_shape_wrt_dilation
csrc/include/natten/cuda/fna_hopper/collective/fna_fusion.hpp:212
↓ 2 callersFunctioncorrect_qkv_shape_wrt_dilation
csrc/include/natten/cuda/fna_blackwell/collective/fna_fusion.hpp:258
↓ 2 callersFunctioncreate_causal_arg_from_bool
(na_dim: int, value: bool)
src/natten/utils/tuples.py:50
↓ 2 callersFunctiondo_profile
( problem: Problem, backend: Optional[str] = None, fmha_backend: Optional[str] = None, q_tile_
src/natten/profiler.py:78
↓ 2 callersFunctiondry_run_for_backend
( problem: Problem, backend: Optional[str], fmha_backend: Optional[str], backprop: bool, t
src/natten/profiling_utils/dry_run.py:70
↓ 2 callersMethodenable_mask
Clears the predicate set efficiently
csrc/include/natten/cuda/fna/iterators/predicated_tile_access_iterator_residual_last.h:948
↓ 2 callersMethodenable_mask
Clears the predicate set efficiently
csrc/include/natten/cuda/fna/iterators/predicated_tile_iterator.h:771
↓ 2 callersFunctionfloor_tuple
csrc/include/natten/cuda/fna_hopper/collective/fna_common.hpp:56
↓ 2 callersFunctionfloor_tuple
csrc/include/natten/cuda/fna_blackwell/collective/fna_fusion.hpp:45
↓ 2 callersFunctionformat_time
(epoch)
scripts/testing/monitor.py:57
↓ 2 callersFunctionfully_block_sparse
csrc/include/natten/cuda/fna_hopper/collective/fna_fusion.hpp:466
↓ 2 callersMethodgQ_strideM
csrc/include/natten/cuda/fmha/kernel_backward.h:639
↓ 2 callersFunctiongetQueryEnd
csrc/include/natten/cuda/fna/kernel_backward.h:2342
↓ 2 callersFunctionget_all_tile_shapes_backward
( input_tensor: Tensor, )
src/natten/backends/configs/cutlass/__init__.py:256
↓ 2 callersFunctionget_all_tile_shapes_forward
( input_tensor: Tensor, )
src/natten/backends/configs/flex/__init__.py:79
↓ 2 callersFunctionget_all_tile_shapes_forward
( input_tensor: Tensor, )
src/natten/backends/configs/cutlass/__init__.py:83
↓ 2 callersFunctionget_compatible_backends
( query: Tensor, key: Tensor, value: Tensor, torch_compile: bool )
src/natten/backends/__init__.py:156
↓ 2 callersFunctionget_compatible_fmha_backends
( query: Tensor, key: Tensor, value: Tensor, is_causal: bool, is_varlen: bool, torch_c
src/natten/backends/__init__.py:175
↓ 2 callersFunctionget_default_backward_config
( input_tensor: Tensor, )
src/natten/backends/configs/cutlass_hopper/__init__.py:268
↓ 2 callersFunctionget_default_backward_config
( input_tensor: Tensor, )
src/natten/backends/configs/cutlass/__init__.py:319
↓ 2 callersFunctionget_default_backward_tile_shapes
( input_tensor: Tensor, )
src/natten/backends/configs/cutlass_blackwell/__init__.py:181
↓ 2 callersFunctionget_default_forward_config
( input_tensor: Tensor, )
src/natten/backends/configs/cutlass_hopper/__init__.py:253
↓ 2 callersFunctionget_default_forward_config
( input_tensor: Tensor, dilation: Optional[DimensionType] = None )
src/natten/backends/configs/cutlass/__init__.py:129
↓ 2 callersFunctionget_default_forward_tile_shapes
(input_tensor: Tensor)
src/natten/backends/configs/flex/__init__.py:101
↓ 2 callersFunctionget_default_forward_tile_shapes
( input_tensor: Tensor, )
src/natten/backends/configs/cutlass_blackwell/__init__.py:160
↓ 2 callersFunctionget_default_kv_splits_backward
( input_tensor: Tensor, kv_tile_shape: DimensionType, deterministic: bool, dilation: Optional[
src/natten/backends/configs/cutlass/backward_knobs.py:131
↓ 2 callersFunctionget_dim_type
(na_dim: int)
scripts/autogen_hopper_fna_bwd.py:149
↓ 2 callersFunctionget_dim_type
(na_dim: int)
scripts/autogen_blackwell_fna.py:137
↓ 2 callersFunctionget_dim_type
(na_dim: int)
scripts/autogen_blackwell_fna_bwd.py:146
↓ 2 callersFunctionget_dim_type
(na_dim: int)
scripts/autogen_reference_fna.py:189
↓ 2 callersFunctionget_dim_type
(na_dim: int)
scripts/autogen_hopper_fna.py:157
↓ 2 callersMethodget_kernel_instance
(self, q_tile_shape, kv_tile_shape)
scripts/autogen_hopper_fna_bwd.py:463
↓ 2 callersMethodget_kernel_instance
(self, q_tile_size, kv_tile_size)
scripts/autogen_blackwell_fmha_bwd.py:383
↓ 2 callersMethodget_kernel_instance
(self, q_tile_shape, kv_tile_shape, persistent)
scripts/autogen_blackwell_fna.py:452
↓ 2 callersMethodget_kernel_instance
(self, q_tile_shape, kv_tile_shape)
scripts/autogen_blackwell_fna_bwd.py:458
↓ 2 callersMethodget_kernel_instance
(self, cm)
scripts/autogen_reference_fna.py:386
↓ 2 callersMethodget_kernel_instance
(self, q_tile_size, kv_tile_size, persistent)
scripts/autogen_blackwell_fmha.py:355
↓ 2 callersMethodget_kernel_instance
(self, q_tile_size, kv_tile_size)
scripts/autogen_hopper_fmha_bwd.py:367
↓ 2 callersMethodget_kernel_instance
(self, q_tile_size, kv_tile_size, kernel_type)
scripts/autogen_hopper_fmha.py:400
↓ 2 callersMethodget_kernel_instance
(self, q_tile_shape, kv_tile_shape, kernel_type)
scripts/autogen_hopper_fna.py:497
↓ 2 callersMethodget_mask
Gets the mask
csrc/include/natten/cuda/fna/iterators/predicated_tile_iterator.h:783
↓ 2 callersMethodget_masked_trip_count
csrc/include/natten/cuda/fmha_blackwell/collective/fmha_fusion.hpp:51
↓ 2 callersMethodget_masked_trip_count
csrc/include/natten/cuda/fmha_hopper/collective/fmha_fusion.hpp:135
↓ 2 callersFunctionget_matching_files
(directory)
setup.py:249
↓ 2 callersFunctionget_max_splits
( input_shape: DimensionType, dilation: DimensionType, kv_tile_shape: DimensionType )
src/natten/backends/configs/cutlass/backward_knobs.py:61
↓ 2 callersMethodget_name
(self, is_backward: bool)
scripts/autogen_fna.py:42
↓ 2 callersMethodget_name
(self, is_backward: bool)
scripts/autogen_fmha.py:40
↓ 2 callersMethodget_target_name
(self, q_tile_shape, kv_tile_shape, persistent)
scripts/autogen_blackwell_fna.py:472
↓ 2 callersMethodget_target_name
(self, q_tile_size, kv_tile_size, persistent)
scripts/autogen_blackwell_fmha.py:372
↓ 2 callersMethodget_unmasked_trip_count
csrc/include/natten/cuda/fmha_hopper/collective/fmha_fusion.hpp:143
↓ 2 callersFunctionget_window_left
csrc/include/natten/cuda/fna_blackwell/collective/fna_fusion.hpp:493
← previousnext →301–400 of 2,059, ranked by callers