MCPcopy Create free account

hub / github.com/SHI-Labs/NATTEN / functions

Functions2,059 in github.com/SHI-Labs/NATTEN

↓ 4 callersMethodget_trip_count
csrc/include/natten/cuda/fna_hopper/collective/fna_fusion.hpp:268
↓ 4 callersMethodis_contributing
csrc/include/natten/cuda/fmha_hopper/collective/fmha_fusion.hpp:227
↓ 4 callersMethodload
csrc/include/natten/cuda/fna/kernel_backward.h:149
↓ 4 callersFunctionparse_env_flag
(env_var: str, default: bool)
src/natten/utils/environment.py:31
↓ 4 callersFunctionreduction_target_n
csrc/include/natten/cuda/fna_hopper/collective/fna_common.hpp:201
↓ 4 callersFunctionreduction_target_n
csrc/include/natten/cuda/fmha_hopper/collective/fmha_common.hpp:158
↓ 4 callersFunctionreference_fna_generic
( query: Tensor, key: Tensor, value: Tensor, kernel_size: DimensionTypeOrDed, stride: Dime
src/natten/backends/reference.py:182
↓ 4 callersFunctionseparator
(char="-", junction="+")
src/natten/profiling_utils/pretty_printer.py:78
↓ 4 callersMethodset_mask
Sets the predicate mask, overriding value stored in predicate iterator
csrc/include/natten/cuda/fna/iterators/predicated_tile_access_iterator_residual_last.h:954
↓ 4 callersMethodset_mask
Sets the predicate mask, overriding value stored in predicate iterator
csrc/include/natten/cuda/fna/iterators/predicated_tile_access_iterator.h:462
↓ 4 callersFunctionstore
Stores a fragment to memory
csrc/include/natten/cuda/fmha/iterators/epilogue_predicated_tile_iterator.h:497
↓ 4 callersMethodstore
csrc/include/natten/cuda/fmha_blackwell/kernel/sm100_fmha_bwd_kernel_tma_warpspecialized.hpp:1183
↓ 4 callersFunctionsupports_bfloat16
(device: torch.device)
src/natten/utils/testing.py:139
↓ 4 callersFunctiontensor_op_mk_v
csrc/include/natten/cuda/fna_hopper/collective/fna_common.hpp:194
↓ 4 callersFunctiontensor_op_mk_v
csrc/include/natten/cuda/fmha_hopper/collective/fmha_common.hpp:151
↓ 4 callersMethodthread_start_row
Need to get the thread start row from the tile iterator
csrc/include/natten/cuda/fna/epilogue/predicated_tile_iterator.h:495
↓ 4 callersFunctiontuple_leq
csrc/include/natten/cuda/fna_blackwell/collective/fna_fusion.hpp:75
↓ 4 callersFunctionwrite_header_file
(content, path, namespaces, extra_includes=None)
scripts/autogen_blackwell_fmha_bwd.py:441
↓ 4 callersFunctionwrite_header_file
(content, path, namespaces, extra_includes=None)
scripts/autogen_reference_fna.py:437
↓ 4 callersFunctionwrite_header_file
(content, path, namespaces, extra_includes=None)
scripts/autogen_blackwell_fmha.py:420
↓ 4 callersFunctionwrite_header_file
(content, path, namespaces, extra_includes=None)
scripts/autogen_hopper_fmha_bwd.py:426
↓ 4 callersFunctionwrite_header_file
(content, path, namespaces, extra_includes=None)
scripts/autogen_hopper_fmha.py:470
↓ 3 callersMethod__init__
( self, embed_dim: int, num_heads: int, kernel_size: Dimension1DTypeOrDed,
src/natten/modules.py:208
↓ 3 callersMethod_file_logger
(self, name: str, path: str, level: str)
tests/test_log.py:171
↓ 3 callersFunction_make_fn
(backend, use_compile, is_causal)
tests/test_varlen_recompile.py:164
↓ 3 callersFunction_reset_everything
(random_seed: int = 42, torch_seed: int = 42)
tests/test_attn_merge.py:43
↓ 3 callersFunction_reset_everything
(random_seed: int = 42, torch_seed: int = 42)
tests/test_compute_delta.py:38
↓ 3 callersMethod_test_against_manual_varlen
( self, batch: int, heads: int, head_dim: int, seqlens_Q_list: List[in
tests/test_fmha_varlen.py:216
↓ 3 callersMethod_test_against_natten_fmha
( self, batch: int, heads: int, head_dim: int, seqlen_q: int,
tests/test_fmha.py:416
↓ 3 callersMethod_test_against_reference
( self, batch, seqlen, heads, head_dim, eps, dtype, dtype_out )
tests/test_compute_delta.py:59
↓ 3 callersMethod_test_against_torch_sdpa
( self, batch: int, heads: int, head_dim: int, seqlen_q: int,
tests/test_fmha.py:353
↓ 3 callersMethod_test_randsweep_against_cutlass_2x
(self, na_dim, torch_compile: bool = False)
tests/test_flex.py:631
↓ 3 callersFunction_universal_tensor_checks
( query: Tensor, key: Tensor, value: Tensor, raise_error: bool = True )
src/natten/utils/checks.py:49
↓ 3 callersFunction_write
(data, path)
scripts/testing/pytest_progress_plugin.py:17
↓ 3 callersFunctionadd_pointer_offset
Adds a pointer offset in units of Element
csrc/include/natten/cuda/fmha/iterators/predicated_tile_access_iterator_residual_last.h:299
↓ 3 callersFunctionalign_up
csrc/include/natten/cuda/fna/gemm_kernel_utils.h:68
↓ 3 callersFunctioncheck_dilation_arg
(na_dim: int, dilation: Any)
src/natten/utils/checks.py:430
↓ 3 callersFunctionconvert_to_gmma_rs
csrc/include/natten/cuda/fna_hopper/collective/fna_common.hpp:220
↓ 3 callersFunctionconvert_to_gmma_rs
csrc/include/natten/cuda/fmha_hopper/collective/fmha_common.hpp:177
↓ 3 callersFunctiondrain_cp_asyncs
csrc/include/natten/cuda/fmha/gemm/mma_from_smem.h:543
↓ 3 callersFunctiondrain_cp_asyncs
csrc/include/natten/cuda/fna/gemm/mma_from_smem.h:544
↓ 3 callersFunctionformat_row
(row, bold_row=False)
src/natten/profiling_utils/pretty_printer.py:71
↓ 3 callersFunctiongenerate_varlen_parameters
( query: Tensor, key: Tensor, value: Tensor, seqlens_Q: Optional[Tensor] = None, seqlens_K
src/natten/utils/varlen.py:33
↓ 3 callersFunctiongetNumParallelBlocksForQuery
Returns how many kernel blocks will write to a given block in `grad_query` This is usually equal to the number of key splits, but can be different for
csrc/include/natten/cuda/fna/kernel_backward.h:2385
↓ 3 callersFunctionget_all_forward_configs
( input_tensor: Tensor, )
src/natten/backends/configs/cutlass_blackwell/__init__.py:105
↓ 3 callersMethodget_dispatcher
(self)
scripts/autogen_blackwell_fmha_bwd.py:298
↓ 3 callersMethodget_dispatcher
(self)
scripts/autogen_blackwell_fmha.py:270
↓ 3 callersMethodget_dispatcher
(self)
scripts/autogen_hopper_fmha_bwd.py:280
↓ 3 callersMethodget_dispatcher
(self)
scripts/autogen_hopper_fmha.py:316
↓ 3 callersMethodget_dispatcher
(self)
scripts/autogen_fmha.py:283
↓ 3 callersFunctionget_memory_usage_preference
()
src/natten/context.py:74
↓ 3 callersMethodget_name
(self)
scripts/autogen_hopper_fna_bwd.py:192
↓ 3 callersMethodget_name
(self)
scripts/autogen_blackwell_fmha_bwd.py:221
↓ 3 callersMethodget_name
(self)
scripts/autogen_blackwell_fna.py:182
↓ 3 callersMethodget_name
(self)
scripts/autogen_blackwell_fna_bwd.py:189
↓ 3 callersMethodget_name
(self)
scripts/autogen_reference_fna.py:215
↓ 3 callersMethodget_name
(self)
scripts/autogen_blackwell_fmha.py:190
↓ 3 callersMethodget_name
(self)
scripts/autogen_hopper_fmha_bwd.py:203
↓ 3 callersMethodget_name
(self)
scripts/autogen_hopper_fmha.py:237
↓ 3 callersMethodget_name
(self)
scripts/autogen_hopper_fna.py:225
↓ 3 callersMethodget_trip_count
csrc/include/natten/cuda/fmha_blackwell/collective/fmha_fusion.hpp:43
↓ 3 callersFunctionget_window_left
csrc/include/natten/cuda/fna_hopper/collective/fna_fusion.hpp:427
↓ 3 callersFunctionget_window_right
csrc/include/natten/cuda/fna_hopper/collective/fna_fusion.hpp:432
↓ 3 callersMethodinfo
(self, *args, **kwargs)
src/natten/utils/log.py:120
↓ 3 callersMethodinit
csrc/include/natten/cuda/fmha_hopper/collective/fmha_collective_softmax.hpp:51
↓ 3 callersFunctioninit_tensors
( problem: Problem, flatten_sequence: bool, heads_last: bool )
src/natten/profiling_utils/profiling.py:46
↓ 3 callersMethodinitialize
Initializes state from arguments.
csrc/include/natten/cuda/fmha_blackwell/device/fmha_bwd_sm100.hpp:360
↓ 3 callersMethodinitialize
Initializes state from arguments.
csrc/include/natten/cuda/fna_hopper/device/fna_bwd_sm90.hpp:341
↓ 3 callersMethodinitialize
Initializes state from arguments.
csrc/include/natten/cuda/fna_blackwell/device/fna_bwd_sm100.hpp:347
↓ 3 callersMethodinitialize
Initializes state from arguments.
csrc/include/natten/cuda/fmha_hopper/device/fmha_bwd_sm90.hpp:315
↓ 3 callersFunctionis_dilated
csrc/include/natten/cuda/fna_hopper/collective/fna_fusion.hpp:530
↓ 3 callersFunctionis_flex_compile_allowed
Returns whether compilation is allowed in `"flex-fna"` and `"flex-fmha"` backends.
src/natten/context.py:149
↓ 3 callersMethodis_safe_to_log
(self)
src/natten/utils/log.py:117
↓ 3 callersFunctioniterable_to_static_cute_tuple
(shape_in)
scripts/autogen_hopper_fna_bwd.py:144
↓ 3 callersFunctioniterable_to_static_cute_tuple
(shape_in)
scripts/autogen_blackwell_fna.py:132
↓ 3 callersFunctioniterable_to_static_cute_tuple
(shape_in)
scripts/autogen_blackwell_fna_bwd.py:141
↓ 3 callersFunctioniterable_to_static_cute_tuple
(shape_in)
scripts/autogen_hopper_fna.py:152
↓ 3 callersFunctionlayout_separate
csrc/include/natten/cuda/fna_hopper/collective/fna_common.hpp:146
↓ 3 callersFunctionlayout_separate
csrc/include/natten/cuda/fmha_hopper/collective/fmha_common.hpp:103
↓ 3 callersMethodload
csrc/include/natten/cuda/fmha/kernel_backward.h:119
↓ 3 callersMethodload_with_byte_offset
Loads a fragment from memory
csrc/include/natten/cuda/fmha/iterators/predicated_tile_iterator_residual_last.h:888
↓ 3 callersFunctionmake_blackwell_fna_ops
(na_dim)
src/natten/_libnatten/torch_wrappers.py:657
↓ 3 callersFunctionmake_cutlass_blackwell_fna_autograd_fn
(na_dim)
src/natten/backends/blackwell_fna.py:73
↓ 3 callersFunctionmake_cutlass_fna_autograd_fn
(na_dim)
src/natten/backends/fna.py:75
↓ 3 callersFunctionmake_cutlass_hopper_fna_autograd_fn
(na_dim)
src/natten/backends/hopper_fna.py:74
↓ 3 callersFunctionmake_cutlass_token_permute_autograd_fn
(na_dim)
src/natten/token_permute/cutlass_impl.py:85
↓ 3 callersFunctionmake_cutlass_token_unpermute_autograd_fn
(na_dim)
src/natten/token_permute/cutlass_impl.py:141
↓ 3 callersFunctionmake_fna_ops
(na_dim)
src/natten/_libnatten/torch_wrappers.py:1021
↓ 3 callersFunctionmake_hopper_fna_ops
(na_dim)
src/natten/_libnatten/torch_wrappers.py:839
↓ 3 callersFunctionmake_reference_fna_autograd_fn
(na_dim)
src/natten/backends/reference.py:66
↓ 3 callersFunctionmake_reference_fna_ops
(na_dim)
src/natten/_libnatten/torch_wrappers.py:1207
↓ 3 callersFunctionmake_token_permute_ops
(na_dim)
src/natten/_libnatten/torch_wrappers.py:1373
↓ 3 callersFunctionparse_env_str
(env_var: str, default: str)
src/natten/utils/environment.py:53
↓ 3 callersFunctionprint_table
(title, headers, values, has_footer=False)
src/natten/profiling_utils/pretty_printer.py:51
↓ 3 callersMethodref
Returns a TensorRef to the operand
csrc/include/natten/cuda/fmha/gemm/custom_mma_base.h:120
↓ 3 callersMethodref
Returns a TensorRef to the operand
csrc/include/natten/cuda/fna/gemm/custom_mma_base.h:120
↓ 3 callersFunctionrender_table
Render the test status table to `out`. Returns done count.
scripts/testing/monitor.py:152
↓ 3 callersMethodsplit_key_device
csrc/include/natten/cuda/fmha/kernel_backward.h:681
↓ 3 callersMethodsplit_key_device_dim
csrc/include/natten/cuda/fna/kernel_backward.h:730
↓ 3 callersMethodstore
csrc/include/natten/cuda/fmha/kernel_backward.h:134
← previousnext →201–300 of 2,059, ranked by callers