Code
Hub
Workspaces
Following
Trending
Connect
MCP
copy
Create free account
hub
/
github.com/SHI-Labs/NATTEN
/ functions
Functions
2,059 in github.com/SHI-Labs/NATTEN
⨍
Functions
2,059
◇
Types & classes
573
↓ 4 callers
Method
get_trip_count
csrc/include/natten/cuda/fna_hopper/collective/fna_fusion.hpp:268
↓ 4 callers
Method
is_contributing
csrc/include/natten/cuda/fmha_hopper/collective/fmha_fusion.hpp:227
↓ 4 callers
Method
load
csrc/include/natten/cuda/fna/kernel_backward.h:149
↓ 4 callers
Function
parse_env_flag
(env_var: str, default: bool)
src/natten/utils/environment.py:31
↓ 4 callers
Function
reduction_target_n
csrc/include/natten/cuda/fna_hopper/collective/fna_common.hpp:201
↓ 4 callers
Function
reduction_target_n
csrc/include/natten/cuda/fmha_hopper/collective/fmha_common.hpp:158
↓ 4 callers
Function
reference_fna_generic
( query: Tensor, key: Tensor, value: Tensor, kernel_size: DimensionTypeOrDed, stride: Dime
src/natten/backends/reference.py:182
↓ 4 callers
Function
separator
(char="-", junction="+")
src/natten/profiling_utils/pretty_printer.py:78
↓ 4 callers
Method
set_mask
Sets the predicate mask, overriding value stored in predicate iterator
csrc/include/natten/cuda/fna/iterators/predicated_tile_access_iterator_residual_last.h:954
↓ 4 callers
Method
set_mask
Sets the predicate mask, overriding value stored in predicate iterator
csrc/include/natten/cuda/fna/iterators/predicated_tile_access_iterator.h:462
↓ 4 callers
Function
store
Stores a fragment to memory
csrc/include/natten/cuda/fmha/iterators/epilogue_predicated_tile_iterator.h:497
↓ 4 callers
Method
store
csrc/include/natten/cuda/fmha_blackwell/kernel/sm100_fmha_bwd_kernel_tma_warpspecialized.hpp:1183
↓ 4 callers
Function
supports_bfloat16
(device: torch.device)
src/natten/utils/testing.py:139
↓ 4 callers
Function
tensor_op_mk_v
csrc/include/natten/cuda/fna_hopper/collective/fna_common.hpp:194
↓ 4 callers
Function
tensor_op_mk_v
csrc/include/natten/cuda/fmha_hopper/collective/fmha_common.hpp:151
↓ 4 callers
Method
thread_start_row
Need to get the thread start row from the tile iterator
csrc/include/natten/cuda/fna/epilogue/predicated_tile_iterator.h:495
↓ 4 callers
Function
tuple_leq
csrc/include/natten/cuda/fna_blackwell/collective/fna_fusion.hpp:75
↓ 4 callers
Function
write_header_file
(content, path, namespaces, extra_includes=None)
scripts/autogen_blackwell_fmha_bwd.py:441
↓ 4 callers
Function
write_header_file
(content, path, namespaces, extra_includes=None)
scripts/autogen_reference_fna.py:437
↓ 4 callers
Function
write_header_file
(content, path, namespaces, extra_includes=None)
scripts/autogen_blackwell_fmha.py:420
↓ 4 callers
Function
write_header_file
(content, path, namespaces, extra_includes=None)
scripts/autogen_hopper_fmha_bwd.py:426
↓ 4 callers
Function
write_header_file
(content, path, namespaces, extra_includes=None)
scripts/autogen_hopper_fmha.py:470
↓ 3 callers
Method
__init__
( self, embed_dim: int, num_heads: int, kernel_size: Dimension1DTypeOrDed,
src/natten/modules.py:208
↓ 3 callers
Method
_file_logger
(self, name: str, path: str, level: str)
tests/test_log.py:171
↓ 3 callers
Function
_make_fn
(backend, use_compile, is_causal)
tests/test_varlen_recompile.py:164
↓ 3 callers
Function
_reset_everything
(random_seed: int = 42, torch_seed: int = 42)
tests/test_attn_merge.py:43
↓ 3 callers
Function
_reset_everything
(random_seed: int = 42, torch_seed: int = 42)
tests/test_compute_delta.py:38
↓ 3 callers
Method
_test_against_manual_varlen
( self, batch: int, heads: int, head_dim: int, seqlens_Q_list: List[in
tests/test_fmha_varlen.py:216
↓ 3 callers
Method
_test_against_natten_fmha
( self, batch: int, heads: int, head_dim: int, seqlen_q: int,
tests/test_fmha.py:416
↓ 3 callers
Method
_test_against_reference
( self, batch, seqlen, heads, head_dim, eps, dtype, dtype_out )
tests/test_compute_delta.py:59
↓ 3 callers
Method
_test_against_torch_sdpa
( self, batch: int, heads: int, head_dim: int, seqlen_q: int,
tests/test_fmha.py:353
↓ 3 callers
Method
_test_randsweep_against_cutlass_2x
(self, na_dim, torch_compile: bool = False)
tests/test_flex.py:631
↓ 3 callers
Function
_universal_tensor_checks
( query: Tensor, key: Tensor, value: Tensor, raise_error: bool = True )
src/natten/utils/checks.py:49
↓ 3 callers
Function
_write
(data, path)
scripts/testing/pytest_progress_plugin.py:17
↓ 3 callers
Function
add_pointer_offset
Adds a pointer offset in units of Element
csrc/include/natten/cuda/fmha/iterators/predicated_tile_access_iterator_residual_last.h:299
↓ 3 callers
Function
align_up
csrc/include/natten/cuda/fna/gemm_kernel_utils.h:68
↓ 3 callers
Function
check_dilation_arg
(na_dim: int, dilation: Any)
src/natten/utils/checks.py:430
↓ 3 callers
Function
convert_to_gmma_rs
csrc/include/natten/cuda/fna_hopper/collective/fna_common.hpp:220
↓ 3 callers
Function
convert_to_gmma_rs
csrc/include/natten/cuda/fmha_hopper/collective/fmha_common.hpp:177
↓ 3 callers
Function
drain_cp_asyncs
csrc/include/natten/cuda/fmha/gemm/mma_from_smem.h:543
↓ 3 callers
Function
drain_cp_asyncs
csrc/include/natten/cuda/fna/gemm/mma_from_smem.h:544
↓ 3 callers
Function
format_row
(row, bold_row=False)
src/natten/profiling_utils/pretty_printer.py:71
↓ 3 callers
Function
generate_varlen_parameters
( query: Tensor, key: Tensor, value: Tensor, seqlens_Q: Optional[Tensor] = None, seqlens_K
src/natten/utils/varlen.py:33
↓ 3 callers
Function
getNumParallelBlocksForQuery
Returns how many kernel blocks will write to a given block in `grad_query` This is usually equal to the number of key splits, but can be different for
csrc/include/natten/cuda/fna/kernel_backward.h:2385
↓ 3 callers
Function
get_all_forward_configs
( input_tensor: Tensor, )
src/natten/backends/configs/cutlass_blackwell/__init__.py:105
↓ 3 callers
Method
get_dispatcher
(self)
scripts/autogen_blackwell_fmha_bwd.py:298
↓ 3 callers
Method
get_dispatcher
(self)
scripts/autogen_blackwell_fmha.py:270
↓ 3 callers
Method
get_dispatcher
(self)
scripts/autogen_hopper_fmha_bwd.py:280
↓ 3 callers
Method
get_dispatcher
(self)
scripts/autogen_hopper_fmha.py:316
↓ 3 callers
Method
get_dispatcher
(self)
scripts/autogen_fmha.py:283
↓ 3 callers
Function
get_memory_usage_preference
()
src/natten/context.py:74
↓ 3 callers
Method
get_name
(self)
scripts/autogen_hopper_fna_bwd.py:192
↓ 3 callers
Method
get_name
(self)
scripts/autogen_blackwell_fmha_bwd.py:221
↓ 3 callers
Method
get_name
(self)
scripts/autogen_blackwell_fna.py:182
↓ 3 callers
Method
get_name
(self)
scripts/autogen_blackwell_fna_bwd.py:189
↓ 3 callers
Method
get_name
(self)
scripts/autogen_reference_fna.py:215
↓ 3 callers
Method
get_name
(self)
scripts/autogen_blackwell_fmha.py:190
↓ 3 callers
Method
get_name
(self)
scripts/autogen_hopper_fmha_bwd.py:203
↓ 3 callers
Method
get_name
(self)
scripts/autogen_hopper_fmha.py:237
↓ 3 callers
Method
get_name
(self)
scripts/autogen_hopper_fna.py:225
↓ 3 callers
Method
get_trip_count
csrc/include/natten/cuda/fmha_blackwell/collective/fmha_fusion.hpp:43
↓ 3 callers
Function
get_window_left
csrc/include/natten/cuda/fna_hopper/collective/fna_fusion.hpp:427
↓ 3 callers
Function
get_window_right
csrc/include/natten/cuda/fna_hopper/collective/fna_fusion.hpp:432
↓ 3 callers
Method
info
(self, *args, **kwargs)
src/natten/utils/log.py:120
↓ 3 callers
Method
init
csrc/include/natten/cuda/fmha_hopper/collective/fmha_collective_softmax.hpp:51
↓ 3 callers
Function
init_tensors
( problem: Problem, flatten_sequence: bool, heads_last: bool )
src/natten/profiling_utils/profiling.py:46
↓ 3 callers
Method
initialize
Initializes state from arguments.
csrc/include/natten/cuda/fmha_blackwell/device/fmha_bwd_sm100.hpp:360
↓ 3 callers
Method
initialize
Initializes state from arguments.
csrc/include/natten/cuda/fna_hopper/device/fna_bwd_sm90.hpp:341
↓ 3 callers
Method
initialize
Initializes state from arguments.
csrc/include/natten/cuda/fna_blackwell/device/fna_bwd_sm100.hpp:347
↓ 3 callers
Method
initialize
Initializes state from arguments.
csrc/include/natten/cuda/fmha_hopper/device/fmha_bwd_sm90.hpp:315
↓ 3 callers
Function
is_dilated
csrc/include/natten/cuda/fna_hopper/collective/fna_fusion.hpp:530
↓ 3 callers
Function
is_flex_compile_allowed
Returns whether compilation is allowed in `"flex-fna"` and `"flex-fmha"` backends.
src/natten/context.py:149
↓ 3 callers
Method
is_safe_to_log
(self)
src/natten/utils/log.py:117
↓ 3 callers
Function
iterable_to_static_cute_tuple
(shape_in)
scripts/autogen_hopper_fna_bwd.py:144
↓ 3 callers
Function
iterable_to_static_cute_tuple
(shape_in)
scripts/autogen_blackwell_fna.py:132
↓ 3 callers
Function
iterable_to_static_cute_tuple
(shape_in)
scripts/autogen_blackwell_fna_bwd.py:141
↓ 3 callers
Function
iterable_to_static_cute_tuple
(shape_in)
scripts/autogen_hopper_fna.py:152
↓ 3 callers
Function
layout_separate
csrc/include/natten/cuda/fna_hopper/collective/fna_common.hpp:146
↓ 3 callers
Function
layout_separate
csrc/include/natten/cuda/fmha_hopper/collective/fmha_common.hpp:103
↓ 3 callers
Method
load
csrc/include/natten/cuda/fmha/kernel_backward.h:119
↓ 3 callers
Method
load_with_byte_offset
Loads a fragment from memory
csrc/include/natten/cuda/fmha/iterators/predicated_tile_iterator_residual_last.h:888
↓ 3 callers
Function
make_blackwell_fna_ops
(na_dim)
src/natten/_libnatten/torch_wrappers.py:657
↓ 3 callers
Function
make_cutlass_blackwell_fna_autograd_fn
(na_dim)
src/natten/backends/blackwell_fna.py:73
↓ 3 callers
Function
make_cutlass_fna_autograd_fn
(na_dim)
src/natten/backends/fna.py:75
↓ 3 callers
Function
make_cutlass_hopper_fna_autograd_fn
(na_dim)
src/natten/backends/hopper_fna.py:74
↓ 3 callers
Function
make_cutlass_token_permute_autograd_fn
(na_dim)
src/natten/token_permute/cutlass_impl.py:85
↓ 3 callers
Function
make_cutlass_token_unpermute_autograd_fn
(na_dim)
src/natten/token_permute/cutlass_impl.py:141
↓ 3 callers
Function
make_fna_ops
(na_dim)
src/natten/_libnatten/torch_wrappers.py:1021
↓ 3 callers
Function
make_hopper_fna_ops
(na_dim)
src/natten/_libnatten/torch_wrappers.py:839
↓ 3 callers
Function
make_reference_fna_autograd_fn
(na_dim)
src/natten/backends/reference.py:66
↓ 3 callers
Function
make_reference_fna_ops
(na_dim)
src/natten/_libnatten/torch_wrappers.py:1207
↓ 3 callers
Function
make_token_permute_ops
(na_dim)
src/natten/_libnatten/torch_wrappers.py:1373
↓ 3 callers
Function
parse_env_str
(env_var: str, default: str)
src/natten/utils/environment.py:53
↓ 3 callers
Function
print_table
(title, headers, values, has_footer=False)
src/natten/profiling_utils/pretty_printer.py:51
↓ 3 callers
Method
ref
Returns a TensorRef to the operand
csrc/include/natten/cuda/fmha/gemm/custom_mma_base.h:120
↓ 3 callers
Method
ref
Returns a TensorRef to the operand
csrc/include/natten/cuda/fna/gemm/custom_mma_base.h:120
↓ 3 callers
Function
render_table
Render the test status table to `out`. Returns done count.
scripts/testing/monitor.py:152
↓ 3 callers
Method
split_key_device
csrc/include/natten/cuda/fmha/kernel_backward.h:681
↓ 3 callers
Method
split_key_device_dim
csrc/include/natten/cuda/fna/kernel_backward.h:730
↓ 3 callers
Method
store
csrc/include/natten/cuda/fmha/kernel_backward.h:134
← previous
next →
201–300 of 2,059, ranked by callers