MCPcopy Create free account

hub / github.com/SHI-Labs/NATTEN / functions

Functions2,059 in github.com/SHI-Labs/NATTEN

↓ 2 callersFunctionget_window_right
csrc/include/natten/cuda/fna_blackwell/collective/fna_fusion.hpp:498
↓ 2 callersMethodheader
(self)
scripts/autogen_fna.py:179
↓ 2 callersMethodheader
(self)
scripts/autogen_fmha.py:160
↓ 2 callersMethodinit
csrc/include/natten/cuda/fna_hopper/collective/fna_collective_softmax.hpp:51
↓ 2 callersFunctionis_dilated
csrc/include/natten/cuda/fna_blackwell/collective/fna_fusion.hpp:596
↓ 2 callersFunctionis_neighbor
csrc/include/natten/cuda/fna_hopper/collective/fna_fusion.hpp:164
↓ 2 callersFunctionis_neighbor
csrc/include/natten/cuda/fna_blackwell/collective/fna_fusion.hpp:210
↓ 2 callersFunctionis_within_bounds
csrc/include/natten/cuda/fna_hopper/collective/fna_fusion.hpp:189
↓ 2 callersFunctionis_within_bounds
csrc/include/natten/cuda/fna_blackwell/collective/fna_fusion.hpp:235
↓ 2 callersFunctionkernel_type_to_str
(kernel_type: KernelType)
scripts/autogen_hopper_fmha.py:198
↓ 2 callersFunctionkernel_type_to_str
(kernel_type: KernelType)
scripts/autogen_hopper_fna.py:162
↓ 2 callersFunctionmake_qkv_stride
csrc/include/natten/cuda/reference/utils.hpp:46
↓ 2 callersFunctionmake_token_permuted_layout
csrc/include/natten/cuda/tokperm/layouts.hpp:43
↓ 2 callersFunctionmap_idx_to_di_coords
csrc/include/natten/cuda/reference/mask.hpp:259
↓ 2 callersFunctionmerge_attentions
Takes multiple attention *outputs* originating from the same query tensor, and their corresponding logsumexps, and merges them as if their context
src/natten/attn_merge.py:193
↓ 2 callersMethodoperator++
csrc/include/natten/cuda/fmha/iterators/predicated_tile_access_iterator_residual_last.h:853
↓ 2 callersMethodoperator++
csrc/include/natten/cuda/fna/iterators/predicated_tile_access_iterator_residual_last.h:935
↓ 2 callersMethodoperator++
csrc/include/natten/cuda/fna/iterators/predicated_tile_access_iterator.h:1621
↓ 2 callersFunctionprint_configs
(title, configs, keys)
src/natten/profiling_utils/dry_run.py:211
↓ 2 callersFunctionprint_summary
(status_dir, test_names, monitor_start, out=sys.stdout)
scripts/testing/monitor.py:202
↓ 2 callersFunctionread_status
(path)
scripts/testing/monitor.py:61
↓ 2 callersMethodreduce
csrc/include/natten/cuda/fna_hopper/collective/fna_collective_tma_warpspecialized.hpp:550
↓ 2 callersFunctionremove_wrapper
(s, wrapper)
src/natten/profiling_utils/formatting.py:227
↓ 2 callersFunctionrun_flex_attn
( q: Tensor, k: Tensor, v: Tensor, block_mask: BlockMask, scale: float, torch_compile:
src/natten/backends/flex.py:139
↓ 2 callersFunctionrun_optimize
(configs, run_backprop: bool)
src/natten/profiling_utils/dry_run.py:451
↓ 2 callersMethodset_iteration_index
Overrides the internal iteration index
csrc/include/natten/cuda/fna/iterators/predicated_tile_access_iterator_residual_last.h:887
↓ 2 callersMethodset_mask
Sets the predicate mask, overriding value stored in predicate iterator
csrc/include/natten/cuda/fna/iterators/predicated_tile_iterator.h:777
↓ 2 callersMethodsource
(self)
scripts/autogen_fna.py:185
↓ 2 callersMethodsource
(self)
scripts/autogen_fmha.py:166
↓ 2 callersMethodsplit_key_device
csrc/include/natten/cuda/fna/kernel_backward.h:722
↓ 2 callersMethodstep_interleave_step
csrc/include/natten/cuda/fna_hopper/collective/fna_collective_softmax.hpp:296
↓ 2 callersMethodstep_interleave_step
csrc/include/natten/cuda/fmha_hopper/collective/fmha_collective_softmax.hpp:229
↓ 2 callersMethodtail
csrc/include/natten/cuda/fna_hopper/collective/fna_collective_softmax.hpp:432
↓ 2 callersMethodtail
csrc/include/natten/cuda/fmha_hopper/collective/fmha_collective_softmax.hpp:331
↓ 2 callersFunctiontuple_sub
csrc/include/natten/cuda/fna_hopper/collective/fna_common.hpp:68
↓ 2 callersFunctiontuple_sub
csrc/include/natten/cuda/fna_blackwell/collective/fna_fusion.hpp:57
↓ 2 callersMethodupdate
Update API is preserved in 3.0, but does not guarantee a lightweight update of params.
csrc/include/natten/cuda/fna_hopper/device/fna_sm90.hpp:194
↓ 2 callersMethodvalid
Returns whether access is valid or not
csrc/include/natten/cuda/fna/iterators/predicated_tile_access_iterator_residual_last.h:966
↓ 2 callersFunctionwarp_uniform
csrc/include/natten/cuda/fna_blackwell/collective/fna_common.hpp:79
↓ 1 callersFunctionCheckIfBatchHeadsMatch
csrc/include/natten/helpers.h:276
↓ 1 callersFunctionCheckIfHeadDimsMatch
csrc/include/natten/helpers.h:293
↓ 1 callersMethod__str__
(self)
src/natten/profiling_utils/problem.py:209
↓ 1 callersFunction_format_time
Source: https://github.com/pytorch/pytorch/blob/orig/release/1.13/torch/autograd/profiler_util.py
src/natten/profiling_utils/formatting.py:88
↓ 1 callersFunction_generate_problem_shapes
Generate random but deterministic problem shapes for the test. Returns heads, head_dim, two normal seqlen lists, two empty seqlen lists, and
tests/test_varlen_recompile.py:51
↓ 1 callersFunction_get_default_tile_shapes_backward
( na_dim: int, )
src/natten/backends/configs/cutlass/__init__.py:241
↓ 1 callersFunction_get_default_tile_shapes_forward
( na_dim: int, )
src/natten/backends/configs/flex/__init__.py:64
↓ 1 callersFunction_get_default_tile_shapes_forward
( na_dim: int, )
src/natten/backends/configs/cutlass/__init__.py:68
↓ 1 callersFunction_get_default_tile_shapes_forward
( na_dim: int, )
src/natten/backends/configs/cutlass_blackwell/__init__.py:90
↓ 1 callersFunction_get_expected_attn_shape
(input_tensor: Tensor, attention_dim: int)
src/natten/utils/tensor.py:30
↓ 1 callersFunction_get_log_level
()
src/natten/utils/log.py:43
↓ 1 callersFunction_get_max_grid_size_allowed
()
src/natten/backends/configs/cutlass/backward_knobs.py:47
↓ 1 callersFunction_make_qkv
(total_seqlen, heads, head_dim)
tests/test_varlen_recompile.py:98
↓ 1 callersFunction_maybe_pad
( tensor: Tensor, tile_shape: DimensionType, dilation: DimensionType )
src/natten/token_permute/torch_impl.py:40
↓ 1 callersFunction_maybe_unpad
(tensor: Tensor, padding: DimensionType)
src/natten/token_permute/torch_impl.py:270
↓ 1 callersFunction_merge_attentions_compile
( outputs: List[Tensor], lse_tensors: List[Tensor] )
src/natten/attn_merge.py:109
↓ 1 callersFunction_print_ascii_table
(header, rows, title=None, footer=None)
src/natten/profiling_utils/pretty_printer.py:65
↓ 1 callersFunction_print_rich_table
(title, headers, values)
src/natten/profiling_utils/pretty_printer.py:43
↓ 1 callersFunction_profile_fmha_with_torch
( problem: Problem, warmup_steps: int, backend: Optional[str] = "cudnn", disable_backward: Opt
src/natten/profiling_utils/profiling.py:438
↓ 1 callersFunction_profile_na_with_torch
( problem: Problem, warmup_steps: int, backend: Optional[str] = None, fmha_backend: Optional[s
src/natten/profiling_utils/profiling.py:229
↓ 1 callersFunction_progress_path
()
scripts/testing/pytest_progress_plugin.py:13
↓ 1 callersFunction_reduce_max_kv_splits
( na_dim: int, kv_splits: DimensionType, max_splits: int, )
src/natten/backends/configs/cutlass/backward_knobs.py:70
↓ 1 callersFunction_reset
(max_recompiles=1)
tests/test_varlen_recompile.py:89
↓ 1 callersFunction_reset_torch_compile
()
tests/test_flex.py:51
↓ 1 callersFunction_run_fwd_bwd
(fn, q, k, v, cumseq_q, cumseq_kv, max_q, max_kv)
tests/test_varlen_recompile.py:122
↓ 1 callersMethod_test
( self, batch: int, heads: int, head_dim: int, seqlen_Q: int,
tests/test_attn_merge.py:124
↓ 1 callersMethod_test_all_dtypes_against_reference
(self, batch, seqlen, heads, head_dim)
tests/test_compute_delta.py:74
↓ 1 callersMethod_test_cutlass_blackwell_fmha_against_natten
( self, batch, heads, head_dim, seqlen_q, seqlen_kv, i
tests/test_fmha.py:821
↓ 1 callersMethod_test_cutlass_blackwell_fmha_against_torch_sdpa
( self, batch, heads, head_dim, seqlen_q, seqlen_kv, i
tests/test_fmha.py:573
↓ 1 callersMethod_test_cutlass_blackwell_fmha_varlen
( self, batch, heads, head_dim, seqlens_Q_list, seqlens_KV_lis
tests/test_fmha_varlen.py:513
↓ 1 callersMethod_test_cutlass_fmha_against_torch_sdpa
( self, batch, heads, head_dim, head_dim_v, seqlen_q,
tests/test_fmha.py:494
↓ 1 callersMethod_test_cutlass_fmha_determinism
( self, batch, heads, head_dim, seqlen_q, seqlen_kv, i
tests/test_fmha.py:749
↓ 1 callersMethod_test_cutlass_fmha_varlen
( self, batch, heads, head_dim, head_dim_v, seqlens_Q_list,
tests/test_fmha_varlen.py:347
↓ 1 callersMethod_test_cutlass_hopper_fmha_against_natten
( self, batch, heads, head_dim, seqlen_q, seqlen_kv, i
tests/test_fmha.py:692
↓ 1 callersMethod_test_cutlass_hopper_fmha_against_torch_sdpa
( self, batch, heads, head_dim, seqlen_q, seqlen_kv, i
tests/test_fmha.py:636
↓ 1 callersMethod_test_cutlass_hopper_fmha_varlen
( self, batch, heads, head_dim, seqlens_Q_list, seqlens_KV_lis
tests/test_fmha_varlen.py:449
↓ 1 callersMethod_test_fmha_module
( self, batch, seqlens_Q, seqlens_KV, num_heads, head_dim, is_causal, atol )
tests/test_torch_compile.py:310
↓ 1 callersMethod_test_na_module
( self, batch, token_layout_shape, num_heads, head_dim, kernel
tests/test_torch_compile.py:195
↓ 1 callersMethod_test_permute_cuda_kernel
( self, B, H, S, D, tile_shape, dilation, flip_dims, eps, dtype, device="cuda" )
tests/test_token_permute.py:280
↓ 1 callersFunction_token_permute
( tensor: Tensor, tile_shape: DimensionType, dilation: DimensionType, flip_tiled_dims: bool, )
src/natten/token_permute/torch_impl.py:99
↓ 1 callersFunction_token_unpermute
( tensor: Tensor, token_layout: DimensionType, tile_shape: DimensionType, dilation: DimensionT
src/natten/token_permute/torch_impl.py:190
↓ 1 callersFunctionadd_pointer_offset
csrc/include/natten/cuda/fna/iterators/predicated_tile_access_iterator_residual_last.h:281
↓ 1 callersMethodadd_pointer_offset
Adds a pointer offset in units of Element
csrc/include/natten/cuda/fna/iterators/predicated_tile_access_iterator.h:1584
↓ 1 callersMethodadd_pointer_offset
Adds a pointer offset in units of Element
csrc/include/natten/cuda/fna/iterators/predicated_tile_iterator_residual_last.h:517
↓ 1 callersMethodappend
(self, dtype: DataType)
scripts/autogen_hopper_fna_bwd.py:309
↓ 1 callersMethodappend
(self, dim: int)
scripts/autogen_hopper_fna_bwd.py:348
↓ 1 callersMethodappend
(self, cm)
scripts/autogen_hopper_fna_bwd.py:393
↓ 1 callersMethodappend
(self, config)
scripts/autogen_hopper_fna_bwd.py:457
↓ 1 callersMethodappend
(self, dim: int)
scripts/autogen_blackwell_fmha_bwd.py:332
↓ 1 callersMethodappend
(self, tile_size)
scripts/autogen_blackwell_fmha_bwd.py:377
↓ 1 callersMethodappend
(self, dtype: DataType)
scripts/autogen_blackwell_fna.py:300
↓ 1 callersMethodappend
(self, dim: int)
scripts/autogen_blackwell_fna.py:339
↓ 1 callersMethodappend
(self, cm)
scripts/autogen_blackwell_fna.py:384
↓ 1 callersMethodappend
(self, tile_shape)
scripts/autogen_blackwell_fna.py:446
↓ 1 callersMethodappend
(self, dtype: DataType)
scripts/autogen_blackwell_fna_bwd.py:304
↓ 1 callersMethodappend
(self, dim: int)
scripts/autogen_blackwell_fna_bwd.py:343
↓ 1 callersMethodappend
(self, cm)
scripts/autogen_blackwell_fna_bwd.py:388
↓ 1 callersMethodappend
(self, tile_shape)
scripts/autogen_blackwell_fna_bwd.py:452
↓ 1 callersMethodappend
(self, dtype: DataType)
scripts/autogen_reference_fna.py:336
↓ 1 callersMethodappend
(self, cm)
scripts/autogen_reference_fna.py:382
↓ 1 callersMethodappend
(self, dim: int)
scripts/autogen_blackwell_fmha.py:302
← previousnext →401–500 of 2,059, ranked by callers