Code
Hub
Workspaces
Following
Trending
Connect
MCP
copy
Create free account
hub
/
github.com/SHI-Labs/NATTEN
/ functions
Functions
2,059 in github.com/SHI-Labs/NATTEN
⨍
Functions
2,059
◇
Types & classes
573
Method
load
Loads a fragment from memory
csrc/include/natten/cuda/fmha/iterators/predicated_tile_iterator_residual_last.h:1874
Method
load
Loads a fragment from memory
csrc/include/natten/cuda/fmha/iterators/predicated_tile_iterator_residual_last.h:2103
Method
load
csrc/include/natten/cuda/fna_blackwell/collective/sm100_fna_load_tma_warpspecialized.hpp:182
Method
load
csrc/include/natten/cuda/fna_blackwell/collective/sm100_fna_fwd_mainloop_tma_warpspecialized.hpp:306
Method
load
Loads a fragment from memory
csrc/include/natten/cuda/fna/iterators/predicated_tile_iterator.h:316
Method
load
Loads a fragment from memory
csrc/include/natten/cuda/fna/iterators/predicated_tile_iterator.h:570
Method
load
Loads a fragment from memory
csrc/include/natten/cuda/fna/iterators/predicated_tile_iterator.h:801
Method
load
Loads a fragment from memory
csrc/include/natten/cuda/fna/iterators/epilogue_predicated_tile_iterator.h:369
Method
load
Loads a fragment from memory
csrc/include/natten/cuda/fna/iterators/predicated_tile_iterator_residual_last.h:332
Method
load
Loads a fragment from memory
csrc/include/natten/cuda/fna/iterators/predicated_tile_iterator_residual_last.h:589
Method
load_kv_maybe_q
csrc/include/natten/cuda/fna_hopper/collective/fna_collective_tma_warpspecialized.hpp:349
Method
load_kv_maybe_q
csrc/include/natten/cuda/fna_hopper/collective/fna_collective_bwd_tma_warpspecialized.hpp:628
Method
load_kv_maybe_q
csrc/include/natten/cuda/fmha_hopper/collective/fmha_collective_tma_warpspecialized.hpp:307
Method
load_kv_maybe_q
csrc/include/natten/cuda/fmha_hopper/collective/fmha_collective_bwd_tma_warpspecialized.hpp:502
Method
load_maybe_q
csrc/include/natten/cuda/fna_hopper/collective/fna_collective_tma_warpspecialized.hpp:514
Method
load_maybe_q
csrc/include/natten/cuda/fna_hopper/collective/fna_collective_bwd_tma_warpspecialized.hpp:848
Method
load_maybe_q
csrc/include/natten/cuda/fmha_hopper/collective/fmha_collective_bwd_tma_warpspecialized.hpp:710
Method
load_with_byte_offset
csrc/include/natten/cuda/fmha/iterators/predicated_tile_iterator_residual_last.h:378
Method
load_with_byte_offset
Loads a fragment from memory
csrc/include/natten/cuda/fmha/iterators/predicated_tile_iterator_residual_last.h:658
Method
load_with_byte_offset
csrc/include/natten/cuda/fmha/iterators/predicated_tile_iterator_residual_last.h:1130
Method
load_with_byte_offset
Loads a fragment from memory
csrc/include/natten/cuda/fmha/iterators/predicated_tile_iterator_residual_last.h:1405
Method
load_with_byte_offset
Loads a fragment from memory
csrc/include/natten/cuda/fmha/iterators/predicated_tile_iterator_residual_last.h:1632
Method
load_with_byte_offset
csrc/include/natten/cuda/fna/iterators/predicated_tile_iterator.h:285
Method
load_with_byte_offset
Loads a fragment from memory
csrc/include/natten/cuda/fna/iterators/predicated_tile_iterator.h:564
Method
load_with_byte_offset
Loads a fragment from memory
csrc/include/natten/cuda/fna/iterators/epilogue_predicated_tile_iterator.h:296
Method
load_with_byte_offset
csrc/include/natten/cuda/fna/iterators/predicated_tile_iterator_residual_last.h:301
Method
load_with_byte_offset
Loads a fragment from memory
csrc/include/natten/cuda/fna/iterators/predicated_tile_iterator_residual_last.h:583
Method
load_with_byte_offset
Loads a fragment from memory
csrc/include/natten/cuda/fna/epilogue/predicated_tile_iterator.h:306
Method
load_with_pointer_offset
csrc/include/natten/cuda/fmha/iterators/predicated_tile_iterator_residual_last.h:372
Method
load_with_pointer_offset
Loads a fragment from memory
csrc/include/natten/cuda/fmha/iterators/predicated_tile_iterator_residual_last.h:652
Method
load_with_pointer_offset
csrc/include/natten/cuda/fmha/iterators/predicated_tile_iterator_residual_last.h:1124
Method
load_with_pointer_offset
Loads a fragment from memory
csrc/include/natten/cuda/fmha/iterators/predicated_tile_iterator_residual_last.h:1399
Method
load_with_pointer_offset
Loads a fragment from memory
csrc/include/natten/cuda/fmha/iterators/predicated_tile_iterator_residual_last.h:1626
Method
load_with_pointer_offset
Loads a fragment from memory
csrc/include/natten/cuda/fmha/iterators/predicated_tile_iterator_residual_last.h:1868
Method
load_with_pointer_offset
Loads a fragment from memory
csrc/include/natten/cuda/fmha/iterators/predicated_tile_iterator_residual_last.h:2097
Method
load_with_pointer_offset
csrc/include/natten/cuda/fna/iterators/predicated_tile_iterator.h:279
Method
load_with_pointer_offset
Loads a fragment from memory
csrc/include/natten/cuda/fna/iterators/predicated_tile_iterator.h:558
Method
load_with_pointer_offset
csrc/include/natten/cuda/fna/iterators/predicated_tile_iterator_residual_last.h:295
Method
load_with_pointer_offset
Loads a fragment from memory
csrc/include/natten/cuda/fna/iterators/predicated_tile_iterator_residual_last.h:577
Function
log_or_raise_error
( msg: str, raise_error: bool = False, exception: Any = RuntimeError )
src/natten/utils/checks.py:40
Function
log_test_name
(request)
tests/conftest.py:7
Function
make_acc_into_op
csrc/include/natten/cuda/fna_hopper/collective/fna_common.hpp:291
Function
make_acc_into_op
csrc/include/natten/cuda/fmha_hopper/collective/fmha_common.hpp:248
Function
make_attn_tensor_from_input
(input_tensor: Tensor, attention_dim: int)
src/natten/utils/tensor.py:35
Method
maximum_active_blocks
Computes the maximum number of active blocks per multiprocessor
csrc/include/natten/cuda/fmha_blackwell/device/fmha_sm100.hpp:108
Method
maximum_active_blocks
Computes the maximum number of active blocks per multiprocessor
csrc/include/natten/cuda/utils/generic_cutlass_device.hpp:109
Method
maximum_active_blocks
Computes the maximum number of active blocks per multiprocessor
csrc/include/natten/cuda/fna_hopper/device/fna_sm90.hpp:108
Method
maximum_active_blocks
Computes the maximum number of active blocks per multiprocessor
csrc/include/natten/cuda/fna_blackwell/device/fna_sm100.hpp:108
Method
maximum_active_blocks
Computes the maximum number of active blocks per multiprocessor
csrc/include/natten/cuda/fmha_hopper/device/fmha_sm90.hpp:108
Method
mma
csrc/include/natten/cuda/fmha_blackwell/collective/sm100_fmha_fwd_mainloop_tma_warpspecialized.hpp:302
Method
mma
csrc/include/natten/cuda/fna_blackwell/collective/sm100_fna_fwd_mainloop_tma_warpspecialized.hpp:330
Function
multi_dim_tiling_mask
( b: IntTensor, h: IntTensor, q_idx: IntTensor, kv_idx: IntTensor, q_t
src/natten/backends/flex.py:409
Function
na1d
Computes 1-D neighborhood attention. GQA/MQA support (`heads != heads_kv`) is available. For now, `blackwell-fna` and `flex-fna` support GQA/
src/natten/functional.py:567
Function
na1d_cutlass_blackwell_fna
( query: Tensor, key: Tensor, value: Tensor, kernel_size: Dimension1DTypeOrDed, stride: Di
src/natten/backends/blackwell_fna.py:401
Function
na1d_cutlass_fna
( query: Tensor, key: Tensor, value: Tensor, kernel_size: Dimension1DTypeOrDed, stride: Di
src/natten/backends/fna.py:312
Function
na1d_cutlass_hopper_fna
( query: Tensor, key: Tensor, value: Tensor, kernel_size: Dimension1DTypeOrDed, stride: Di
src/natten/backends/hopper_fna.py:413
Function
na1d_flex
( query: Tensor, key: Tensor, value: Tensor, kernel_size: Dimension1DTypeOrDed, stride: Di
src/natten/backends/flex.py:712
Function
na1d_reference
( query: Tensor, key: Tensor, value: Tensor, kernel_size: Dimension1DTypeOrDed, stride: Di
src/natten/backends/reference.py:262
Function
na2d
Computes 2-D neighborhood attention. GQA/MQA support (`heads != heads_kv`) is available. For now, `blackwell-fna` and `flex-fna` support GQA/
src/natten/functional.py:756
Function
na2d_cutlass_blackwell_fna
( query: Tensor, key: Tensor, value: Tensor, kernel_size: Dimension2DTypeOrDed, stride: Di
src/natten/backends/blackwell_fna.py:435
Function
na2d_cutlass_fna
( query: Tensor, key: Tensor, value: Tensor, kernel_size: Dimension2DTypeOrDed, stride: Di
src/natten/backends/fna.py:348
Function
na2d_cutlass_hopper_fna
( query: Tensor, key: Tensor, value: Tensor, kernel_size: Dimension2DTypeOrDed, stride: Di
src/natten/backends/hopper_fna.py:447
Function
na2d_flex
( query: Tensor, key: Tensor, value: Tensor, kernel_size: Dimension2DTypeOrDed, stride: Di
src/natten/backends/flex.py:742
Function
na2d_reference
( query: Tensor, key: Tensor, value: Tensor, kernel_size: Dimension2DTypeOrDed, stride: Di
src/natten/backends/reference.py:290
Function
na3d
Computes 3-D neighborhood attention. GQA/MQA support (`heads != heads_kv`) is available. For now, `blackwell-fna` and `flex-fna` support GQA/
src/natten/functional.py:955
Function
na3d_cutlass_blackwell_fna
( query: Tensor, key: Tensor, value: Tensor, kernel_size: Dimension3DTypeOrDed, stride: Di
src/natten/backends/blackwell_fna.py:469
Function
na3d_cutlass_fna
( query: Tensor, key: Tensor, value: Tensor, kernel_size: Dimension3DTypeOrDed, stride: Di
src/natten/backends/fna.py:384
Function
na3d_cutlass_hopper_fna
( query: Tensor, key: Tensor, value: Tensor, kernel_size: Dimension3DTypeOrDed, stride: Di
src/natten/backends/hopper_fna.py:481
Function
na3d_flex
( query: Tensor, key: Tensor, value: Tensor, kernel_size: Dimension3DTypeOrDed, stride: Di
src/natten/backends/flex.py:772
Function
na3d_reference
( query: Tensor, key: Tensor, value: Tensor, kernel_size: Dimension3DTypeOrDed, stride: Di
src/natten/backends/reference.py:318
Method
op_name_formatted
(self)
src/natten/profiling_utils/formatting.py:158
Function
operator%
csrc/include/natten/cuda/fmha_blackwell/common/pow_2.hpp:72
Function
operator%
csrc/include/natten/cuda/fna_blackwell/common/pow_2.hpp:72
Function
operator()
Perform a threadblock-scoped matrix multiply-accumulate
csrc/include/natten/cuda/fmha/gemm/custom_mma_pipelined.h:256
Function
operator()
Perform a threadblock-scoped matrix multiply-accumulate
csrc/include/natten/cuda/fmha/gemm/mma_from_smem.h:547
Function
operator()
Perform a threadblock-scoped matrix multiply-accumulate
csrc/include/natten/cuda/fmha/gemm/custom_mma_multistage.h:459
Function
operator()
Perform a threadblock-scoped matrix multiply-accumulate
csrc/include/natten/cuda/fna/gemm/custom_mma_pipelined.h:256
Function
operator()
Perform a threadblock-scoped matrix multiply-accumulate
csrc/include/natten/cuda/fna/gemm/mma_from_smem.h:548
Function
operator()
Perform a threadblock-scoped matrix multiply-accumulate
csrc/include/natten/cuda/fna/gemm/custom_mma_multistage.h:459
Method
operator()
Launches the kernel after first constructing Params internal state from supplied arguments.
csrc/include/natten/cuda/fmha_blackwell/device/fmha_sm100.hpp:262
Method
operator()
csrc/include/natten/cuda/fmha_blackwell/kernel/sm100_fmha_fwd_kernel_tma_warpspecialized.hpp:284
Method
operator()
csrc/include/natten/cuda/fmha_blackwell/kernel/fmha_kernel_bwd_convert.hpp:184
Method
operator()
csrc/include/natten/cuda/fmha_blackwell/kernel/sm100_fmha_bwd_kernel_tma_warpspecialized.hpp:1880
Method
operator()
csrc/include/natten/cuda/fmha_blackwell/kernel/fmha_kernel_bwd_sum_OdO.hpp:124
Method
operator()
Launches the kernel after first constructing Params internal state from supplied arguments.
csrc/include/natten/cuda/utils/generic_cutlass_device.hpp:264
Method
operator()
csrc/include/natten/cuda/fmha/epilogue/epilogue_rescale_output.h:157
Method
operator()
csrc/include/natten/cuda/fmha/epilogue/epilogue_pipelined.h:242
Method
operator()
csrc/include/natten/cuda/fmha/epilogue/epilogue_thread_apply_logsumexp.h:66
Method
operator()
csrc/include/natten/cuda/fmha/epilogue/epilogue_thread_apply_logsumexp.h:82
Method
operator()
csrc/include/natten/cuda/fmha/epilogue/epilogue_thread_apply_logsumexp.h:150
Method
operator()
csrc/include/natten/cuda/fna_hopper/collective/fna_common.hpp:140
Method
operator()
Launches the kernel after first constructing Params internal state from supplied arguments.
csrc/include/natten/cuda/fna_hopper/device/fna_sm90.hpp:263
Method
operator()
Launches the kernel after first constructing Params internal state from supplied arguments.
csrc/include/natten/cuda/fna_blackwell/device/fna_sm100.hpp:262
Method
operator()
csrc/include/natten/cuda/fna_blackwell/kernel/sm100_fna_bwd_kernel_tma_warpspecialized.hpp:1881
Method
operator()
csrc/include/natten/cuda/fna_blackwell/kernel/sm100_fna_fwd_kernel_tma_warpspecialized.hpp:241
Method
operator()
csrc/include/natten/cuda/fna/epilogue/epilogue_rescale_output.h:157
Method
operator()
csrc/include/natten/cuda/fna/epilogue/epilogue_pipelined.h:242
Method
operator()
csrc/include/natten/cuda/fna/epilogue/epilogue_thread_apply_logsumexp.h:66
Method
operator()
csrc/include/natten/cuda/fna/epilogue/epilogue_thread_apply_logsumexp.h:82
Method
operator()
csrc/include/natten/cuda/fna/epilogue/epilogue_thread_apply_logsumexp.h:150
← previous
next →
1,501–1,600 of 2,059, ranked by callers