MCPcopy Create free account

hub / github.com/SHI-Labs/NATTEN / functions

Functions2,059 in github.com/SHI-Labs/NATTEN

Methodload
Loads a fragment from memory
csrc/include/natten/cuda/fmha/iterators/predicated_tile_iterator_residual_last.h:1874
Methodload
Loads a fragment from memory
csrc/include/natten/cuda/fmha/iterators/predicated_tile_iterator_residual_last.h:2103
Methodload
csrc/include/natten/cuda/fna_blackwell/collective/sm100_fna_load_tma_warpspecialized.hpp:182
Methodload
csrc/include/natten/cuda/fna_blackwell/collective/sm100_fna_fwd_mainloop_tma_warpspecialized.hpp:306
Methodload
Loads a fragment from memory
csrc/include/natten/cuda/fna/iterators/predicated_tile_iterator.h:316
Methodload
Loads a fragment from memory
csrc/include/natten/cuda/fna/iterators/predicated_tile_iterator.h:570
Methodload
Loads a fragment from memory
csrc/include/natten/cuda/fna/iterators/predicated_tile_iterator.h:801
Methodload
Loads a fragment from memory
csrc/include/natten/cuda/fna/iterators/epilogue_predicated_tile_iterator.h:369
Methodload
Loads a fragment from memory
csrc/include/natten/cuda/fna/iterators/predicated_tile_iterator_residual_last.h:332
Methodload
Loads a fragment from memory
csrc/include/natten/cuda/fna/iterators/predicated_tile_iterator_residual_last.h:589
Methodload_kv_maybe_q
csrc/include/natten/cuda/fna_hopper/collective/fna_collective_tma_warpspecialized.hpp:349
Methodload_kv_maybe_q
csrc/include/natten/cuda/fna_hopper/collective/fna_collective_bwd_tma_warpspecialized.hpp:628
Methodload_kv_maybe_q
csrc/include/natten/cuda/fmha_hopper/collective/fmha_collective_tma_warpspecialized.hpp:307
Methodload_kv_maybe_q
csrc/include/natten/cuda/fmha_hopper/collective/fmha_collective_bwd_tma_warpspecialized.hpp:502
Methodload_maybe_q
csrc/include/natten/cuda/fna_hopper/collective/fna_collective_tma_warpspecialized.hpp:514
Methodload_maybe_q
csrc/include/natten/cuda/fna_hopper/collective/fna_collective_bwd_tma_warpspecialized.hpp:848
Methodload_maybe_q
csrc/include/natten/cuda/fmha_hopper/collective/fmha_collective_bwd_tma_warpspecialized.hpp:710
Methodload_with_byte_offset
csrc/include/natten/cuda/fmha/iterators/predicated_tile_iterator_residual_last.h:378
Methodload_with_byte_offset
Loads a fragment from memory
csrc/include/natten/cuda/fmha/iterators/predicated_tile_iterator_residual_last.h:658
Methodload_with_byte_offset
csrc/include/natten/cuda/fmha/iterators/predicated_tile_iterator_residual_last.h:1130
Methodload_with_byte_offset
Loads a fragment from memory
csrc/include/natten/cuda/fmha/iterators/predicated_tile_iterator_residual_last.h:1405
Methodload_with_byte_offset
Loads a fragment from memory
csrc/include/natten/cuda/fmha/iterators/predicated_tile_iterator_residual_last.h:1632
Methodload_with_byte_offset
csrc/include/natten/cuda/fna/iterators/predicated_tile_iterator.h:285
Methodload_with_byte_offset
Loads a fragment from memory
csrc/include/natten/cuda/fna/iterators/predicated_tile_iterator.h:564
Methodload_with_byte_offset
Loads a fragment from memory
csrc/include/natten/cuda/fna/iterators/epilogue_predicated_tile_iterator.h:296
Methodload_with_byte_offset
csrc/include/natten/cuda/fna/iterators/predicated_tile_iterator_residual_last.h:301
Methodload_with_byte_offset
Loads a fragment from memory
csrc/include/natten/cuda/fna/iterators/predicated_tile_iterator_residual_last.h:583
Methodload_with_byte_offset
Loads a fragment from memory
csrc/include/natten/cuda/fna/epilogue/predicated_tile_iterator.h:306
Methodload_with_pointer_offset
csrc/include/natten/cuda/fmha/iterators/predicated_tile_iterator_residual_last.h:372
Methodload_with_pointer_offset
Loads a fragment from memory
csrc/include/natten/cuda/fmha/iterators/predicated_tile_iterator_residual_last.h:652
Methodload_with_pointer_offset
csrc/include/natten/cuda/fmha/iterators/predicated_tile_iterator_residual_last.h:1124
Methodload_with_pointer_offset
Loads a fragment from memory
csrc/include/natten/cuda/fmha/iterators/predicated_tile_iterator_residual_last.h:1399
Methodload_with_pointer_offset
Loads a fragment from memory
csrc/include/natten/cuda/fmha/iterators/predicated_tile_iterator_residual_last.h:1626
Methodload_with_pointer_offset
Loads a fragment from memory
csrc/include/natten/cuda/fmha/iterators/predicated_tile_iterator_residual_last.h:1868
Methodload_with_pointer_offset
Loads a fragment from memory
csrc/include/natten/cuda/fmha/iterators/predicated_tile_iterator_residual_last.h:2097
Methodload_with_pointer_offset
csrc/include/natten/cuda/fna/iterators/predicated_tile_iterator.h:279
Methodload_with_pointer_offset
Loads a fragment from memory
csrc/include/natten/cuda/fna/iterators/predicated_tile_iterator.h:558
Methodload_with_pointer_offset
csrc/include/natten/cuda/fna/iterators/predicated_tile_iterator_residual_last.h:295
Methodload_with_pointer_offset
Loads a fragment from memory
csrc/include/natten/cuda/fna/iterators/predicated_tile_iterator_residual_last.h:577
Functionlog_or_raise_error
( msg: str, raise_error: bool = False, exception: Any = RuntimeError )
src/natten/utils/checks.py:40
Functionlog_test_name
(request)
tests/conftest.py:7
Functionmake_acc_into_op
csrc/include/natten/cuda/fna_hopper/collective/fna_common.hpp:291
Functionmake_acc_into_op
csrc/include/natten/cuda/fmha_hopper/collective/fmha_common.hpp:248
Functionmake_attn_tensor_from_input
(input_tensor: Tensor, attention_dim: int)
src/natten/utils/tensor.py:35
Methodmaximum_active_blocks
Computes the maximum number of active blocks per multiprocessor
csrc/include/natten/cuda/fmha_blackwell/device/fmha_sm100.hpp:108
Methodmaximum_active_blocks
Computes the maximum number of active blocks per multiprocessor
csrc/include/natten/cuda/utils/generic_cutlass_device.hpp:109
Methodmaximum_active_blocks
Computes the maximum number of active blocks per multiprocessor
csrc/include/natten/cuda/fna_hopper/device/fna_sm90.hpp:108
Methodmaximum_active_blocks
Computes the maximum number of active blocks per multiprocessor
csrc/include/natten/cuda/fna_blackwell/device/fna_sm100.hpp:108
Methodmaximum_active_blocks
Computes the maximum number of active blocks per multiprocessor
csrc/include/natten/cuda/fmha_hopper/device/fmha_sm90.hpp:108
Methodmma
csrc/include/natten/cuda/fmha_blackwell/collective/sm100_fmha_fwd_mainloop_tma_warpspecialized.hpp:302
Methodmma
csrc/include/natten/cuda/fna_blackwell/collective/sm100_fna_fwd_mainloop_tma_warpspecialized.hpp:330
Functionmulti_dim_tiling_mask
( b: IntTensor, h: IntTensor, q_idx: IntTensor, kv_idx: IntTensor, q_t
src/natten/backends/flex.py:409
Functionna1d
Computes 1-D neighborhood attention. GQA/MQA support (`heads != heads_kv`) is available. For now, `blackwell-fna` and `flex-fna` support GQA/
src/natten/functional.py:567
Functionna1d_cutlass_blackwell_fna
( query: Tensor, key: Tensor, value: Tensor, kernel_size: Dimension1DTypeOrDed, stride: Di
src/natten/backends/blackwell_fna.py:401
Functionna1d_cutlass_fna
( query: Tensor, key: Tensor, value: Tensor, kernel_size: Dimension1DTypeOrDed, stride: Di
src/natten/backends/fna.py:312
Functionna1d_cutlass_hopper_fna
( query: Tensor, key: Tensor, value: Tensor, kernel_size: Dimension1DTypeOrDed, stride: Di
src/natten/backends/hopper_fna.py:413
Functionna1d_flex
( query: Tensor, key: Tensor, value: Tensor, kernel_size: Dimension1DTypeOrDed, stride: Di
src/natten/backends/flex.py:712
Functionna1d_reference
( query: Tensor, key: Tensor, value: Tensor, kernel_size: Dimension1DTypeOrDed, stride: Di
src/natten/backends/reference.py:262
Functionna2d
Computes 2-D neighborhood attention. GQA/MQA support (`heads != heads_kv`) is available. For now, `blackwell-fna` and `flex-fna` support GQA/
src/natten/functional.py:756
Functionna2d_cutlass_blackwell_fna
( query: Tensor, key: Tensor, value: Tensor, kernel_size: Dimension2DTypeOrDed, stride: Di
src/natten/backends/blackwell_fna.py:435
Functionna2d_cutlass_fna
( query: Tensor, key: Tensor, value: Tensor, kernel_size: Dimension2DTypeOrDed, stride: Di
src/natten/backends/fna.py:348
Functionna2d_cutlass_hopper_fna
( query: Tensor, key: Tensor, value: Tensor, kernel_size: Dimension2DTypeOrDed, stride: Di
src/natten/backends/hopper_fna.py:447
Functionna2d_flex
( query: Tensor, key: Tensor, value: Tensor, kernel_size: Dimension2DTypeOrDed, stride: Di
src/natten/backends/flex.py:742
Functionna2d_reference
( query: Tensor, key: Tensor, value: Tensor, kernel_size: Dimension2DTypeOrDed, stride: Di
src/natten/backends/reference.py:290
Functionna3d
Computes 3-D neighborhood attention. GQA/MQA support (`heads != heads_kv`) is available. For now, `blackwell-fna` and `flex-fna` support GQA/
src/natten/functional.py:955
Functionna3d_cutlass_blackwell_fna
( query: Tensor, key: Tensor, value: Tensor, kernel_size: Dimension3DTypeOrDed, stride: Di
src/natten/backends/blackwell_fna.py:469
Functionna3d_cutlass_fna
( query: Tensor, key: Tensor, value: Tensor, kernel_size: Dimension3DTypeOrDed, stride: Di
src/natten/backends/fna.py:384
Functionna3d_cutlass_hopper_fna
( query: Tensor, key: Tensor, value: Tensor, kernel_size: Dimension3DTypeOrDed, stride: Di
src/natten/backends/hopper_fna.py:481
Functionna3d_flex
( query: Tensor, key: Tensor, value: Tensor, kernel_size: Dimension3DTypeOrDed, stride: Di
src/natten/backends/flex.py:772
Functionna3d_reference
( query: Tensor, key: Tensor, value: Tensor, kernel_size: Dimension3DTypeOrDed, stride: Di
src/natten/backends/reference.py:318
Methodop_name_formatted
(self)
src/natten/profiling_utils/formatting.py:158
Functionoperator%
csrc/include/natten/cuda/fmha_blackwell/common/pow_2.hpp:72
Functionoperator%
csrc/include/natten/cuda/fna_blackwell/common/pow_2.hpp:72
Functionoperator()
Perform a threadblock-scoped matrix multiply-accumulate
csrc/include/natten/cuda/fmha/gemm/custom_mma_pipelined.h:256
Functionoperator()
Perform a threadblock-scoped matrix multiply-accumulate
csrc/include/natten/cuda/fmha/gemm/mma_from_smem.h:547
Functionoperator()
Perform a threadblock-scoped matrix multiply-accumulate
csrc/include/natten/cuda/fmha/gemm/custom_mma_multistage.h:459
Functionoperator()
Perform a threadblock-scoped matrix multiply-accumulate
csrc/include/natten/cuda/fna/gemm/custom_mma_pipelined.h:256
Functionoperator()
Perform a threadblock-scoped matrix multiply-accumulate
csrc/include/natten/cuda/fna/gemm/mma_from_smem.h:548
Functionoperator()
Perform a threadblock-scoped matrix multiply-accumulate
csrc/include/natten/cuda/fna/gemm/custom_mma_multistage.h:459
Methodoperator()
Launches the kernel after first constructing Params internal state from supplied arguments.
csrc/include/natten/cuda/fmha_blackwell/device/fmha_sm100.hpp:262
Methodoperator()
csrc/include/natten/cuda/fmha_blackwell/kernel/sm100_fmha_fwd_kernel_tma_warpspecialized.hpp:284
Methodoperator()
csrc/include/natten/cuda/fmha_blackwell/kernel/fmha_kernel_bwd_convert.hpp:184
Methodoperator()
csrc/include/natten/cuda/fmha_blackwell/kernel/sm100_fmha_bwd_kernel_tma_warpspecialized.hpp:1880
Methodoperator()
csrc/include/natten/cuda/fmha_blackwell/kernel/fmha_kernel_bwd_sum_OdO.hpp:124
Methodoperator()
Launches the kernel after first constructing Params internal state from supplied arguments.
csrc/include/natten/cuda/utils/generic_cutlass_device.hpp:264
Methodoperator()
csrc/include/natten/cuda/fmha/epilogue/epilogue_rescale_output.h:157
Methodoperator()
csrc/include/natten/cuda/fmha/epilogue/epilogue_pipelined.h:242
Methodoperator()
csrc/include/natten/cuda/fmha/epilogue/epilogue_thread_apply_logsumexp.h:66
Methodoperator()
csrc/include/natten/cuda/fmha/epilogue/epilogue_thread_apply_logsumexp.h:82
Methodoperator()
csrc/include/natten/cuda/fmha/epilogue/epilogue_thread_apply_logsumexp.h:150
Methodoperator()
csrc/include/natten/cuda/fna_hopper/collective/fna_common.hpp:140
Methodoperator()
Launches the kernel after first constructing Params internal state from supplied arguments.
csrc/include/natten/cuda/fna_hopper/device/fna_sm90.hpp:263
Methodoperator()
Launches the kernel after first constructing Params internal state from supplied arguments.
csrc/include/natten/cuda/fna_blackwell/device/fna_sm100.hpp:262
Methodoperator()
csrc/include/natten/cuda/fna_blackwell/kernel/sm100_fna_bwd_kernel_tma_warpspecialized.hpp:1881
Methodoperator()
csrc/include/natten/cuda/fna_blackwell/kernel/sm100_fna_fwd_kernel_tma_warpspecialized.hpp:241
Methodoperator()
csrc/include/natten/cuda/fna/epilogue/epilogue_rescale_output.h:157
Methodoperator()
csrc/include/natten/cuda/fna/epilogue/epilogue_pipelined.h:242
Methodoperator()
csrc/include/natten/cuda/fna/epilogue/epilogue_thread_apply_logsumexp.h:66
Methodoperator()
csrc/include/natten/cuda/fna/epilogue/epilogue_thread_apply_logsumexp.h:82
Methodoperator()
csrc/include/natten/cuda/fna/epilogue/epilogue_thread_apply_logsumexp.h:150
← previousnext →1,501–1,600 of 2,059, ranked by callers