Code
Hub
Workspaces
Following
Trending
Connect
MCP
copy
Create free account
hub
/
github.com/Tencent/hpc-ops
/ functions
Functions
10,542 in github.com/Tencent/hpc-ops
⨍
Functions
10,542
◇
Types & classes
8,025
↳
Endpoints
13
↓ 1 callers
Method
get_buffer
( self, rank: int, *sizes: Any, dtype: Optional[torch.dtype] = None, storage_offset: int = 0 )
hpc/multicast_handle.py:102
↓ 1 callers
Function
get_checks
()
conftest.py:141
↓ 1 callers
Method
get_flat_coord
3rd/cutlass/include/cute/tensor_impl.hpp:334
↓ 1 callers
Method
get_grid_dims
Returns the grid extents in thread blocks to launch
3rd/cutlass/include/cutlass/gemm/kernel/gemm_streamk_with_fused_epilogue.h:1631
↓ 1 callers
Method
get_grid_shape
Computes the grid shape
3rd/cutlass/include/cutlass/gemm/device/rank_2k.h:494
↓ 1 callers
Method
get_grid_shape
Computes the grid shape
3rd/cutlass/include/cutlass/gemm/device/rank_k.h:457
↓ 1 callers
Method
get_grid_shape
Computes the grid shape
3rd/cutlass/include/cutlass/gemm/device/symm.h:549
↓ 1 callers
Method
get_mask
Gets the mask
3rd/cutlass/include/cutlass/transform/threadblock/predicated_tile_access_iterator_triangular_matrix.h:870
↓ 1 callers
Method
get_mask
Gets the mask
3rd/cutlass/include/cutlass/transform/threadblock/predicated_tile_access_iterator_2dthreadtile.h:816
↓ 1 callers
Method
get_scale_aux
3rd/cutlass/include/cutlass/epilogue/thread/linear_combination_generic_with_scaling.h:316
↓ 1 callers
Method
get_scale_d
3rd/cutlass/include/cutlass/epilogue/thread/linear_combination_generic_with_scaling.h:311
↓ 1 callers
Function
get_scaled_fp8_quant
()
benchmark/fused_moe/backends/base.py:85
↓ 1 callers
Method
get_smem_size
3rd/cutlass/include/cutlass/conv/device/direct_convolution.h:261
↓ 1 callers
Function
get_tilem
(avg_bs)
tests/test_group_gemm_blockwise.py:111
↓ 1 callers
Function
get_version
()
setup.py:44
↓ 1 callers
Function
get_work_k_tile_start
3rd/cutlass/include/cutlass/gemm/kernel/sm100_tile_scheduler_stream_k.hpp:572
↓ 1 callers
Function
get_workspace_size
3rd/cutlass/include/cutlass/transform/kernel/sm90_sparse_gemm_compressor.hpp:164
↓ 1 callers
Function
half_root_two<double>
3rd/cutlass/include/cutlass/constants.h:263
↓ 1 callers
Function
half_root_two<float>
3rd/cutlass/include/cutlass/constants.h:487
↓ 1 callers
Function
has_sk_work
Returns whether the current parameters contain any stream-K work
3rd/cutlass/include/cutlass/gemm/kernel/sm100_tile_scheduler_stream_k.hpp:908
↓ 1 callers
Function
heuristic_permutation
3rd/cutlass/include/cute/algorithm/cooperative_copy.hpp:79
↓ 1 callers
Method
hook
(self)
conftest.py:125
↓ 1 callers
Function
host_precompute
3rd/cutlass/include/cutlass/gemm/kernel/grouped_problem_visitor.h:421
↓ 1 callers
Method
imaginary_data
Returns the pointer to referenced data
3rd/cutlass/include/cutlass/tensor_ref_planar_complex.h:224
↓ 1 callers
Method
imaginary_stride
Gets the stride to an imaginary element
3rd/cutlass/include/cutlass/tensor_ref_planar_complex.h:246
↓ 1 callers
Function
increment
3rd/cutlass/include/cute/stride.hpp:490
↓ 1 callers
Method
init_workspace
Assign and initialize the specified workspace buffer. Assumes the memory allocated to workspace is at least as large as get_workspace_size().
3rd/cutlass/include/cutlass/gemm/kernel/params_universal_base.h:172
↓ 1 callers
Method
initialize
Initializes SYRK state from arguments.
3rd/cutlass/include/cutlass/gemm/device/rank_2k.h:235
↓ 1 callers
Method
initialize
Initializes SYRK state from arguments.
3rd/cutlass/include/cutlass/gemm/device/rank_k.h:212
↓ 1 callers
Method
initialize
Initializes TRMM state from arguments.
3rd/cutlass/include/cutlass/gemm/device/trmm.h:400
↓ 1 callers
Method
initialize
Initializes GEMM state from arguments.
3rd/cutlass/include/cutlass/gemm/device/gemm_splitk_parallel.h:271
↓ 1 callers
Method
initialize
Initializes GEMM state from arguments.
3rd/cutlass/include/cutlass/gemm/device/gemm_batched.h:366
↓ 1 callers
Method
initialize
Initializes GEMM state from arguments.
3rd/cutlass/include/cutlass/gemm/device/gemm_blockwise.h:414
↓ 1 callers
Method
initialize
Initializes GEMM state from arguments.
3rd/cutlass/include/cutlass/gemm/device/gemm.h:403
↓ 1 callers
Method
initialize
Initializes GEMM state from arguments.
3rd/cutlass/include/cutlass/gemm/device/gemm_universal_adapter.h:312
↓ 1 callers
Method
initialize
Initializes GEMM state from arguments.
3rd/cutlass/include/cutlass/gemm/device/gemm_complex.h:362
↓ 1 callers
Method
initialize
Initializes GEMM state from arguments.
3rd/cutlass/include/cutlass/gemm/device/gemm_array.h:371
↓ 1 callers
Method
initialize
Initializes SYMM state from arguments.
3rd/cutlass/include/cutlass/gemm/device/symm.h:242
↓ 1 callers
Function
initialize_preferred_cluster_launch
Cluster launch utility
3rd/cutlass/include/cute/arch/cluster_sm100.hpp:47
↓ 1 callers
Function
initialize_workspace
3rd/cutlass/include/cutlass/transform/kernel/sm90_sparse_gemm_compressor.hpp:172
↓ 1 callers
Function
inner_product
3rd/cutlass/include/cute/int_tuple.hpp:304
↓ 1 callers
Method
is_C_load_needed
3rd/cutlass/include/cutlass/epilogue/fusion/sm90_visitor_compute_tma_warpspecialized.hpp:279
↓ 1 callers
Function
is_dp_only
Returns whether the current parameters contain only data-parallel tiles
3rd/cutlass/include/cutlass/gemm/kernel/sm100_tile_scheduler_stream_k.hpp:894
↓ 1 callers
Method
is_final_split
3rd/cutlass/include/cutlass/gemm/kernel/sm90_tile_scheduler_stream_k.hpp:149
↓ 1 callers
Method
is_inf
3rd/cutlass/include/cutlass/exmy_base.h:467
↓ 1 callers
Method
is_nan
3rd/cutlass/include/cutlass/exmy_base.h:487
↓ 1 callers
Function
is_same_row
3rd/cutlass/include/cutlass/pipeline/sm100_pipeline.hpp:432
↓ 1 callers
Function
is_split_k
Returns whether the current parameters are for a split-K decomposition
3rd/cutlass/include/cutlass/gemm/kernel/sm100_tile_scheduler_stream_k.hpp:901
↓ 1 callers
Method
isfinite
Is finite implementation
3rd/cutlass/include/cutlass/float8.h:178
↓ 1 callers
Function
lift_dice
3rd/cutlass/include/cute/underscore.hpp:144
↓ 1 callers
Function
lift_slice
3rd/cutlass/include/cute/underscore.hpp:101
↓ 1 callers
Function
listen_tcp
src/communicator/listener.cc:27
↓ 1 callers
Function
listen_unix
src/communicator/listener.cc:62
↓ 1 callers
Function
load
Loads a fragment from memory at the location pointed to by the iterator.
3rd/cutlass/include/cutlass/gemm/warp/mma_complex_tensor_op_tile_iterator_sm80.h:246
↓ 1 callers
Function
load
Loads a fragment from memory at the location pointed to by the iterator.
3rd/cutlass/include/cutlass/gemm/warp/mma_tensor_op_tile_iterator_sm70.h:294
↓ 1 callers
Function
load
Loads a fragment from memory at the location pointed to by the iterator.
3rd/cutlass/include/cutlass/gemm/warp/mma_simt_tile_iterator.h:501
↓ 1 callers
Method
load_MK
3rd/cutlass/include/cutlass/gemm/collective/sm120_sparse_mma_tma.hpp:492
↓ 1 callers
Method
load_NK
3rd/cutlass/include/cutlass/gemm/collective/sm120_sparse_mma_tma.hpp:562
↓ 1 callers
Method
load_init
3rd/cutlass/include/cutlass/conv/collective/sm90_implicit_gemm_gmma_ss_warpspecialized.hpp:521
↓ 1 callers
Function
load_sf_update
3rd/cutlass/include/cutlass/gemm/collective/sm100_mma_array_warpspecialized_blockwise_scaling.hpp:646
↓ 1 callers
Function
load_symmetric_with_byte_offset
Loads a fragment on the diagonal of a symmetric kernel to memory
3rd/cutlass/include/cutlass/epilogue/threadblock/predicated_tile_iterator_blas3.h:322
↓ 1 callers
Function
load_with_byte_offset
Loads a fragment from memory
3rd/cutlass/include/cutlass/epilogue/threadblock/predicated_tile_iterator_direct_conv.h:289
↓ 1 callers
Function
load_with_byte_offset
Loads a fragment from memory
3rd/cutlass/include/cutlass/epilogue/threadblock/predicated_tile_iterator_blas3.h:263
↓ 1 callers
Function
load_with_byte_offset
Loads a fragment from memory
3rd/cutlass/include/cutlass/epilogue/threadblock/predicated_tile_iterator_conv.h:307
↓ 1 callers
Method
load_with_byte_offset
Loads a fragment from memory
3rd/cutlass/include/cutlass/transform/threadblock/ell_predicated_tile_iterator.h:877
↓ 1 callers
Method
load_with_byte_offset
Loads a fragment from memory
3rd/cutlass/include/cutlass/transform/threadblock/predicated_tile_iterator_triangular_matrix.h:782
↓ 1 callers
Function
load_with_pointer_offset
Loads a fragment from memory
3rd/cutlass/include/cutlass/transform/threadblock/predicated_scale_bias_vector_iterator.h:164
↓ 1 callers
Function
load_with_pointer_offset
Loads a fragment from memory
3rd/cutlass/include/cutlass/epilogue/threadblock/shared_load_iterator_pitch_linear.h:153
↓ 1 callers
Function
load_with_pointer_offset
Load
3rd/cutlass/include/cutlass/epilogue/warp/tile_iterator_tensor_op_mixed.h:281
↓ 1 callers
Method
load_with_pointer_offset
Loads a fragment from memory
3rd/cutlass/include/cutlass/transform/threadblock/predicated_tile_iterator_triangular_matrix.h:776
↓ 1 callers
Method
load_with_pointer_offset
Loads a fragment from memory
3rd/cutlass/include/cutlass/transform/threadblock/predicated_tile_iterator_2dthreadtile.h:757
↓ 1 callers
Method
load_with_pointer_offset
Loads a fragment from memory
3rd/cutlass/include/cutlass/epilogue/threadblock/shared_load_iterator_mixed.h:196
↓ 1 callers
Function
local_partition
3rd/cutlass/include/cute/tensor_impl.hpp:1077
↓ 1 callers
Function
log_2
3rd/cutlass/include/cute/numeric/math.hpp:319
↓ 1 callers
Function
main
()
benchmark/fuse_allreduce_rmsorm/benchmark_fuse_allreduce_rmsnorm.py:711
↓ 1 callers
Function
main
()
benchmark/fused_moe/benchmark_fuse_moe.py:456
↓ 1 callers
Function
main
()
benchmark/fused_moe/worker.py:19
↓ 1 callers
Function
main
()
benchmark/route_gemm/benchmark_gemm_bf16xfp32.py:394
↓ 1 callers
Function
main
()
benchmark/sampler/benchmark_sampler.py:459
↓ 1 callers
Function
main
()
benchmark/attention_decode/bench_attention_decode_fp8.py:503
↓ 1 callers
Method
make
3rd/cutlass/include/cute/atom/mma_traits_sm100.hpp:480
↓ 1 callers
Function
make_im2col_tma_atom_A_sm100
3rd/cutlass/include/cute/atom/copy_traits_sm100_im2col.hpp:385
↓ 1 callers
Function
make_im2col_tma_atom_B_sm100
3rd/cutlass/include/cute/atom/copy_traits_sm100_im2col.hpp:443
↓ 1 callers
Function
make_im2col_tma_copy_desc
3rd/cutlass/include/cute/atom/copy_traits_sm90_im2col.hpp:390
↓ 1 callers
Function
make_inputs
(N)
benchmark/fuse_allreduce_rmsorm/benchmark_fuse_allreduce_rmsnorm.py:545
↓ 1 callers
Function
make_inttuple_iter
3rd/cutlass/include/cute/numeric/arithmetic_tuple.hpp:209
↓ 1 callers
Function
make_rmem_ptr
3rd/cutlass/include/cute/pointer.hpp:236
↓ 1 callers
Function
make_runner
(key, N)
benchmark/fuse_allreduce_rmsorm/benchmark_fuse_allreduce_rmsnorm.py:558
↓ 1 callers
Function
make_tensor
3rd/cutlass/include/cute/pointer_flagged.hpp:58
↓ 1 callers
Function
make_tiler_impl
3rd/cutlass/include/cute/atom/partitioner.hpp:102
↓ 1 callers
Function
make_tma_atom_A_sm100
3rd/cutlass/include/cute/atom/copy_traits_sm100_tma.hpp:402
↓ 1 callers
Function
make_transform_iter
3rd/cutlass/include/cute/pointer_base.hpp:288
↓ 1 callers
Function
masked_softmax
3rd/cutlass/include/cutlass/epilogue/fusion/sm90_visitor_topk_softmax.hpp:310
↓ 1 callers
Method
mma
3rd/cutlass/include/cutlass/conv/collective/sm90_implicit_gemm_gmma_ss_warpspecialized.hpp:653
↓ 1 callers
Function
mma_unpack
3rd/cutlass/include/cute/atom/mma_traits.hpp:111
↓ 1 callers
Method
mnk
Obtains a Coord<3> from GemmCoord
3rd/cutlass/include/cutlass/gemm_coord.h:144
↓ 1 callers
Method
multiply
Elementwise multiply operator (1-by-2)
3rd/cutlass/include/cutlass/matrix.h:332
↓ 1 callers
Function
naive_act_mul_and_blockwise_quant
(gate_up_out)
tests/test_fuse_moe_blockwise.py:139
↓ 1 callers
Function
naive_act_mul_and_quant
(gate_up, scale, use_bf16_mul)
tests/test_fuse_moe_cp_async.py:90
↓ 1 callers
Function
naive_act_mul_and_quant
(gate_up, scale)
tests/test_fuse_moe_pertensor.py:95
← previous
next →
1,201–1,300 of 10,542, ranked by callers