Code
Hub
Workspaces
Following
Trending
Connect
MCP
copy
Create free account
hub
/
github.com/Tencent/hpc-ops
/ functions
Functions
10,542 in github.com/Tencent/hpc-ops
⨍
Functions
10,542
◇
Types & classes
8,025
↳
Endpoints
13
↓ 1 callers
Method
begin_step
Called at the start of one step before starting accumulator exchange
3rd/cutlass/include/cutlass/epilogue/threadblock/epilogue_with_visitor.h:110
↓ 1 callers
Function
bench_cuda_events
(fn, flush, warmup, iters)
benchmark/route_gemm/benchmark_gemm_bf16xfp32.py:39
↓ 1 callers
Function
bench_event
(call_fn, *, warmup: int, iters: int)
benchmark/sampler/benchmark_sampler.py:216
↓ 1 callers
Function
bench_nsys_provider
(provider, m, n, k, args, out_dir: Path)
benchmark/route_gemm/benchmark_gemm_bf16xfp32.py:144
↓ 1 callers
Function
bench_stable
(bench_fn, step_fn, warmup, iters, rounds, nvtx_label=None)
benchmark/fuse_allreduce_rmsorm/benchmark_fuse_allreduce_rmsnorm.py:171
↓ 1 callers
Function
bench_us_mean
(fn, warmup: int, iters: int)
benchmark/attention_decode/bench_attention_decode_fp8.py:249
↓ 1 callers
Function
benchmark_shape
(m, n, k, providers, args, flush)
benchmark/route_gemm/benchmark_gemm_bf16xfp32.py:202
↓ 1 callers
Function
build_block_ids
(kv_lens: torch.Tensor, block_size: int, max_num_blocks: int)
benchmark/attention_decode/bench_attention_decode_fp8.py:117
↓ 1 callers
Function
build_cells
(args)
benchmark/sampler/benchmark_sampler.py:337
↓ 1 callers
Method
build_extension
(self, ext)
setup.py:24
↓ 1 callers
Function
build_inputs
(batch: int, device: str = "cuda")
benchmark/sampler/benchmark_sampler.py:67
↓ 1 callers
Function
build_inputs
Create the common input bundle used by correctness tests.
tests/test_group_gemm_cp_async.py:18
↓ 1 callers
Function
build_routing
Sample uniform `topk_ids` and a normalized `topk_weights`. Returns: topk_ids : (num_seq, num_topk) int32, sorted along topk axis
benchmark/fused_moe/backends/base.py:133
↓ 1 callers
Function
calculate_errors
Calculate various error metrics between reference and real tensors Args: ref_tensor: Reference tensor (PyTorch tensor) real_
tests/utils.py:4
↓ 1 callers
Function
call
3rd/cutlass/include/cute/atom/copy_atom.hpp:92
↓ 1 callers
Function
can_implement
3rd/cutlass/include/cutlass/transform/kernel/sm90_sparse_gemm_compressor.hpp:153
↓ 1 callers
Function
ceil_div
3rd/cutlass/include/cute/int_tuple.hpp:324
↓ 1 callers
Function
check_runner
(key, r, N, ref_norm)
benchmark/fuse_allreduce_rmsorm/benchmark_fuse_allreduce_rmsnorm.py:567
↓ 1 callers
Function
clear
3rd/cutlass/include/cute/algorithm/clear.hpp:44
↓ 1 callers
Method
clear
< Efficiently disables all accesses guarded by mask
3rd/cutlass/include/cutlass/epilogue/threadblock/predicated_tile_iterator_predicates.h:145
↓ 1 callers
Method
clear
< Efficiently disables all accesses guarded by mask
3rd/cutlass/include/cutlass/epilogue/threadblock/predicated_tile_iterator_blas3.h:151
↓ 1 callers
Method
clear
Efficient clear method
3rd/cutlass/include/cute/container/array_subbyte.hpp:432
↓ 1 callers
Function
clear_mask
Clears the predicates
3rd/cutlass/include/cutlass/conv/threadblock/conv3d_dgrad_output_gradient_tile_access_iterator_optimized.h:414
↓ 1 callers
Function
clear_mask
Clears the predicates
3rd/cutlass/include/cutlass/conv/threadblock/conv3d_fprop_activation_tile_access_iterator_optimized.h:408
↓ 1 callers
Method
clear_mask
Clears the predicate set efficiently
3rd/cutlass/include/cutlass/transform/threadblock/predicated_tile_access_iterator_triangular_matrix.h:858
↓ 1 callers
Method
clear_mask
Clears the predicate set efficiently
3rd/cutlass/include/cutlass/transform/threadblock/predicated_tile_access_iterator_2dthreadtile.h:804
↓ 1 callers
Function
clear_mask_
Clears the predicates
3rd/cutlass/include/cutlass/conv/threadblock/conv3d_dgrad_output_gradient_tile_access_iterator_optimized.h:296
↓ 1 callers
Function
clear_mask_
Clears the predicates
3rd/cutlass/include/cutlass/conv/threadblock/conv3d_fprop_activation_tile_access_iterator_optimized.h:292
↓ 1 callers
Function
clz
3rd/cutlass/include/cutlass/fast_math.h:233
↓ 1 callers
Function
coalesce_256
3rd/cutlass/include/cute/atom/copy_traits_sm90_tma.hpp:719
↓ 1 callers
Function
compute_epilogue
Returns whether the block assigned this work should compute the epilogue for the corresponding output tile. For the case of stream-K, this should only
3rd/cutlass/include/cutlass/gemm/kernel/sm100_tile_scheduler_stream_k.hpp:377
↓ 1 callers
Function
conj_impl
3rd/cutlass/include/cutlass/complex.h:511
↓ 1 callers
Function
connect_tcp
src/communicator/connector.cc:57
↓ 1 callers
Function
connect_unix
src/communicator/connector.cc:31
↓ 1 callers
Function
consumer_release
Consumer signalling Producer of completion Ensures all blocks in the Same Row and Column get notifed.
3rd/cutlass/include/cutlass/pipeline/sm90_pipeline.hpp:627
↓ 1 callers
Function
consumer_release_2x1SM
3rd/cutlass/include/cutlass/pipeline/sm100_pipeline.hpp:271
↓ 1 callers
Method
consumer_test_wait
3rd/cutlass/include/cutlass/pipeline/sm90_pipeline.hpp:1150
↓ 1 callers
Function
consumer_wait
Wait for producer to commit transactions (done by TMA)
3rd/cutlass/include/cutlass/pipeline/sm90_pipeline.hpp:610
↓ 1 callers
Function
continue_current_work
Returns whether the current work_tile_info passed in should continue to be used.
3rd/cutlass/include/cutlass/gemm/kernel/sm100_tile_scheduler_stream_k.hpp:316
↓ 1 callers
Function
copy_tiles_and_advance
3rd/cutlass/include/cutlass/conv/threadblock/depthwise_fprop_direct_conv_multistage.h:88
↓ 1 callers
Method
copy_unpack_
3rd/cutlass/include/cute/atom/copy_traits_sm90_tma.hpp:507
↓ 1 callers
Function
cosize
3rd/cutlass/include/cute/layout_composed.hpp:272
↓ 1 callers
Function
countl_zero
3rd/cutlass/include/cute/numeric/math.hpp:226
↓ 1 callers
Function
crd2idx_horner
3rd/cutlass/include/cute/stride.hpp:129
↓ 1 callers
Function
crd2idx_itt
3rd/cutlass/include/cute/stride.hpp:77
↓ 1 callers
Function
crd2idx_ttt
3rd/cutlass/include/cute/stride.hpp:67
↓ 1 callers
Method
data
Returns the pointer to referenced data
3rd/cutlass/include/cutlass/tensor_ref_planar_complex.h:220
↓ 1 callers
Method
data
Returns a pointer to the shared memory buffer
3rd/cutlass/include/cutlass/epilogue/threadblock/epilogue_depthwise.h:164
↓ 1 callers
Method
data
Returns a pointer to the shared memory buffer
3rd/cutlass/include/cutlass/epilogue/threadblock/epilogue_planar_complex.h:181
↓ 1 callers
Function
depth
3rd/cutlass/include/cute/int_tuple.hpp:194
↓ 1 callers
Function
device_breakpoint
Triggers a breakpoint on the device
3rd/cutlass/include/cutlass/arch/arch.h:117
↓ 1 callers
Method
dot
Computes the dot product with anotherCoord object
3rd/cutlass/include/cutlass/coord.h:257
↓ 1 callers
Function
dry_run
Try a 1-step kernel call at minimal shape to confirm the backend actually works. Returns (ok, message).
benchmark/fused_moe/benchmark_fuse_moe.py:199
↓ 1 callers
Function
dump_test_py
(file_name, test_before_file, test_after_file, pypath)
conftest.py:27
↓ 1 callers
Function
elem_scale
3rd/cutlass/include/cute/int_tuple.hpp:423
↓ 1 callers
Method
enable
< CUTLASS_HOST_DEVICE enables all accesses guarded by mask
3rd/cutlass/include/cutlass/epilogue/threadblock/predicated_tile_iterator_predicates.h:153
↓ 1 callers
Method
enable
< CUTLASS_HOST_DEVICE enables all accesses guarded by mask
3rd/cutlass/include/cutlass/epilogue/threadblock/predicated_tile_iterator_affine.h:235
↓ 1 callers
Method
enable
< CUTLASS_HOST_DEVICE enables all accesses guarded by mask
3rd/cutlass/include/cutlass/epilogue/threadblock/predicated_tile_iterator_direct_conv.h:162
↓ 1 callers
Method
enable
< CUTLASS_HOST_DEVICE enables all accesses guarded by mask
3rd/cutlass/include/cutlass/epilogue/threadblock/predicated_tile_iterator_strided_dgrad.h:163
↓ 1 callers
Method
enable
< CUTLASS_HOST_DEVICE enables all accesses guarded by mask
3rd/cutlass/include/cutlass/epilogue/threadblock/predicated_tile_iterator_blas3.h:159
↓ 1 callers
Method
enable
< CUTLASS_HOST_DEVICE enables all accesses guarded by mask
3rd/cutlass/include/cutlass/epilogue/threadblock/predicated_tile_iterator_conv.h:201
↓ 1 callers
Method
enable_mask
Clears the predicate set efficiently
3rd/cutlass/include/cutlass/transform/threadblock/predicated_tile_access_iterator_triangular_matrix.h:862
↓ 1 callers
Method
enable_mask
Clears the predicate set efficiently
3rd/cutlass/include/cutlass/transform/threadblock/predicated_tile_access_iterator_2dthreadtile.h:808
↓ 1 callers
Method
end_epilogue
Called after all steps have been completed
3rd/cutlass/include/cutlass/epilogue/threadblock/epilogue_with_visitor.h:145
↓ 1 callers
Method
end_row
Called at the end of a row
3rd/cutlass/include/cutlass/epilogue/threadblock/epilogue_with_visitor.h:133
↓ 1 callers
Method
end_step
Called after all accumulator elements have been visited
3rd/cutlass/include/cutlass/epilogue/threadblock/epilogue_with_visitor.h:139
↓ 1 callers
Function
enrich_rows
(rows)
benchmark/fuse_allreduce_rmsorm/benchmark_fuse_allreduce_rmsnorm.py:420
↓ 1 callers
Function
epilogue_no_predication
3rd/cutlass/include/cute/algorithm/cooperative_gemm.hpp:106
↓ 1 callers
Function
equal_impl
3rd/cutlass/include/cute/container/tuple.hpp:566
↓ 1 callers
Function
error_metrics
(out, ref)
benchmark/route_gemm/benchmark_gemm_bf16xfp32.py:100
↓ 1 callers
Function
errors_to_string
Convert error calculation results to a human-readable string Args: error_results: Dictionary returned by calculate_errors function
tests/utils.py:93
↓ 1 callers
Method
exponent
3rd/cutlass/include/cutlass/exmy_base.h:540
↓ 1 callers
Method
exponent_bits
3rd/cutlass/include/cutlass/exmy_base.h:532
↓ 1 callers
Function
export_nvtx_csv
(report_file: Path, output_dir: Path, tag: str)
benchmark/sampler/benchmark_sampler.py:308
↓ 1 callers
Method
extra_metadata
(self)
benchmark/fused_moe/backends/base.py:53
↓ 1 callers
Function
extract_nvtx
(report_file: str)
benchmark/fused_moe/benchmark_fuse_moe.py:277
↓ 1 callers
Function
extract_nvtx_us
(report_prefix: Path)
benchmark/route_gemm/benchmark_gemm_bf16xfp32.py:74
↓ 1 callers
Function
extract_nvtx_us
(report_prefix: Path)
benchmark/attention_decode/bench_attention_decode_fp8.py:286
↓ 1 callers
Function
fast_assign_work
The fast path to get current output tile index then update fields of work tile info when continuing current work tile is needed, since k tile starting
3rd/cutlass/include/cutlass/gemm/kernel/sm90_tile_scheduler_stream_k.hpp:1045
↓ 1 callers
Function
fast_masked_softmax
3rd/cutlass/include/cutlass/epilogue/fusion/sm90_visitor_topk_softmax.hpp:271
↓ 1 callers
Function
fence_view_shared
3rd/cutlass/include/cutlass/arch/barrier.h:742
↓ 1 callers
Function
fetch_next_work
Kernel helper function to get next work tile
3rd/cutlass/include/cutlass/gemm/kernel/sm90_tile_scheduler_stream_k.hpp:705
↓ 1 callers
Function
fill
3rd/cutlass/include/cute/container/array.hpp:359
↓ 1 callers
Method
fill
3rd/cutlass/include/cute/container/array_subbyte.hpp:440
↓ 1 callers
Function
fill_int_tuple_from
3rd/cutlass/include/cute/int_tuple.hpp:663
↓ 1 callers
Method
find
3rd/cutlass/include/cute/container/type_list.hpp:55
↓ 1 callers
Function
fmt_speedup
(value)
benchmark/attention_decode/bench_attention_decode_fp8.py:441
↓ 1 callers
Function
fold_first
3rd/cutlass/include/cute/algorithm/tuple_algorithms.hpp:411
↓ 1 callers
Function
format_table
(hidden, rows, fi_backend)
benchmark/fuse_allreduce_rmsorm/benchmark_fuse_allreduce_rmsnorm.py:380
↓ 1 callers
Function
fuse_moe_cp_async_entry
src/fuse_moe/entry.cc:123
↓ 1 callers
Function
gemm_iters
Perform the specified number of threadblock mainloop iterations of matrix multiply-accumulate. Assumes prologue has been initiated.
3rd/cutlass/include/cutlass/gemm/threadblock/mma_multistage.h:613
↓ 1 callers
Function
generate_block_sparse_mask
Block-level sparse mask. True = attend. Diagonal always kept under causal.
tests/test_attention_blocksparse_qkpertoken_perhead_vperhead_fp8.py:17
↓ 1 callers
Function
generate_block_sparse_mask
Block-level sparse mask. True = attend.
tests/test_attention_blocksparse_qpertoken_perhead_kvpertensor_fp8.py:21
↓ 1 callers
Function
generate_cos_sin_cache
(max_position, head_dim, base=10000.0)
tests/test_rope.py:14
↓ 1 callers
Function
get
Returns a pointer to the vector starting at the current coordinate
3rd/cutlass/include/cutlass/conv/threadblock/conv2d_dgrad_output_gradient_tile_access_iterator_analytic.h:301
↓ 1 callers
Method
get
Returns a pointer
3rd/cutlass/include/cutlass/transform/threadblock/regular_tile_access_iterator_pitch_linear_direct_conv.h:553
↓ 1 callers
Method
get
Returns a pointer
3rd/cutlass/include/cutlass/transform/threadblock/regular_tile_access_iterator_pitch_linear.h:374
↓ 1 callers
Method
get
Returns a pointer
3rd/cutlass/include/cutlass/transform/threadblock/predicated_tile_access_iterator_triangular_matrix.h:827
↓ 1 callers
Method
get
Returns a pointer
3rd/cutlass/include/cutlass/transform/threadblock/predicated_tile_access_iterator_2dthreadtile.h:773
↓ 1 callers
Method
get
Returns a pointer
3rd/cutlass/include/cutlass/transform/threadblock/regular_tile_iterator_pitch_linear.h:406
← previous
next →
1,101–1,200 of 10,542, ranked by callers