MCPcopy Create free account

hub / github.com/Tencent/hpc-ops / functions

Functions10,542 in github.com/Tencent/hpc-ops

↓ 1 callersMethodrun
Runs the kernel using initialized state.
3rd/cutlass/include/cutlass/gemm/device/gemm_blockwise.h:483
↓ 1 callersMethodrun
Runs the kernel using initialized state.
3rd/cutlass/include/cutlass/gemm/device/gemm.h:473
↓ 1 callersMethodrun
Primary run() entry point API that is static allowing users to create and manage their own params. Supplied params struct must be construct by calling
3rd/cutlass/include/cutlass/gemm/device/gemm_universal_adapter.h:374
↓ 1 callersMethodrun
Runs the kernel using initialized state.
3rd/cutlass/include/cutlass/gemm/device/gemm_complex.h:428
↓ 1 callersMethodrun
Runs the kernel using initialized state.
3rd/cutlass/include/cutlass/gemm/device/gemm_array.h:429
↓ 1 callersMethodrun
Runs the kernel using initialized state.
3rd/cutlass/include/cutlass/gemm/device/symm.h:321
↓ 1 callersFunctionrun_method_c
Run warmup, graph capture, replay warmup, and timed graph replays. Timing is read from Nsight Systems' NVTX GPU projected duration. Do not sy
benchmark/fused_moe/backends/base.py:184
↓ 1 callersFunctionrun_method_c
FusedMoE-style timing worker: warmup, graph capture, replay under NVTX.
benchmark/attention_decode/bench_attention_decode_fp8.py:263
↓ 1 callersFunctionrun_nsys_driver
(args: argparse.Namespace)
benchmark/attention_decode/bench_attention_decode_fp8.py:408
↓ 1 callersFunctionrun_nsys_profile
Run worker.py under nsys profile.
benchmark/fused_moe/benchmark_fuse_moe.py:239
↓ 1 callersFunctionrun_nsys_profile
(args, cells: list[dict], output_dir: Path, tag: str)
benchmark/sampler/benchmark_sampler.py:259
↓ 1 callersFunctionrun_nsys_steps
(fn, warmup, iters)
benchmark/route_gemm/benchmark_gemm_bf16xfp32.py:58
↓ 1 callersFunctionrun_nsys_steps
(call_fn, *, warmup: int, iters: int, range_name: str = "step")
benchmark/sampler/benchmark_sampler.py:202
↓ 1 callersFunctionrun_nsys_worker
(args)
benchmark/route_gemm/benchmark_gemm_bf16xfp32.py:133
↓ 1 callersFunctionrun_nsys_worker
(args: argparse.Namespace)
benchmark/attention_decode/bench_attention_decode_fp8.py:313
↓ 1 callersFunctionrun_worker
(args)
benchmark/sampler/benchmark_sampler.py:231
↓ 1 callersFunctionsanitizer_check
(file_name, check)
conftest.py:74
↓ 1 callersFunctionselect_config
src/gemm/sm90/entry.cc:51
↓ 1 callersFunctionselect_split_k_by_work
src/gemm/sm90/entry.cc:29
↓ 1 callersFunctionselect_tile16_wgn
src/gemm/sm90/entry.cc:42
↓ 1 callersMethodset
Set a bit within the predicate vector.
3rd/cutlass/include/cutlass/predicate_vector.h:486
↓ 1 callersMethodset
3rd/cutlass/include/cutlass/gemm/device/ell_gemm.h:422
↓ 1 callersFunctionset_activation_coord
3rd/cutlass/include/cutlass/conv/threadblock/depthwise_fprop_activation_tile_access_iterator_direct_conv_optimized.h:68
↓ 1 callersFunctionset_iteration_index
Overrides the internal iteration index
3rd/cutlass/include/cutlass/conv/threadblock/conv2d_fprop_filter_tile_access_iterator_optimized.h:206
↓ 1 callersFunctionset_iteration_index
Overrides the internal iteration index
3rd/cutlass/include/cutlass/conv/threadblock/conv2d_fprop_activation_tile_access_iterator_analytic.h:183
↓ 1 callersFunctionset_iteration_index
Overrides the internal iteration index
3rd/cutlass/include/cutlass/conv/threadblock/conv2d_fprop_filter_tile_access_iterator_analytic.h:183
↓ 1 callersFunctionset_iteration_index
Overrides the internal iteration index
3rd/cutlass/include/cutlass/conv/threadblock/conv2d_dgrad_output_gradient_tile_access_iterator_analytic.h:225
↓ 1 callersFunctionset_iteration_index
Overrides the internal iteration index
3rd/cutlass/include/cutlass/conv/threadblock/conv2d_wgrad_output_gradient_tile_access_iterator_optimized.h:180
↓ 1 callersFunctionset_iteration_index
Overrides the internal iteration index
3rd/cutlass/include/cutlass/conv/threadblock/conv3d_wgrad_output_gradient_tile_access_iterator_optimized.h:187
↓ 1 callersFunctionset_iteration_index
3rd/cutlass/include/cutlass/transform/threadblock/regular_tile_access_iterator_tensor_op_sm80.h:65
↓ 1 callersMethodset_iteration_index
Overrides the internal iteration index
3rd/cutlass/include/cutlass/transform/threadblock/regular_tile_access_iterator_pitch_linear_direct_conv.h:537
↓ 1 callersMethodset_iteration_index
Overrides the internal iteration index
3rd/cutlass/include/cutlass/transform/threadblock/regular_tile_access_iterator_pitch_linear.h:364
↓ 1 callersMethodset_iteration_index
Overrides the internal iteration index
3rd/cutlass/include/cutlass/transform/threadblock/regular_scale_bias_vector_access_iterator.h:119
↓ 1 callersMethodset_iteration_index
Overrides the internal iteration index
3rd/cutlass/include/cutlass/transform/threadblock/predicated_tile_access_iterator_triangular_matrix.h:810
↓ 1 callersMethodset_iteration_index
Overrides the internal iteration index
3rd/cutlass/include/cutlass/transform/threadblock/predicated_tile_access_iterator_2dthreadtile.h:756
↓ 1 callersMethodset_iteration_index
Overrides the internal iteration index
3rd/cutlass/include/cutlass/epilogue/threadblock/predicated_tile_iterator.h:1005
↓ 1 callersFunctionset_mask
Sets the predicate mask, overriding value stored in predicate iterator
3rd/cutlass/include/cutlass/transform/threadblock/ell_predicated_tile_access_iterator.h:480
↓ 1 callersMethodset_mask
Sets the predicate mask, overriding value stored in predicate iterator
3rd/cutlass/include/cutlass/transform/threadblock/predicated_tile_access_iterator_triangular_matrix.h:866
↓ 1 callersMethodset_mask
Sets the predicate mask, overriding value stored in predicate iterator
3rd/cutlass/include/cutlass/transform/threadblock/predicated_tile_access_iterator_2dthreadtile.h:812
↓ 1 callersMethodset_metadata_bits
3rd/cutlass/include/cutlass/transform/kernel/sm90_sparse_gemm_compressor.hpp:228
↓ 1 callersMethodsetup
Build tensors and return the timed call_fn.
benchmark/fused_moe/backends/base.py:46
↓ 1 callersMethodsetup_separate_reduction
3rd/cutlass/include/cutlass/gemm/kernel/sm90_tile_scheduler_stream_k.hpp:130
↓ 1 callersFunctionshiftl
3rd/cutlass/include/cute/numeric/math.hpp:287
↓ 1 callersMethodshuffle_down_sync
Warp shuffle reduction
3rd/cutlass/include/cutlass/epilogue/fusion/sm90_visitor_topk_softmax.hpp:434
↓ 1 callersMethodshuffle_up_sync
3rd/cutlass/include/cutlass/epilogue/fusion/sm90_visitor_topk_softmax.hpp:377
↓ 1 callersMethodshuffle_xor_sync
Butterfly reduction
3rd/cutlass/include/cutlass/epilogue/fusion/sm90_visitor_topk_softmax.hpp:403
↓ 1 callersMethodsign_bit
3rd/cutlass/include/cutlass/exmy_base.h:514
↓ 1 callersMethodsignbit
Returns the sign bit
3rd/cutlass/include/cutlass/tfloat32.h:167
↓ 1 callersMethodsignbit
Returns the sign bit
3rd/cutlass/include/cutlass/bfloat16.h:205
↓ 1 callersMethodsignificand
3rd/cutlass/include/cutlass/exmy_base.h:617
↓ 1 callersFunctionsilu
(x)
tests/test_fuse_moe_blockwise.py:181
↓ 1 callersFunctionsilu
(x)
tests/test_fuse_moe_pertensor.py:97
↓ 1 callersFunctionsize
3rd/cutlass/include/cute/tensor_impl.hpp:550
↓ 1 callersMethodsize
3rd/cutlass/include/cute/container/array.hpp:159
↓ 1 callersFunctionslice
3rd/cutlass/include/cute/layout_composed.hpp:337
↓ 1 callersFunctionslice_and_offset
3rd/cutlass/include/cute/layout_composed.hpp:328
↓ 1 callersFunctionspec_from_args
(args)
benchmark/fused_moe/backends/base.py:233
↓ 1 callersFunctionstages_member
3rd/cutlass/include/cutlass/gemm/device/gemm_universal_adapter.h:104
↓ 1 callersFunctionstem_oam_gemm
Compute dim128 block_logits via OAM GEMM with fused causal mask epilogue. Performs block_logits = FrobScale * (Qflat @ Kflat^T) + V_bias, with
hpc/stem.py:123
↓ 1 callersFunctionstem_oam_prep_paged_kv
Precompute K_flat and V_bias from paged FP8 KV cache. First stage of the Stem sparse scoring pipeline: - K_flat (BF16): group-summed K vect
hpc/stem.py:16
↓ 1 callersFunctionstem_oam_prep_varlen_q
Precompute dim128 Q_flat from packed FP8 Q tensor. Computes weighted group-sum of Q tokens using per-token qscale (Q FP8 dequantization scale
hpc/stem.py:81
↓ 1 callersFunctionstem_tpd
Generate sparse block mask via top-k policy denoising. Fuses per-row budget (3-regime k_schedule + linear decay, both keyed on the full promp
hpc/stem.py:173
↓ 1 callersMethodstep
(self)
benchmark/fuse_allreduce_rmsorm/benchmark_fuse_allreduce_rmsnorm.py:304
↓ 1 callersFunctionstore
Stores a fragment to memory
3rd/cutlass/include/cutlass/epilogue/threadblock/predicated_tile_iterator.h:474
↓ 1 callersMethodstore_invalid_response
3rd/cutlass/include/cutlass/gemm/kernel/sm100_tile_scheduler.hpp:606
↓ 1 callersFunctionstore_symmetric_with_byte_offset
Stores a fragment on the diagonal of a symmetric kernel to memory
3rd/cutlass/include/cutlass/epilogue/threadblock/predicated_tile_iterator_blas3.h:473
↓ 1 callersFunctionstore_with_byte_offset
Stores a fragment to memory
3rd/cutlass/include/cutlass/epilogue/threadblock/predicated_tile_iterator_direct_conv.h:335
↓ 1 callersFunctionstore_with_byte_offset
Stores a fragment to memory
3rd/cutlass/include/cutlass/epilogue/threadblock/predicated_tile_iterator_blas3.h:420
↓ 1 callersFunctionstore_with_byte_offset
Stores a fragment to memory
3rd/cutlass/include/cutlass/epilogue/threadblock/predicated_tile_iterator_conv.h:365
↓ 1 callersMethodstore_with_byte_offset
Store a fragment to memory
3rd/cutlass/include/cutlass/transform/threadblock/ell_predicated_tile_iterator.h:905
↓ 1 callersMethodstore_with_byte_offset
Store a fragment to memory
3rd/cutlass/include/cutlass/transform/threadblock/predicated_tile_iterator_triangular_matrix.h:800
↓ 1 callersMethodstore_with_pointer_offset
Store a fragment to memory
3rd/cutlass/include/cutlass/transform/threadblock/predicated_tile_iterator_triangular_matrix.h:794
↓ 1 callersMethodstore_with_pointer_offset
Store a fragment to memory
3rd/cutlass/include/cutlass/transform/threadblock/predicated_tile_iterator_2dthreadtile.h:769
↓ 1 callersMethodstore_with_pointer_offset
Store
3rd/cutlass/include/cutlass/epilogue/warp/tile_iterator_tensor_op_mixed.h:1033
↓ 1 callersFunctionswap
3rd/cutlass/include/cute/container/array.hpp:366
↓ 1 callersMethodswap
Swaps the managed objects with *this and another unique_ptr
3rd/cutlass/include/cutlass/platform/platform.h:815
↓ 1 callersMethodswap
3rd/cutlass/include/cute/container/array.hpp:185
↓ 1 callersMethodsync
3rd/cutlass/include/cutlass/barrier.h:58
↓ 1 callersFunctionsynclog_condition_print
3rd/cutlass/include/cutlass/arch/synclog.hpp:236
↓ 1 callersFunctionsynclog_emit_cluster_barrier_arrive
3rd/cutlass/include/cutlass/arch/synclog.hpp:492
↓ 1 callersFunctionsynclog_emit_cluster_barrier_arrive_cluster
3rd/cutlass/include/cutlass/arch/synclog.hpp:469
↓ 1 callersFunctionsynclog_emit_cluster_barrier_init
3rd/cutlass/include/cutlass/arch/synclog.hpp:386
↓ 1 callersFunctionsynclog_emit_cluster_barrier_test_wait
3rd/cutlass/include/cutlass/arch/synclog.hpp:426
↓ 1 callersFunctionsynclog_emit_cluster_barrier_try_wait
3rd/cutlass/include/cutlass/arch/synclog.hpp:449
↓ 1 callersFunctionsynclog_emit_cluster_barrier_wait
3rd/cutlass/include/cutlass/arch/synclog.hpp:406
↓ 1 callersFunctionsynclog_emit_cluster_transaction_barrier_arrive_and_expect_tx
3rd/cutlass/include/cutlass/arch/synclog.hpp:526
↓ 1 callersFunctionsynclog_emit_cluster_transaction_barrier_complete_transaction
3rd/cutlass/include/cutlass/arch/synclog.hpp:592
↓ 1 callersFunctionsynclog_emit_cluster_transaction_barrier_expect_transaction
3rd/cutlass/include/cutlass/arch/synclog.hpp:572
↓ 1 callersFunctionsynclog_emit_cp_async
3rd/cutlass/include/cutlass/arch/synclog.hpp:751
↓ 1 callersFunctionsynclog_emit_cp_async_fence
3rd/cutlass/include/cutlass/arch/synclog.hpp:687
↓ 1 callersFunctionsynclog_emit_cp_async_nan
3rd/cutlass/include/cutlass/arch/synclog.hpp:700
↓ 1 callersFunctionsynclog_emit_cp_async_wait
3rd/cutlass/include/cutlass/arch/synclog.hpp:657
↓ 1 callersFunctionsynclog_emit_cp_async_wait_all
3rd/cutlass/include/cutlass/arch/synclog.hpp:674
↓ 1 callersFunctionsynclog_emit_cp_async_zfill
3rd/cutlass/include/cutlass/arch/synclog.hpp:724
↓ 1 callersFunctionsynclog_emit_fence_barrier_init
3rd/cutlass/include/cutlass/arch/synclog.hpp:618
↓ 1 callersFunctionsynclog_emit_fence_view_shared
3rd/cutlass/include/cutlass/arch/synclog.hpp:644
↓ 1 callersFunctionsynclog_emit_tma_store_arrive
3rd/cutlass/include/cutlass/arch/synclog.hpp:823
↓ 1 callersFunctionsynclog_emit_tma_store_wait
3rd/cutlass/include/cutlass/arch/synclog.hpp:836
↓ 1 callersFunctionsynclog_emit_warpgroup_arrive
3rd/cutlass/include/cutlass/arch/synclog.hpp:853
↓ 1 callersFunctionsynclog_emit_warpgroup_commit_batch
3rd/cutlass/include/cutlass/arch/synclog.hpp:884
← previousnext →1,401–1,500 of 10,542, ranked by callers