MCPcopy Create free account

hub / github.com/Tencent/hpc-ops / functions

Functions10,542 in github.com/Tencent/hpc-ops

↓ 2 callersMethodgridDim
Commonly used utility functions
3rd/cutlass/include/cutlass/cluster_launch.hpp:96
↓ 2 callersFunctionimplicit_gemm_k_iterations_per_channel
3rd/cutlass/include/cutlass/conv/conv2d_problem_size.h:493
↓ 2 callersMethodinf_with_sign
3rd/cutlass/include/cutlass/exmy_base.h:603
↓ 2 callersMethodinit_params
Initialize params member
3rd/cutlass/include/cutlass/gemm/device/gemm_universal_base.h:208
↓ 2 callersMethodinverse
Computes the inverse of a 2-by-2 matrix given the matrix's determinant
3rd/cutlass/include/cutlass/matrix.h:3240
↓ 2 callersMethodis_C_load_needed
3rd/cutlass/include/cutlass/epilogue/fusion/sm90_visitor_compute_tma_warpspecialized.hpp:143
↓ 2 callersMethodis_f16c_supported
3rd/cutlass/include/cutlass/half.h:149
↓ 2 callersMethodis_producer_load_needed
3rd/cutlass/include/cutlass/epilogue/fusion/sm90_visitor_compute_tma_warpspecialized.hpp:138
↓ 2 callersMethodis_producer_load_needed
3rd/cutlass/include/cutlass/epilogue/collective/sm90_epilogue_tma_warpspecialized.hpp:413
↓ 2 callersMethodis_reduction_unit
3rd/cutlass/include/cutlass/gemm/kernel/sm90_tile_scheduler_stream_k.hpp:115
↓ 2 callersFunctionis_same_row_or_col
3rd/cutlass/include/cutlass/pipeline/sm100_pipeline.hpp:442
↓ 2 callersMethodis_source1_needed
3rd/cutlass/include/cutlass/epilogue/thread/linear_combination_tensor_broadcast.hpp:200
↓ 2 callersMethodis_source_needed
3rd/cutlass/include/cutlass/epilogue/collective/sm70_epilogue_vectorized.hpp:262
↓ 2 callersFunctionis_valid
3rd/cutlass/include/cute/util/type_traits.hpp:285
↓ 2 callersMethodis_valid
3rd/cutlass/include/cutlass/gemm/kernel/sm90_tile_scheduler_group.hpp:72
↓ 2 callersMethodis_zero
3rd/cutlass/include/cutlass/epilogue/fusion/sm90_visitor_load_tma_warpspecialized.hpp:313
↓ 2 callersFunctionisfinite
3rd/cutlass/include/cutlass/tfloat32.h:208
↓ 2 callersFunctionisfinite
3rd/cutlass/include/cutlass/half.h:510
↓ 2 callersFunctionisinf
3rd/cutlass/include/cutlass/half.h:521
↓ 2 callersMethodkn
Obtains a Coord<2> from GemmCoord
3rd/cutlass/include/cutlass/gemm_coord.h:186
↓ 2 callersFunctionlcm
3rd/cutlass/include/cutlass/fast_math.h:182
↓ 2 callersFunctionlcm_cxx11
3rd/cutlass/include/cutlass/fast_math.h:203
↓ 2 callersFunctionload
Loads a fragment
3rd/cutlass/include/cutlass/epilogue/threadblock/shared_load_iterator.h:210
↓ 2 callersMethodload_A
3rd/cutlass/include/cutlass/gemm/collective/sm100_mma_warpspecialized_mixed_input.hpp:744
↓ 2 callersMethodload_B
3rd/cutlass/include/cutlass/gemm/collective/sm100_mma_warpspecialized_mixed_input.hpp:825
↓ 2 callersMethodload_ab_tail
Perform a Producer Epilogue to prevent early exit of ctas in a Cluster
3rd/cutlass/include/cutlass/gemm/collective/sm100_mma_warpspecialized_blockwise_scaling.hpp:877
↓ 2 callersFunctionload_init
3rd/cutlass/include/cutlass/gemm/collective/sm100_mma_array_warpspecialized.hpp:506
↓ 2 callersMethodload_sf_tail
Perform a Producer Epilogue to prevent early exit of ctas in a Cluster
3rd/cutlass/include/cutlass/gemm/collective/sm100_mma_warpspecialized_blockwise_scaling.hpp:982
↓ 2 callersMethodload_tail
Perform a Producer Epilogue to prevent early exit of blocks in a Cluster
3rd/cutlass/include/cutlass/conv/collective/sm90_implicit_gemm_gmma_ss_warpspecialized.hpp:634
↓ 2 callersFunctionload_with_pointer_offset
Load
3rd/cutlass/include/cutlass/epilogue/warp/tile_iterator_volta_tensor_op.h:216
↓ 2 callersFunctionmac_loop_iter
Perform a threadblock mainloop iteration of matrix multiply-accumulate
3rd/cutlass/include/cutlass/gemm/threadblock/mma_multistage.h:495
↓ 2 callersMethodmake_fragment_C
3rd/cutlass/include/cute/atom/mma_atom.hpp:130
↓ 2 callersFunctionmake_inputs
(m, n, k, scale, device)
benchmark/route_gemm/benchmark_gemm_bf16xfp32.py:115
↓ 2 callersFunctionmake_inputs
( kv_lengths: Iterable[int], quant_name: str, num_seq_q: int, num_head_kv: int, num_head_q
benchmark/attention_decode/bench_attention_decode_fp8.py:145
↓ 2 callersFunctionmake_layout_like
3rd/cutlass/include/cute/layout.hpp:439
↓ 2 callersFunctionmake_ordered_layout
3rd/cutlass/include/cute/layout.hpp:423
↓ 2 callersFunctionmake_swizzle_strides
3rd/cutlass/include/cute/swizzle_layout.hpp:184
↓ 2 callersFunctionmake_task_map
(kv_lens: torch.Tensor, num_head_kv: int, num_seq_q: int, min_process_len: int)
benchmark/attention_decode/bench_attention_decode_fp8.py:128
↓ 2 callersFunctionmake_tiled_copy_C_atom
3rd/cutlass/include/cute/atom/copy_atom.hpp:451
↓ 2 callersFunctionmake_tma_copy_C_sm90
3rd/cutlass/include/cute/atom/copy_traits_sm90_tma.hpp:1580
↓ 2 callersMethodmk
Obtains a Coord<2> from GemmCoord
3rd/cutlass/include/cutlass/gemm_coord.h:168
↓ 2 callersFunctionnaive_group_gemm
(x, w, cu_seqlens, scale, expert_ids)
tests/test_fuse_moe_cp_async.py:65
↓ 2 callersFunctionnaive_group_gemm
(x, w, num_tokens_per_expert, cu_num_tokens_per_expert, xscale, wscale)
tests/test_fuse_moe_blockwise.py:78
↓ 2 callersFunctionnaive_group_gemm
(x, w, cu_seqlens, scale, expert_ids)
tests/test_fuse_moe_pertensor.py:73
↓ 2 callersFunctionnorm
3rd/cutlass/include/cutlass/quaternion.h:410
↓ 2 callersFunctionnullspace
3rd/cutlass/include/cute/layout.hpp:1483
↓ 2 callersMethodoperator++
3rd/cutlass/include/cutlass/transform/threadblock/predicated_tile_access_iterator.h:1025
↓ 2 callersMethodoperator++
3rd/cutlass/include/cutlass/transform/threadblock/regular_tile_access_iterator_pitch_linear_direct_conv.h:573
↓ 2 callersMethodoperator++
3rd/cutlass/include/cutlass/transform/threadblock/regular_tile_access_iterator_tensor_op.h:319
↓ 2 callersFunctionparse_int_list
(value: str)
benchmark/fused_moe/benchmark_fuse_moe.py:130
↓ 2 callersFunctionparse_proto
src/communicator/protocol.cc:10
↓ 2 callersFunctionparse_tcp
src/communicator/protocol.cc:23
↓ 2 callersFunctionparse_unix
src/communicator/protocol.cc:57
↓ 2 callersFunctionpenalty_temperature
()
benchmark/sampler/benchmark_sampler.py:134
↓ 2 callersFunctionpercentile
(values, pct)
benchmark/route_gemm/benchmark_gemm_bf16xfp32.py:29
↓ 2 callersFunctionpick_tile_m
src/group_gemm/cp_async/entry.cc:17
↓ 2 callersFunctionpopcount
3rd/cutlass/include/cute/numeric/math.hpp:270
↓ 2 callersFunctionprefetch
3rd/cutlass/include/cute/algorithm/prefetch.hpp:75
↓ 2 callersMethodprefetch_tma_descriptors
Issue Tma Descriptor Prefetch -- ideally from a single thread for best performance
3rd/cutlass/include/cutlass/conv/collective/sm90_implicit_gemm_gmma_ss_warpspecialized.hpp:506
↓ 2 callersFunctionprepare_inputs
Build all tensors required for rope_norm_store_kv[_fp8] tests. For prefill (is_prefill=True): variable Q tokens per request (random suffix sampl
tests/test_rope.py:148
↓ 2 callersFunctionpretty_print
3rd/cutlass/include/cute/util/print.hpp:207
↓ 2 callersFunctionpretty_print_float_exmy_base
3rd/cutlass/include/cute/numeric/numeric_types.hpp:186
↓ 2 callersFunctionprint_table
(rows: list[dict])
benchmark/attention_decode/bench_attention_decode_fp8.py:437
↓ 2 callersFunctionproduct_like
3rd/cutlass/include/cute/int_tuple.hpp:256
↓ 2 callersFunctionprologue
GEMM prologue. Bootstrap the global->shared memory pipeline by fetching the global fragments needed by the first kStages-1 threadblock mainloop itera
3rd/cutlass/include/cutlass/gemm/threadblock/mma_pipelined.h:245
↓ 2 callersMethodreduce_final
3rd/cutlass/include/cutlass/epilogue/fusion/sm90_visitor_topk_softmax.hpp:398
↓ 2 callersFunctionref_temperature_sample
PyTorch reference for temperature-only Gumbel-max sampling.
tests/test_sampler.py:431
↓ 2 callersFunctionreference_torch_rmsnorm_with_scale
(x, weight, scale, eps)
tests/test_normalization.py:13
↓ 2 callersFunctionreplace_back
3rd/cutlass/include/cute/algorithm/tuple_algorithms.hpp:686
↓ 2 callersFunctionreplace_front
3rd/cutlass/include/cute/algorithm/tuple_algorithms.hpp:671
↓ 2 callersFunctionrequires_separate_reduction
3rd/cutlass/include/cutlass/gemm/kernel/sm90_tile_scheduler_stream_k.hpp:390
↓ 2 callersFunctionretile
3rd/cutlass/include/cute/atom/copy_atom.hpp:285
↓ 2 callersFunctionrope_norm_ref
Unified PyTorch reference: RoPE + optional RMSNorm + paged KV write. Handles prefill, decode (mtp=0), and MTP decode (mtp>=1) uniformly via q_ind
tests/test_rope.py:47
↓ 2 callersFunctionrotr
3rd/cutlass/include/cute/numeric/math.hpp:214
↓ 2 callersFunctionrun_nsys_profile
(args: argparse.Namespace, case_name: str, quant_name: str, variant: str, out_dir: Path)
benchmark/attention_decode/bench_attention_decode_fp8.py:343
↓ 2 callersFunctionsave_data
(file_name, module_name, func_name, ret, args, kwargs)
conftest.py:15
↓ 2 callersFunctionset_activation_coord
3rd/cutlass/include/cutlass/conv/threadblock/depthwise_fprop_activation_tile_access_iterator_direct_conv_fixed_stride_dilation.h:71
↓ 2 callersFunctionset_iteration_index
Overrides the internal iteration index
3rd/cutlass/include/cutlass/conv/threadblock/depthwise_fprop_activation_tile_access_iterator_direct_conv_fixed_stride_dilation.h:199
↓ 2 callersMethodset_iteration_num
Overrides the internal iteration index
3rd/cutlass/include/cutlass/transform/threadblock/regular_tile_access_iterator_pitch_linear_direct_conv.h:541
↓ 2 callersMethodset_mask
Sets the predicate mask, overriding value stored in predicate iterator
3rd/cutlass/include/cutlass/transform/threadblock/predicated_tile_iterator_triangular_matrix.h:764
↓ 2 callersMethodset_mask
Sets the predicate mask, overriding value stored in predicate iterator
3rd/cutlass/include/cutlass/transform/threadblock/predicated_tile_iterator_2dthreadtile.h:745
↓ 2 callersMethodset_predicates
3rd/cutlass/include/cutlass/transform/threadblock/predicated_tile_access_iterator.h:186
↓ 2 callersMethodset_smem_base_address
Set base smem address
3rd/cutlass/include/cutlass/epilogue/threadblock/shared_load_iterator_mixed.h:239
↓ 2 callersFunctionsetup_call
(scene: str, provider: str, batch: int)
benchmark/sampler/benchmark_sampler.py:111
↓ 2 callersMethodsetup_initial_status
3rd/cutlass/include/cutlass/conv/warp/mma_depthwise_simt_tile_iterator.h:451
↓ 2 callersFunctionshape
3rd/cutlass/include/cute/tensor_impl.hpp:532
↓ 2 callersFunctionshared_load<2>
3rd/cutlass/include/cutlass/arch/memory.h:492
↓ 2 callersFunctionsignbit
3rd/cutlass/include/cutlass/half.h:495
↓ 2 callersFunctionsilu
(x)
tests/test_fuse_moe_cp_async.py:91
↓ 2 callersFunctionstore_with_pointer_offset
Store
3rd/cutlass/include/cutlass/epilogue/warp/tile_iterator_volta_tensor_op.h:182
↓ 2 callersMethodstride
Returns the layout object's stride vector
3rd/cutlass/include/cutlass/tensor_ref_planar_complex.h:256
↓ 2 callersFunctionstrided_dgrad_starting_coords
Computes starting Dx coord (h, w) for given starting filter postion
3rd/cutlass/include/cutlass/conv/conv2d_problem_size.h:634
↓ 2 callersMethodswapped_matrices
Returns arguments for the transposed matrices
3rd/cutlass/include/cutlass/gemm/kernel/trmm_universal.h:172
↓ 2 callersFunctionsync_internal
3rd/cutlass/include/cutlass/arch/barrier.h:324
↓ 2 callersFunctionsynclog_emit_cpasync_barrier_arrive
3rd/cutlass/include/cutlass/arch/synclog.hpp:938
↓ 2 callersFunctionsynclog_emit_fence_view_async_shared
3rd/cutlass/include/cutlass/arch/synclog.hpp:631
↓ 2 callersFunctionsynclog_emit_named_barrier_arrive
3rd/cutlass/include/cutlass/arch/synclog.hpp:366
↓ 2 callersFunctionsynclog_emit_named_barrier_arrive_and_wait
3rd/cutlass/include/cutlass/arch/synclog.hpp:346
↓ 2 callersMethodtensormaps_init
3rd/cutlass/include/cutlass/gemm/collective/sm90_mma_array_tma_gmma_ss_warpspecialized.hpp:627
↓ 2 callersMethodtile_finished
Whether this block will perform the last iteration of this tile
3rd/cutlass/include/cutlass/gemm/kernel/gemm_streamk_with_fused_epilogue.h:1695
← previousnext →901–1,000 of 10,542, ranked by callers