MCPcopy Create free account

hub / github.com/Tencent/hpc-ops / functions

Functions10,542 in github.com/Tencent/hpc-ops

↓ 1 callersFunctionsynclog_emit_warpgroup_wait
3rd/cutlass/include/cutlass/arch/synclog.hpp:867
↓ 1 callersFunctiontensormaps_cp_fence_release
3rd/cutlass/include/cutlass/gemm/collective/sm100_mma_array_warpspecialized_blockwise_scaling.hpp:1306
↓ 1 callersFunctiontensormaps_cp_fence_release_ab
3rd/cutlass/include/cutlass/gemm/collective/sm103_blockscaled_mma_array_warpspecialized.hpp:1524
↓ 1 callersFunctiontensormaps_cp_fence_release_sf
3rd/cutlass/include/cutlass/gemm/collective/sm103_blockscaled_mma_array_warpspecialized.hpp:1653
↓ 1 callersFunctiontensormaps_fence_acquire
3rd/cutlass/include/cutlass/gemm/collective/sm100_blockscaled_mma_array_warpspecialized.hpp:1338
↓ 1 callersFunctiontensormaps_fence_acquire
3rd/cutlass/include/cutlass/gemm/collective/sm100_mma_array_warpspecialized_planar_complex.hpp:939
↓ 1 callersFunctiontensormaps_fence_acquire
3rd/cutlass/include/cutlass/gemm/collective/sm100_mma_array_warpspecialized_blockwise_scaling.hpp:1322
↓ 1 callersFunctiontensormaps_fence_acquire
3rd/cutlass/include/cutlass/gemm/collective/sm100_mma_array_warpspecialized.hpp:915
↓ 1 callersFunctiontensormaps_fence_acquire
3rd/cutlass/include/cutlass/gemm/collective/sm100_mma_array_warpspecialized_rcggemm.hpp:879
↓ 1 callersFunctiontensormaps_fence_acquire
3rd/cutlass/include/cutlass/gemm/collective/sm100_blockscaled_mma_array_warpspecialized_rcggemm.hpp:1267
↓ 1 callersFunctiontensormaps_init
3rd/cutlass/include/cutlass/gemm/collective/sm100_blockscaled_mma_array_warpspecialized.hpp:1138
↓ 1 callersFunctiontensormaps_init
3rd/cutlass/include/cutlass/gemm/collective/sm100_mma_array_warpspecialized_blockwise_scaling.hpp:1192
↓ 1 callersFunctiontensormaps_init
3rd/cutlass/include/cutlass/gemm/collective/sm100_blockscaled_mma_array_warpspecialized_rcggemm.hpp:1103
↓ 1 callersFunctiontensormaps_init_ab
Methods to perform different parts of TMA/Tensormap modifications
3rd/cutlass/include/cutlass/gemm/collective/sm103_blockscaled_mma_array_warpspecialized.hpp:1410
↓ 1 callersFunctiontensormaps_init_sf
SF tensormap ops
3rd/cutlass/include/cutlass/gemm/collective/sm103_blockscaled_mma_array_warpspecialized.hpp:1540
↓ 1 callersFunctiontensormaps_replace_global_address
Replace address for the global tensor (to be done by single thread)
3rd/cutlass/include/cutlass/gemm/collective/sm100_blockscaled_mma_array_warpspecialized.hpp:1192
↓ 1 callersFunctiontensormaps_replace_global_address
Replace address for the global tensor (to be done by single thread)
3rd/cutlass/include/cutlass/gemm/collective/sm100_mma_array_warpspecialized_planar_complex.hpp:880
↓ 1 callersFunctiontensormaps_replace_global_address
Replace address for the global tensor (to be done by single thread)
3rd/cutlass/include/cutlass/gemm/collective/sm100_mma_array_warpspecialized_blockwise_scaling.hpp:1219
↓ 1 callersFunctiontensormaps_replace_global_address
Replace address for the global tensor (to be done by single thread)
3rd/cutlass/include/cutlass/gemm/collective/sm100_mma_array_warpspecialized_rcggemm.hpp:785
↓ 1 callersFunctiontensormaps_replace_global_address
Replace address for the global tensor (to be done by single thread)
3rd/cutlass/include/cutlass/gemm/collective/sm100_blockscaled_mma_array_warpspecialized_rcggemm.hpp:1151
↓ 1 callersFunctiontensormaps_replace_global_address_ab
Replace address for the global tensor (to be done by single thread)
3rd/cutlass/include/cutlass/gemm/collective/sm103_blockscaled_mma_array_warpspecialized.hpp:1438
↓ 1 callersFunctiontensormaps_replace_global_address_sf
Replace address for the global tensor (to be done by single thread)
3rd/cutlass/include/cutlass/gemm/collective/sm103_blockscaled_mma_array_warpspecialized.hpp:1567
↓ 1 callersFunctiontensormaps_replace_global_tensor_properties
3rd/cutlass/include/cutlass/gemm/collective/sm100_blockscaled_mma_array_warpspecialized.hpp:1212
↓ 1 callersFunctiontensormaps_replace_global_tensor_properties
3rd/cutlass/include/cutlass/gemm/collective/sm100_mma_array_warpspecialized_blockwise_scaling.hpp:1234
↓ 1 callersFunctiontensormaps_replace_global_tensor_properties
3rd/cutlass/include/cutlass/gemm/collective/sm100_mma_array_warpspecialized_rcggemm.hpp:798
↓ 1 callersFunctiontensormaps_replace_global_tensor_properties
3rd/cutlass/include/cutlass/gemm/collective/sm100_blockscaled_mma_array_warpspecialized_rcggemm.hpp:1167
↓ 1 callersFunctiontensormaps_replace_global_tensor_properties_ab
3rd/cutlass/include/cutlass/gemm/collective/sm103_blockscaled_mma_array_warpspecialized.hpp:1453
↓ 1 callersFunctiontensormaps_replace_global_tensor_properties_sf
3rd/cutlass/include/cutlass/gemm/collective/sm103_blockscaled_mma_array_warpspecialized.hpp:1582
↓ 1 callersFunctiontest_fuse_allreduce_rmsnorm_high_throughput
(world_size, N, H, num_max_blocks)
tests/test_fuse_allreduce_rmsnorm_high_throughput.py:106
↓ 1 callersFunctiontest_fuse_allreduce_rmsnorm_low_latency
(world_size, N, H, num_max_blocks)
tests/test_fuse_allreduce_rmsnorm_low_latency.py:125
↓ 1 callersFunctiontflops
(m, n, k, us)
benchmark/route_gemm/benchmark_gemm_bf16xfp32.py:35
↓ 1 callersFunctionthread
3rd/cutlass/include/cute/util/debug.hpp:132
↓ 1 callersFunctiontidfrg_D
3rd/cutlass/include/cute/atom/copy_atom.hpp:240
↓ 1 callersFunctiontidfrg_S
3rd/cutlass/include/cute/atom/copy_atom.hpp:219
↓ 1 callersMethodtile_count
3rd/cutlass/include/cutlass/gemm/kernel/grouped_problem_visitor.h:168
↓ 1 callersMethodtile_finished
Whether this block will perform the last iteration of this tile
3rd/cutlass/include/cutlass/gemm/kernel/gemm_universal_streamk.h:499
↓ 1 callersMethodtile_finished
3rd/cutlass/include/cutlass/gemm/kernel/gemm_universal_with_visitor_streamk.h:360
↓ 1 callersFunctiontile_peer_range
Returns the starting and ending peer ID of this tile
3rd/cutlass/include/cutlass/gemm/kernel/sm90_tile_scheduler_stream_k.hpp:1064
↓ 1 callersFunctiontmem_load_to_store
3rd/cutlass/include/cute/atom/copy_traits_sm100.hpp:3276
↓ 1 callersFunctionto_CUtensorMapFloatOOBfill
3rd/cutlass/include/cute/arch/copy_sm90_desc.hpp:265
↓ 1 callersFunctionto_CUtensorMapL2promotion
3rd/cutlass/include/cute/arch/copy_sm90_desc.hpp:274
↓ 1 callersFunctionto_atuple_i
3rd/cutlass/include/cute/numeric/arithmetic_tuple.hpp:294
↓ 1 callersFunctionto_linear_idx
3rd/cutlass/include/cutlass/gemm/kernel/sm100_tile_scheduler_stream_k.hpp:752
↓ 1 callersFunctionto_mixed_bits
3rd/cutlass/include/cute/swizzle.hpp:432
↓ 1 callersFunctiontopk_logsumexp
3rd/cutlass/include/cutlass/epilogue/fusion/sm90_visitor_topk_softmax.hpp:257
↓ 1 callersMethodtransform
3rd/cutlass/include/cutlass/transform/thread/transpose.h:60
↓ 1 callersMethodtruncate_to_cluster_size_m
Divides dividend by the cluster size in the M dimension
3rd/cutlass/include/cutlass/gemm/kernel/tile_scheduler_params.h:488
↓ 1 callersMethodtruncate_to_cluster_size_n
Divides dividend by the cluster size in the N dimension
3rd/cutlass/include/cutlass/gemm/kernel/tile_scheduler_params.h:500
↓ 1 callersFunctionumma_arrive_multicast
UMMA arrive for MMA_1sm + TMA_LOAD_MULTICAST combination
3rd/cutlass/include/cutlass/arch/barrier.h:828
↓ 1 callersFunctionunflatten
3rd/cutlass/include/cute/layout.hpp:542
↓ 1 callersFunctionunflatten_impl
3rd/cutlass/include/cute/algorithm/tuple_algorithms.hpp:584
↓ 1 callersMethodupdate
Lightweight update given a subset of arguments
3rd/cutlass/include/cutlass/gemm/device/rank_2k.h:293
↓ 1 callersMethodupdate
Lightweight update given a subset of arguments
3rd/cutlass/include/cutlass/gemm/device/ell_gemm.h:480
↓ 1 callersMethodupdate
Lightweight update given a subset of arguments
3rd/cutlass/include/cutlass/gemm/device/rank_k.h:270
↓ 1 callersMethodupdate
Lightweight update given a subset of arguments
3rd/cutlass/include/cutlass/gemm/device/gemm_splitk_parallel.h:317
↓ 1 callersMethodupdate
Lightweight update given a subset of arguments
3rd/cutlass/include/cutlass/gemm/device/gemm_batched.h:396
↓ 1 callersMethodupdate
Lightweight update given a subset of arguments
3rd/cutlass/include/cutlass/gemm/device/gemm_blockwise.h:463
↓ 1 callersMethodupdate
Lightweight update given a subset of arguments
3rd/cutlass/include/cutlass/gemm/device/gemm.h:454
↓ 1 callersMethodupdate
Update API is preserved in 3.0, but does not guarantee a lightweight update of params.
3rd/cutlass/include/cutlass/gemm/device/gemm_universal_adapter.h:359
↓ 1 callersMethodupdate
Lightweight update given a subset of arguments
3rd/cutlass/include/cutlass/gemm/device/gemm_complex.h:410
↓ 1 callersMethodupdate
Lightweight update given a subset of arguments
3rd/cutlass/include/cutlass/gemm/device/gemm_array.h:401
↓ 1 callersMethodupdate
Lightweight update given a subset of arguments
3rd/cutlass/include/cutlass/gemm/device/symm.h:301
↓ 1 callersFunctionvalid
Returns true if the current coordinate is within the output tensor Dy
3rd/cutlass/include/cutlass/conv/threadblock/conv2d_dgrad_output_gradient_tile_access_iterator_analytic.h:288
↓ 1 callersFunctionvalid
Returns whether access is valid or not
3rd/cutlass/include/cutlass/transform/threadblock/predicated_tile_access_iterator.h:668
↓ 1 callersMethodvalid
Returns whether access is valid or not
3rd/cutlass/include/cutlass/transform/threadblock/predicated_tile_access_iterator_triangular_matrix.h:880
↓ 1 callersMethodvalid
Returns whether access is valid or not
3rd/cutlass/include/cutlass/transform/threadblock/predicated_tile_access_iterator_2dthreadtile.h:820
↓ 1 callersFunctionvalid_
Returns true if the coord is within the output gradient tensor Dy
3rd/cutlass/include/cutlass/conv/threadblock/conv2d_wgrad_output_gradient_tile_access_iterator_optimized.h:240
↓ 1 callersFunctionvalid_
Returns true if the coord is within the output gradient tensor Dy
3rd/cutlass/include/cutlass/conv/threadblock/conv3d_wgrad_output_gradient_tile_access_iterator_optimized.h:243
↓ 1 callersMethodvisit
Called after accumulators have been exchanged for each accumulator vector
3rd/cutlass/include/cutlass/epilogue/threadblock/epilogue_with_visitor.h:122
↓ 1 callersFunctionwait
Waits until the semaphore is equal to the given value
3rd/cutlass/include/cutlass/semaphore.h:90
↓ 1 callersFunctionwarp_mma_planar_complex
3rd/cutlass/include/cutlass/gemm/threadblock/mma_planar_complex_pipelined.h:87
↓ 1 callersFunctionwarp_mma_planar_complex
3rd/cutlass/include/cutlass/gemm/threadblock/mma_planar_complex_multistage.h:310
↓ 1 callersFunctionweakly_congruent
3rd/cutlass/include/cute/int_tuple.hpp:454
↓ 1 callersFunctionwith
3rd/cutlass/include/cute/atom/copy_atom.hpp:78
↓ 1 callersMethodwork_tile_to_cluster_coord_mnkl
Convert CTA-level work tile info to cluster-level tile coord
3rd/cutlass/include/cutlass/gemm/kernel/sm100_tile_scheduler.hpp:379
↓ 1 callersFunctionwrite_csv
(path, rows, hidden, fi_backend, timing)
benchmark/fuse_allreduce_rmsorm/benchmark_fuse_allreduce_rmsnorm.py:443
↓ 1 callersFunctionwrite_csv
(path: str, rows: list[dict])
benchmark/fused_moe/benchmark_fuse_moe.py:436
↓ 1 callersFunctionwrite_csv
(path, rows)
benchmark/route_gemm/benchmark_gemm_bf16xfp32.py:320
↓ 1 callersFunctionwrite_csv
(path: str, rows: list[dict])
benchmark/sampler/benchmark_sampler.py:423
↓ 1 callersFunctionwrite_jsonl
(path, rows, hidden, fi_backend, timing)
benchmark/fuse_allreduce_rmsorm/benchmark_fuse_allreduce_rmsnorm.py:484
↓ 1 callersFunctionwrite_jsonl
(path: str, rows: list[dict])
benchmark/fused_moe/benchmark_fuse_moe.py:445
↓ 1 callersFunctionwrite_jsonl
(path, rows)
benchmark/route_gemm/benchmark_gemm_bf16xfp32.py:360
↓ 1 callersFunctionwrite_jsonl
(path: str, rows: list[dict])
benchmark/sampler/benchmark_sampler.py:432
FunctionAccumulatorPipelineState fixup
3rd/cutlass/include/cutlass/gemm/kernel/sm100_tile_scheduler_stream_k.hpp:640
MethodAccumulatorPipelineState fixup
3rd/cutlass/include/cutlass/gemm/kernel/sm100_tile_scheduler.hpp:553
MethodAccumulatorPipelineState fixup
3rd/cutlass/include/cutlass/gemm/kernel/sm100_tile_scheduler_group.hpp:253
MethodAccumulatorPipelineState fixup
3rd/cutlass/include/cutlass/gemm/kernel/sm100_static_tile_scheduler.hpp:197
FunctionAccumulatorPipelineState tmem_fixup
3rd/cutlass/include/cutlass/gemm/kernel/sm100_tile_scheduler_stream_k.hpp:923
MethodAffineRank2ColumnMajor
Ctor
3rd/cutlass/include/cutlass/layout/matrix.h:764
MethodAffineRank2RowMajor
Ctor
3rd/cutlass/include/cutlass/layout/matrix.h:872
MethodAffineRankN
Ctor
3rd/cutlass/include/cutlass/layout/matrix.h:624
MethodAllgather
src/communicator/communicator.cc:147
MethodAllocator1Sm
3rd/cutlass/include/cute/arch/tmem_allocator_sm100.hpp:65
MethodAllocator2Sm
3rd/cutlass/include/cute/arch/tmem_allocator_sm100.hpp:122
MethodApplySoftmaxFinalReduction
3rd/cutlass/include/cutlass/reduction/kernel/reduce_softmax_final.h:170
MethodArguments
Default ctor
3rd/cutlass/include/cutlass/gemm/device/ell_gemm.h:677
MethodArguments
Default ctor
3rd/cutlass/include/cutlass/gemm/device/gemm_splitk_parallel.h:208
MethodArguments
Default ctor
3rd/cutlass/include/cutlass/gemm/device/gemm_splitk_parallel.h:524
MethodArguments
Default ctor
3rd/cutlass/include/cutlass/gemm/device/gemm_batched.h:297
MethodArguments
Default ctor
3rd/cutlass/include/cutlass/gemm/device/gemm_batched.h:589
← previousnext →1,501–1,600 of 10,542, ranked by callers