MCPcopy Create free account

hub / github.com/ByteDance-Seed/Triton-distributed / functions

Functions5,103 in github.com/ByteDance-Seed/Triton-distributed

↓ 1 callersFunctionbuild_kernel
(kernel_name="preprocess")
python/little_kernel/design/flashcomm_combine.py:547
↓ 1 callersFunctionbuild_kernel
()
python/little_kernel/design/flashcomm_dispatch_chunk.py:364
↓ 1 callersFunctionbuild_kernel
()
python/little_kernel/benchmark/gemm_sm100/gemm_level1.py:194
↓ 1 callersFunctionbuild_kernel
()
python/little_kernel/benchmark/gemm_sm100/gemm_level8.py:325
↓ 1 callersFunctionbuild_kernel
()
python/little_kernel/benchmark/gemm_sm100/gemm_level2.py:201
↓ 1 callersFunctionbuild_kernel
()
python/little_kernel/benchmark/gemm_sm100/gemm_level7.py:298
↓ 1 callersFunctionbuild_kernel
()
python/little_kernel/benchmark/gemm_sm100/gemm_level3.py:207
↓ 1 callersFunctionbuild_kernel
()
python/little_kernel/benchmark/gemm_sm100/gemm_level9.py:371
↓ 1 callersFunctionbuild_kernel
()
python/little_kernel/benchmark/gemm_sm100/gemm_level4.py:193
↓ 1 callersFunctionbuild_kernel
()
python/little_kernel/benchmark/gemm_sm100/gemm_level5.py:221
↓ 1 callersFunctionbuild_kernel
()
python/little_kernel/benchmark/gemm_sm100/gemm_level6.py:307
↓ 1 callersFunctionbuild_kernel
()
python/little_kernel/benchmark/gemm_sm90/gemm_v2.py:117
↓ 1 callersFunctionbuild_kernel
()
python/little_kernel/benchmark/gemm_sm90/gemm_v8.py:192
↓ 1 callersFunctionbuild_kernel
()
python/little_kernel/benchmark/gemm_sm90/gemm_v9.py:221
↓ 1 callersFunctionbuild_kernel
()
python/little_kernel/benchmark/gemm_sm90/gemm_v6.py:174
↓ 1 callersFunctionbuild_kernel
()
python/little_kernel/benchmark/gemm_sm90/gemm_v7.py:179
↓ 1 callersFunctionbuild_kernel
()
python/little_kernel/benchmark/gemm_sm90/gemm_v4.py:136
↓ 1 callersFunctionbuild_kernel
()
python/little_kernel/benchmark/gemm_sm90/gemm_v5.py:154
↓ 1 callersFunctionbuild_kernel
()
python/little_kernel/benchmark/gemm_sm90/gemm_v3.py:158
↓ 1 callersFunctionbuild_kernel
()
python/little_kernel/benchmark/gemm_sm90/gemm_v10.py:228
↓ 1 callersFunctionbuild_kernel
Build a launchable kernel.
python/little_kernel/benchmark/gemm_sm90/gemm_v1.py:162
↓ 1 callersMethodbuild_pymaca_project
(self, ext, project_dir)
python/setup.py:570
↓ 1 callersMethodbuild_pymxshmem_project
(self, ext, project_dir)
python/setup.py:594
↓ 1 callersFunctionbuiltin
(eval_return_type: Union[LLType, Callable], codegen_func: Callable, eval_arg_type: Optional[Callab
python/little_kernel/language/builtin_base.py:84
↓ 1 callersFunctioncalc_full_scatter_indices
(exp_indices, max_tokens, world_size)
python/triton_dist/test/amd/test_ep_a2a.py:174
↓ 1 callersFunctioncalc_full_scatter_indices
(exp_indices)
python/triton_dist/test/nvidia/test_ep_a2a.py:142
↓ 1 callersFunctioncalc_gather_index
( scatter_index: torch.Tensor, row_start: int, row_end: int, BLOCK_SIZE: int = 1024, )
python/triton_dist/test/nvidia/test_ep_moe_inference.py:81
↓ 1 callersFunctioncalc_gather_index
( exp_indices: torch.Tensor, row_start: int, row_end: int, BLOCK_SIZE: int = 1024, )
tutorials/04-deepseek-infer-all2all.py:288
↓ 1 callersFunctioncalc_gather_index_stable
(choosed_experts: torch.Tensor, topk, ntokens)
python/triton_dist/test/nvidia/test_ep_a2a.py:195
↓ 1 callersFunctioncalc_rank_in_warp_and_accumulate
(value: ll.int32, hist_ptr: ll.ptr[ll.int32])
python/little_kernel/design/flashcomm_compute.py:69
↓ 1 callersFunctioncalc_scatter_index_stable
(chosen_experts: torch.Tensor)
python/triton_dist/test/amd/test_ep_a2a.py:170
↓ 1 callersFunctioncalc_scatter_index_stable
(chosen_experts: torch.Tensor)
python/triton_dist/test/nvidia/test_ep_moe_inference.py:127
↓ 1 callersFunctioncalc_sorted_gather_index
( topk_ids: torch.Tensor, tp_size, num_experts: int, block_size_m: int, )
python/triton_dist/kernels/nvidia/allgather_group_gemm.py:169
↓ 1 callersMethodcheck_context
(self, bs, seq, q_nheads, kv_nheads, k_head_dim, v_head_dim)
python/triton_dist/kernels/nvidia/ulysses_sp_dispatch.py:528
↓ 1 callersFunctioncheck_correctness
(sp_group, args)
python/triton_dist/test/nvidia/test_llm_ulysess_gemm_all2all_intra_node.py:204
↓ 1 callersFunctioncheck_correctness
(sp_group, args)
python/triton_dist/test/nvidia/test_llm_ulysess_all2all_gemm_intra_node.py:179
↓ 1 callersFunctioncheck_correctness
(sp_group, args)
python/triton_dist/test/nvidia/test_llm_ulysess_post_attn_all2all_intra_node.py:159
↓ 1 callersFunctioncheck_correctness
Check correctness with random shapes (similar to flux)
python/triton_dist/test/nvidia/test_all_to_all_single_gemm.py:470
↓ 1 callersFunctioncheck_correctness
(sp_group, args)
python/triton_dist/test/nvidia/test_llm_ulysess_pre_attn_all2all_intra_node.py:234
↓ 1 callersMethodcheck_ctx
(self, seq_len, out_feat)
python/triton_dist/kernels/nvidia/ulysses_sp_infer_gemm_a2a.py:136
↓ 1 callersFunctioncheck_license
(file_path)
scripts/check_license.py:187
↓ 1 callersFunctioncheck_res
(res: tuple[torch.Tensor], ref: tuple[torch.Tensor], s: str)
python/triton_dist/kernels/nvidia/gdn.py:1027
↓ 1 callersFunctioncheck_swizzled
python/triton_dist/kernels/nvidia/threadblock_swizzle_ag_moe.cc:280
↓ 1 callersFunctioncheck_swizzled
(swizzled: List[Tuple[int, int]], token_cnts_per_rank_per_expert)
python/triton_dist/kernels/nvidia/threadblock_swizzle_ag_moe.py:164
↓ 1 callersFunctioncheck_with_golden
(golden_results: tuple, received_out_data: torch.Tensor, received_out_splits: torch.Tensor,
python/triton_dist/test/nvidia/test_all_to_all_vdev_2d_offset_inter_node.py:173
↓ 1 callersFunctionchunk_fwd_o
( q: torch.Tensor, k: torch.Tensor, v: torch.Tensor, h: torch.Tensor, g: Optional[torch.Te
python/triton_dist/kernels/nvidia/gdn.py:863
↓ 1 callersFunctionchunk_gated_delta_rule_fwd
( q: torch.Tensor, k: torch.Tensor, v: torch.Tensor, g: torch.Tensor, beta: torch.Tensor,
python/triton_dist/kernels/nvidia/gdn.py:926
↓ 1 callersFunctionchunk_gated_delta_rule_fwd_h
Params: initial_state: Return: final_state: fp32 [N, H, K, V] if output_final_state is True, otherwise None
python/triton_dist/kernels/nvidia/gdn.py:635
↓ 1 callersFunctionchunk_kkt_inv_ut_fused_triton
varlen mode k: [1, T, H, K] v: [1, T, H, V] beta: [1, T, H] g: [1, T, H] cu_seqlens: [B + 1] chunk_indices: [NT * 2] wher
python/triton_dist/kernels/nvidia/gdn.py:308
↓ 1 callersFunctioncodegen_cpp
(tree: ast.AST, ctx=None, emit_header=True)
python/little_kernel/codegen/codegen_base.py:275
↓ 1 callersFunctioncolorize_latency
Color latency - always green
python/triton_dist/profiler_utils.py:429
↓ 1 callersFunctioncolorize_memory
Color memory - always blue
python/triton_dist/profiler_utils.py:433
↓ 1 callersMethodcombine_preprocess
(self, M_recv=None)
python/triton_dist/layers/nvidia/ep_a2a_fused_layer.py:481
↓ 1 callersMethodcombine_token_intra_node_and_send
( self, input: torch.Tensor, ep_a2a_layout_desc: EPAllToAllLayoutDesc, )
python/triton_dist/layers/amd/ep_a2a_layer.py:490
↓ 1 callersMethodcombine_token_intra_node_and_send
(self, input: torch.Tensor, ep_a2a_layout_desc: EPAllToAllLayoutDesc)
python/triton_dist/layers/nvidia/ep_a2a_layer.py:516
↓ 1 callersMethodcomm_only_perf
(self, input: torch.Tensor, iters: int = 100, warmup_iters: int = 10, profile=False)
python/triton_dist/test/amd/test_ag_gemm_intra_node.py:144
↓ 1 callersFunctioncompile_cuda
Compile CUDA code to PTX or CUBIN using nvcc. Parameters ---------- code : str The CUDA source code. kernel_name : s
python/little_kernel/runtime/compiler.py:75
↓ 1 callersFunctioncompile_kernel
(func: triton.JITFunction, workspace: Path)
python/triton_dist/tools/compile_aot.py:251
↓ 1 callersMethodcompute_dispatch_layout
Compute dispatch layout: all-gather splits, cumsum -> scatter indices + send mask. Parameters ---------- topk_indice
python/little_kernel/design/flashcomm_ep_kernels.py:189
↓ 1 callersMethodcompute_offset_kernel
(self)
python/little_kernel/design/flashcomm_ep_kernels.py:115
↓ 1 callersMethodcompute_token_offset
Compute stable within-expert token offsets and expert counts. Parameters ---------- topk_indices : torch.Tensor[int3
python/little_kernel/design/flashcomm_ep_kernels.py:143
↓ 1 callersFunctionconsume_token
(token, ptr, _semantic=None)
python/triton_dist/kernels/nvidia/ep_all2all_fused.py:57
↓ 1 callersFunctionconsumer_all_reduce
(symm_buf, tile_signal, BLOCK_SIZE_M=16, BLOCK_SIZE_N=64, GROUP_SIZE_M=1, NUM_COMM_SMS=16)
python/triton_dist/kernels/amd/gemm_allreduce.py:337
↓ 1 callersFunctionconsumer_all_reduce_kernel
( symm_input_ptr, symm_ar_out_ptr, ar_out_ptr, # gemm_barrier_ptr, multi_st_barrier_ptr,
python/triton_dist/kernels/nvidia/gemm_allreduce.py:140
↓ 1 callersFunctionconsumer_all_reduce_load_store_kernel
(symm_input_ptr, symm_ar_out_ptr, ar_out_ptr, # gemm_barrier_ptr, M
python/triton_dist/kernels/nvidia/gemm_allreduce.py:227
↓ 1 callersFunctioncooperative_barrier_on_this_grid
triton implementation of cooperative_group::this_grid().sync() WARNING: use with care. better launch triton with launch_cooperative_grid=True to
python/triton_dist/kernels/amd/common_ops.py:107
↓ 1 callersFunctioncopy_1d_tilewise_kernel
(src_ptr, dst_ptr, # nelems, # BLOCK_SIZE: tl.conste
python/triton_dist/kernels/nvidia/gemm_allreduce.py:351
↓ 1 callersMethodcopy_and_reset_and_barrier_all_triton
(self, local_data)
python/triton_dist/kernels/nvidia/allgather_group_gemm.py:320
↓ 1 callersMethodcopy_externals
()
python/setup.py:202
↓ 1 callersFunctioncopy_file
Make source path a copy of the target path. if both source/target under base_dir, create link with relative path. why it's tricky for that?
python/build_helpers.py:69
↓ 1 callersFunctioncopy_kernel
( rank, local_buf_ptr, global_buf_ptr, M_per_rank, N, stride_local_m, stride_local
python/triton_dist/kernels/nvidia/allgather_gemm.py:48
↓ 1 callersFunctioncos_f32
Cosine function (approximate).
python/little_kernel/language/intrin/cuda_asm.py:249
↓ 1 callersFunctioncp_engine_producer_all_gather_full_mesh_pull
( rank, num_ranks, local_tensor: torch.Tensor, remote_tensor_buffers: List[torch.Tensor],
tutorials/02-intra-node-allgather.py:70
↓ 1 callersFunctioncp_engine_producer_all_gather_full_mesh_pull_inter_node
( rank, local_world_size, world_size, local_tensor: torch.Tensor, ag_buffer: list[torch.Te
python/triton_dist/kernels/nvidia/allgather.py:387
↓ 1 callersFunctioncp_engine_producer_all_gather_full_mesh_push
( rank, num_ranks, local_tensor: torch.Tensor, remote_tensor_buffers: List[torch.Tensor],
python/triton_dist/kernels/metax/allgather_gemm.py:44
↓ 1 callersFunctioncp_engine_producer_all_gather_numa_node_push
( rank, num_ranks, numa_world_size, local_tensor: torch.Tensor, remote_tensor_buffers: Lis
python/triton_dist/kernels/metax/allgather_gemm.py:102
↓ 1 callersFunctioncp_engine_producer_all_gather_put
(local_tensor, ag_buffer, signal_buffer, M_per_rank, N, signal_target, r
tutorials/07-overlapping-allgather-gemm.py:308
↓ 1 callersFunctioncp_engine_producer_all_gather_ring_push_2d_inter_node
( rank, num_local_ranks, num_ranks, local_tensor: torch.Tensor, remote_tensor_buffers: Lis
python/triton_dist/kernels/nvidia/allgather.py:232
↓ 1 callersFunctioncp_engine_producer_kv_all_gather
( k_shard: torch.Tensor, # [total_kv_shard, kv_head, head_dim] v_shard: torch.Tensor, # [total_kv_sh
python/triton_dist/kernels/nvidia/sp_ag_attention_intra_node.py:106
↓ 1 callersFunctioncp_engine_producer_kv_all_gather
( k_shard: torch.Tensor, # [total_kv_shard, kv_head, head_dim] v_shard: torch.Tensor, # [total_kv_sh
python/triton_dist/kernels/nvidia/sp_ag_attention_inter_node.py:193
↓ 1 callersMethodcreate
(ep_config: EPConfig, capacity: int = 2)
python/triton_dist/layers/amd/ep_a2a_layer.py:119
↓ 1 callersMethodcreate
(ep_config: EPConfig, capacity: int = 2)
python/triton_dist/layers/nvidia/ep_a2a_layer.py:119
↓ 1 callersMethodcreate_2d_descriptor
Create a 2D TMA descriptor. Parameters ---------- tensor : torch.Tensor Input tensor gme
python/little_kernel/runtime/tma_descriptor.py:168
↓ 1 callersFunctioncreate_all_to_all_context
( max_m: int, hidden: int, rank: int, num_tot_experts: int, WORLD_SIZE: int, experts_p
python/triton_dist/kernels/amd/low_latency_all_to_all.py:193
↓ 1 callersFunctioncreate_all_to_all_single_2d_context
( max_m: int, hidden_dim: int, rank: int, world_size: int, dtype=torch.bfloat16, )
python/triton_dist/kernels/nvidia/all_to_all_single_2d.py:145
↓ 1 callersFunctioncreate_gemm_rs_context
(max_M, N, rank, world_size,
tutorials/08-overlapping-gemm-reduce-scatter.py:98
↓ 1 callersFunctioncreate_ll_gemm_ar_context
(rank, world_size, local_world_size, max_M, N, dtype, MIN_BLOCK_SIZE_M=16, MIN_B
python/triton_dist/kernels/nvidia/gemm_allreduce.py:127
↓ 1 callersFunctioncreate_moe_ar_context
(rank, world_size, local_world_size, max_token_num, hidden_dim, num_experts, topk, input_dtype,
python/triton_dist/kernels/nvidia/moe_reduce_ar.py:316
↓ 1 callersFunctioncreate_reduce_scater_2d_ctx
for num_reduction_sms: tunable param, 16 sms are enough for H800 For H800, we overlap local reduce and inter-node p2p with intra-
tutorials/06-inter-node-reduce-scatter.py:170
↓ 1 callersFunctioncreate_sp_ag_attention_context_inter_node
( batch_size, q_head, kv_head, max_seqlen_k, max_q_shard_len, head_dim, input_dtyp
python/triton_dist/kernels/nvidia/sp_ag_attention_inter_node.py:57
↓ 1 callersFunctioncreate_sp_ag_attention_context_intra_node
( batch_size, q_head, kv_head, max_seqlen_k, max_q_shard_len, head_dim, input_dtyp
python/triton_dist/kernels/nvidia/sp_ag_attention_intra_node.py:60
↓ 1 callersFunctioncreate_tensor
Dont call this in main(), avoid call mxshmem_free(ptr) after mxshmem_finalize
shmem/mxshmem_bind/pymxshmem/src/pymxshmem.cc:88
↓ 1 callersFunctioncreate_ulysses_sp_pre_attn_comm_context
(bs: int, max_seq: int, q_nheads: int, k_head_dim: int, v_head_dim: int,
python/triton_dist/kernels/nvidia/ulysses_sp_dispatch.py:546
↓ 1 callersFunctioncuda_occupancy_max_activate_blocks_per_multiprocessor
(triton_func, num_warps, *func_args, **func_kwargs)
python/triton_dist/utils.py:847
↓ 1 callersFunctioncutlass_scaled_gemm_ar
(A, weight, scale_a, scale_b, out_dtype, tp_group)
python/triton_dist/test/nvidia/test_gemm_ar.py:79
↓ 1 callersFunctiondebug_wrapper
Wrap a pass function with debug output.
python/little_kernel/core/passes/pass_registry.py:39
↓ 1 callersFunctiondecorator
(fn: T)
python/triton_dist/jit.py:304
↓ 1 callersFunctiondecorator
(cls: type)
python/little_kernel/codegen/registries/special_struct_registry.py:79
↓ 1 callersMethoddeinit_gemm_barriers
(self)
python/triton_dist/kernels/nvidia/sp_ulysess_qkv_gemm_all2all.py:653
↓ 1 callersMethoddeinit_group_sync_barrier
(self)
python/triton_dist/kernels/nvidia/sp_ulysess_qkv_gemm_all2all.py:640
← previousnext →1,301–1,400 of 5,103, ranked by callers