MCPcopy Create free account

hub / github.com/ByteDance-Seed/Triton-distributed / functions

Functions5,103 in github.com/ByteDance-Seed/Triton-distributed

↓ 5 callersFunctionget_dram_gbps
(device=None)
python/triton_dist/kernels/nvidia/gemm_perf_model.py:214
↓ 5 callersFunctionget_nvshmem_home
()
python/triton_dist/utils.py:586
↓ 5 callersFunctionget_source_lines
Get source lines from a function or source code string. Args: func_or_source: Either a function object or a string containing so
python/little_kernel/core/passes/utils/error_report.py:40
↓ 5 callersMethodinit_triton_dist_ctx
(self, max_M: int = 4096)
python/triton_dist/models/dense.py:169
↓ 5 callersFunctionipc_share_buffer
Share a RawCudaBuffer across processes via CUDA IPC. Uses int-array serialization to avoid null-byte truncation. Returns list[world_size] of d
python/little_kernel/design/test_flashcomm_torchrun.py:148
↓ 5 callersFunctionload_shared
(ptr: ll.Tensor[ll.uint64], offset: int)
python/little_kernel/benchmark/memory/bench_shared_memory.py:43
↓ 5 callersFunctionload_v4
(ptr, suffix: core.constexpr, _semantic=None)
python/triton_dist/kernels/nvidia/memory_ops.py:96
↓ 5 callersFunctionmake_ptr_tensor
(ptrs, device)
python/little_kernel/design/test_flashcomm_torchrun.py:185
↓ 5 callersMethodmake_rms_norm
(self, input: torch.Tensor, rms_weight: torch.Tensor, output: torch.Tensor, rms_eps: float = 1e-6,
python/triton_dist/mega_triton_kernel/models/model_builder.py:436
↓ 5 callersMethodmatch
Checks if a given integer value matches the filter's rule. Args: val: The integer value to check. Returns:
python/triton_dist/tools/tune/find_topk.py:80
↓ 5 callersMethodrecv
(self, ctx, rank, src_rank)
python/triton_dist/test/nvidia/test_pp.py:206
↓ 5 callersFunctionregister_test
(name)
python/triton_dist/test/metax/test_ag_gemm_intra_node.py:20
↓ 5 callersFunctionring_reduce
( input, # [M_per_node, N] output, # [M_per_rank, N] begin_idx, num_splits, num_sms=-1,
python/triton_dist/kernels/nvidia/reduce_scatter.py:780
↓ 5 callersFunctionrun_benchmark
A helper function to encapsulate the benchmarking logic.
python/triton_dist/test/nvidia/test_tp_mlp.py:96
↓ 5 callersFunctionst_v4_b32
(ptr, val0, val1, val2, val3, scope="", semantic="", _semantic=None)
python/triton_dist/language/extra/cuda/language_extra.py:248
↓ 5 callersFunctionstore_v4
(ptr, val0, val1, val2, val3, suffix: core.constexpr, _semantic=None)
python/triton_dist/kernels/nvidia/memory_ops.py:132
↓ 5 callersFunctiontile_kernel_moe_grouped_gemm_nk_const
( pid, num_pid, counter_ptr, barriers_ptr, # symm buf; producer: [max_num_tiles_m * num_block
python/triton_dist/kernels/nvidia/ep_all2all_fused.py:599
↓ 5 callersMethodupdate_phase
(self)
python/triton_dist/kernels/nvidia/gemm_allreduce.py:89
↓ 5 callersFunctionupdate_symlink
(link_path, source_path, materialization=False)
python/setup.py:361
↓ 4 callersFunctionCreateNVSHMEMOp
lib/Conversion/TritonDistributedToLLVM/NVIDIA/DistributedOpToLLVM.cpp:50
↓ 4 callersMethod__init__
(self, inner_type: LLType)
python/little_kernel/core/type_system.py:268
↓ 4 callersFunction__shfl_down_sync_i32
Shuffle down: each lane reads from (laneid + delta), clamped to 63.
python/triton_dist/language/extra/hip/language_extra.py:523
↓ 4 callersFunction__shfl_sync_with_mode_i32
( mask, value, delta, mode: core.constexpr = "up", c: core.constexpr = 31, _semantic=N
python/triton_dist/language/extra/cuda/language_extra.py:819
↓ 4 callersFunction_add_noise_workload_debug
()
python/triton_dist/kernels/nvidia/allgather.py:75
↓ 4 callersFunction_attn_fwd_inner
(acc, l_i, m_i, q, # desc_k, desc_v, # dtype: tl.constexpr, start_m,
python/triton_dist/mega_triton_kernel/kernels/flash_attn.py:33
↓ 4 callersFunction_barrier_all_intra_node_non_atomic_once_block
(local_rank, rank, local_world_size, symm_flags, target_value)
python/triton_dist/kernels/nvidia/common_ops.py:172
↓ 4 callersFunction_build_on_gpu
Build a kernel ensuring it's loaded on the specified GPU.
python/little_kernel/design/test_flashcomm_multi_gpu.py:301
↓ 4 callersFunction_bw
(nbytes, ms)
python/triton_dist/test/amd/test_ep_a2a.py:457
↓ 4 callersFunction_ce_p2p
no check dtype. no check device/host. no check tensor size. no check contiguous.
python/triton_dist/kernels/nvidia/ulysses_sp_infer_gemm_a2a.py:380
↓ 4 callersFunction_check
(out: torch.Tensor, ref: torch.Tensor, msg: str = "Triton")
python/triton_dist/test/amd/test_all_to_all.py:420
↓ 4 callersFunction_check
(out: torch.Tensor, ref: torch.Tensor, msg: str = "Triton")
python/triton_dist/test/nvidia/test_all_to_all.py:418
↓ 4 callersFunction_cp_engine_copy_data
(dst_ptr, src_ptr, cp_size, stream)
python/triton_dist/kernels/nvidia/sp_ag_attention_intra_node.py:130
↓ 4 callersFunction_ds_bpermute_b32
Low-level ds_bpermute_b32: read *value* from the lane at *byte_offset/4*.
python/triton_dist/language/extra/hip/language_extra.py:489
↓ 4 callersFunction_element_at
(x: tl.tensor, idx)
python/triton_dist/kernels/amd/ep_all2all_fused.py:698
↓ 4 callersFunction_i32
(lst)
python/triton_dist/kernels/amd/ep_all2all_fused.py:543
↓ 4 callersMethod_init_AR_ctx
(self, max_M, method: AllReduceMethod, dtype=torch.bfloat16)
python/triton_dist/layers/nvidia/tp_mlp.py:169
↓ 4 callersMethod_init_gemm_ar_ctx
(self, max_M, dtype=torch.bfloat16)
python/triton_dist/layers/nvidia/tp_mlp.py:198
↓ 4 callersMethod_init_triton_backend
Lazy initialization of Triton-dist backend
python/triton_dist/layers/nvidia/pp_block.py:135
↓ 4 callersFunction_is_cta_master
()
python/triton_dist/kernels/nvidia/common_ops.py:45
↓ 4 callersMethod_lltype_to_cpp
Convert an LLType instance to a valid C++ type string.
python/little_kernel/codegen/codegen_cuda.py:55
↓ 4 callersMethod_make_fc
(self, op_type: str, input: torch.Tensor, weight: torch.Tensor, output: torch.Tensor, layer_i
python/triton_dist/mega_triton_kernel/models/model_builder.py:216
↓ 4 callersFunction_memcpy_async_unsafe
no check dtype. no check device/host. no check tensor size. no check contiguous.
python/triton_dist/kernels/nvidia/allreduce.py:67
↓ 4 callersFunction_ptx_suffix_to_constraint
(suffix: core.constexpr, _semantic=None)
python/triton_dist/kernels/nvidia/memory_ops.py:35
↓ 4 callersFunction_putmem_impl
(dest, source, nbytes, pe, SCOPE_SUFFIX: core.constexpr, NBI: core.constexpr = core.constexpr(""),
python/triton_dist/language/extra/maca/libmxshmem_device.py:124
↓ 4 callersFunction_putmem_signal_impl
(dest, source, nbytes, sig_addr, signal, sig_op, pe, SCOPE_SUFFIX: core.constexpr, NBI
python/triton_dist/language/extra/maca/libmxshmem_device.py:167
↓ 4 callersFunction_set_cos_sin_cache
Precomputes cosine and sine cache for rotary position embeddings.
python/triton_dist/layers/amd/tp_attn.py:62
↓ 4 callersFunction_target_to_load_expr
Convert assignment target (Store ctx) to expression for use as Call argument (Load ctx). Required for A[0] etc.: ast.Name(id='A[0]') is invalid; m
python/little_kernel/core/passes/insert_mem_alloc.py:45
↓ 4 callersFunction_test_bisect
(side="left", aligned=False)
python/triton_dist/test/nvidia/test_common_ops.py:172
↓ 4 callersFunction_tile_ratio
(m, block_size_m)
python/triton_dist/kernels/amd/gemm.py:626
↓ 4 callersMethod_transform_special_struct_call
Transform special struct constructor call into struct declaration. Example: scheduler = Scheduler(...) b
python/little_kernel/core/passes/special_struct_materialize_pass.py:324
↓ 4 callersFunction_translate_scope
(scope)
python/triton_dist/language/extra/hip/language_extra.py:31
↓ 4 callersFunction_translate_semantic
(semantic)
python/triton_dist/language/extra/hip/language_extra.py:39
↓ 4 callersFunction_wave_ratio
(block_size_m, block_size_n)
python/triton_dist/kernels/amd/gemm.py:630
↓ 4 callersFunction_write_if_changed
(path: Path, content: str)
python/triton_dist/tools/compile_aot.py:135
↓ 4 callersFunctionag_gemm_intra_node
allgather gemm for intra-node return C = all_gather(A) @ B.T Args: A (torch.Tensor<float>): local matmul A matrix. shape: [M_per_ran
python/triton_dist/kernels/amd/allgather_gemm.py:1121
↓ 4 callersFunctionall_to_all_post_process
( ctx: AllToAllContext, input_splits: torch.Tensor, recv_buffer: torch.Tensor, scale_buffer: O
python/triton_dist/kernels/nvidia/low_latency_all_to_all.py:260
↓ 4 callersFunctionall_to_all_v_offset_op
(ctx: AllToAllContext, rank_in_row: bool, input: torch.Tensor = None, output: torch
python/triton_dist/kernels/nvidia/all_to_all_vdev_2d_offset.py:637
↓ 4 callersMethodapply_rotary_pos_emb
Applies Rotary Position Embedding inplace.
python/triton_dist/layers/nvidia/tp_attn.py:165
↓ 4 callersFunctionbarrier_all_on_stream
barrier_all_on_stream does not support CUDAGraph
python/triton_dist/kernels/nvidia/common_ops.py:250
↓ 4 callersFunctionbuild_kernel_on_gpu
(gpu_id, build_fn)
python/little_kernel/design/test_flashcomm_torchrun.py:189
↓ 4 callersFunctionbuild_tile_desc
(full_shape: List[int], tile_sizes: List[int], tile_id: int, return_valid_size=False)
python/triton_dist/mega_triton_kernel/tasks/utils.py:29
↓ 4 callersFunctioncdiv
(a, b)
python/triton_dist/test/nvidia/test_ep_moe_inference.py:231
↓ 4 callersFunctioncdiv
(n, m)
python/triton_dist/kernels/nvidia/ag_gemm_threadblock_swizzle.py:186
↓ 4 callersFunctioncdiv
(x, y)
python/triton_dist/kernels/nvidia/threadblock_swizzle_ag_moe.py:34
↓ 4 callersFunctioncheck_alignment
(tensors)
python/triton_dist/mega_triton_kernel/models/model_builder.py:79
↓ 4 callersFunctioncheck_allclose
Check if two tensors are close within a tolerance.
python/triton_dist/test/amd/test_tp_mlp.py:51
↓ 4 callersFunctioncheck_contiguous
(tensors)
python/triton_dist/mega_triton_kernel/models/model_builder.py:72
↓ 4 callersFunctioncheck_tensor_shape
(tensor, shape)
python/triton_dist/mega_triton_kernel/models/model_builder.py:51
↓ 4 callersMethodcombine
(self, input, ep_a2a_layout_desc: EPAllToAllLayoutDesc)
python/triton_dist/layers/amd/ep_a2a_layer.py:537
↓ 4 callersFunctionconsumer_gemm
(A, B, C, rank, num_ranks, barrier, needs_wait=True)
python/triton_dist/test/nvidia/test_distributed_wait.py:166
↓ 4 callersFunctionconsumer_gemm
(A, B, C, rank, num_ranks, barrier, needs_wait=True)
python/triton_dist/test/metax/test_distributed_wait.py:143
↓ 4 callersFunctioncreate_ag_gemm_inter_node_context
create context for allgather gemm inter-node Args: tensor_A (torch.Tensor<float>): local matmul A matrix. shape: [M_per_rank, K]
python/triton_dist/kernels/metax/allgather_gemm.py:1172
↓ 4 callersFunctioncreate_ag_gemm_intra_node_context
create context for allgather gemm intra-node Args: max_M: max number of M shape N(int): N K(int): K input_dtype(t
python/triton_dist/kernels/amd/allgather_gemm.py:909
↓ 4 callersFunctioncreate_allreduce_ctx
symmetric buffer requirement for input tensor x with x.nbytes = N. method | symmetric buffer size double_tree
python/triton_dist/kernels/nvidia/allreduce.py:109
↓ 4 callersFunctioncreate_gemm_rs_intra_node_context
create context for gemm reduce-scatter intra-node Args: max_M (int): max M N(int): N output_dtype(torch.dtype): dtype of
python/triton_dist/kernels/amd/gemm_reduce_scatter.py:391
↓ 4 callersMethoddata_ptr
(self)
python/little_kernel/design/test_flashcomm_multi_gpu.py:115
↓ 4 callersFunctiondequant_fp8_bf16
(q_tensor: torch.Tensor, scales: torch.Tensor)
python/triton_dist/test/amd/ep_a2a_utils.py:87
↓ 4 callersFunctiondequant_fp8_bf16
(q_tensor: torch.Tensor, scales: torch.Tensor)
python/triton_dist/test/nvidia/ep_a2a_utils.py:87
↓ 4 callersFunctiondist_test
Usage: def test_xxx(dist_test): dist_test(my_worker_fn, world_size=2)
python/triton_dist/test/ascend/conftest.py:57
↓ 4 callersFunctiondot_k_const
( a_ptrs, b_ptrs, c_ptrs, M, N, K: tl.constexpr, stride_ak: tl.constexpr, stri
python/triton_dist/kernels/nvidia/group_gemm.py:159
↓ 4 callersMethodep_barrier_all
(self)
python/triton_dist/layers/nvidia/ep_a2a_fused_layer.py:491
↓ 4 callersMethodexit_scope
Exit the current scope.
python/little_kernel/core/passes/utils/scope_manager.py:177
↓ 4 callersFunctionfill_tensor
(tensor: torch.Tensor, value: Any, num_sms: int = -1, eager=False)
python/triton_dist/kernels/nvidia/memory_ops.py:582
↓ 4 callersMethodforward
(self, A: torch.Tensor, # [M, local_K] weight: torch.Tensor, # [N, local_K]
python/triton_dist/test/amd/test_gemm_rs_intra_node.py:81
↓ 4 callersFunctionfused_ep_moe
Functional (inference) entry: full fused MoE forward; returns ``[num_tokens, hidden]``.
python/triton_dist/function/amd/ep_moe_fused.py:48
↓ 4 callersFunctiongemm
()
python/triton_dist/test/metax/test_ag_gemm_intra_node.py:255
↓ 4 callersFunctiongemm_rs
GEMM Reduce-Scatter for Multi-Node computes local GEMM (A x B) to generate partial results, followed by `reduce_scatter` to produce c Args:
python/triton_dist/kernels/nvidia/gemm_reduce_scatter.py:754
↓ 4 callersFunctiongemm_rs_intra_node
GEMM Reduce-Scatter for Intra-Node return C = reduce_scatter(A @ B.T) Args: A (torch.Tensor<bfloat16/float16>): local matmul A matri
python/triton_dist/kernels/amd/gemm_reduce_scatter.py:432
↓ 4 callersMethodgetNumCTAs
lib/Conversion/TritonDistributedToTritonGPU/TritonDistributedToTritonGPU.cpp:975
↓ 4 callersMethodgetNumWarps
lib/Conversion/TritonDistributedToTritonGPU/TritonDistributedToTritonGPU.cpp:973
↓ 4 callersFunctionget_all_registered_enums
Get all registered Enum classes.
python/little_kernel/codegen/registries/enum_registry.py:96
↓ 4 callersFunctionget_auto_all_gather_method
(num_ranks, num_local_ranks, pg: torch.distributed.ProcessGroup | None = None)
python/triton_dist/kernels/nvidia/allgather.py:57
↓ 4 callersMethodget_bin_op
Get C++ operator string for binary operator.
python/little_kernel/codegen/registries/operator_codegen.py:78
↓ 4 callersFunctionget_config_space
(persistent=True, device_id=0)
python/triton_dist/kernels/nvidia/gemm.py:382
↓ 4 callersFunctionget_device_max_shared_memory_size
(device_id)
python/triton_dist/utils.py:574
↓ 4 callersFunctionget_intranode_max_speed_gbps
(gpu_index=0, with_scale: bool = False)
python/triton_dist/nv_utils.py:309
↓ 4 callersFunctionget_max_gpu_clock_rate_in_khz
(device_id=0)
python/triton_dist/nv_utils.py:80
↓ 4 callersMethodget_nvshmem_size_gb
Get the total nvshmem memory size in GB.
python/triton_dist/layers/nvidia/ep_a2a_fused_layer.py:249
↓ 4 callersFunctionget_operator_codegen_registry
Get the global operator codegen registry.
python/little_kernel/codegen/registries/operator_codegen.py:104
↓ 4 callersMethodget_pybind11_cmake_args
(self)
python/setup.py:641
← previousnext →301–400 of 5,103, ranked by callers