MCPcopy Create free account

hub / github.com/ByteDance-Seed/Triton-distributed / functions

Functions5,103 in github.com/ByteDance-Seed/Triton-distributed

↓ 1 callersMethod_transform_special_struct_method_call
Transform special struct method calls. This method marks the call node for special handling in codegen. For methods
python/little_kernel/core/passes/special_struct_materialize_pass.py:632
↓ 1 callersFunction_triton_dist_impl
(input, weight, seq_lens_cpu, input_scale, weight_scale, bias)
python/triton_dist/test/nvidia/test_llm_ulysess_gemm_all2all_intra_node.py:265
↓ 1 callersFunction_triton_dist_impl
(input, weight, bias, seq_lens_cpu)
python/triton_dist/test/nvidia/test_llm_ulysess_all2all_gemm_intra_node.py:231
↓ 1 callersFunction_triton_dist_impl
(input, seq_lens_cpu)
python/triton_dist/test/nvidia/test_llm_ulysess_post_attn_all2all_intra_node.py:194
↓ 1 callersFunction_triton_dist_impl
(input, seq_lens_cpu)
python/triton_dist/test/nvidia/test_llm_ulysess_pre_attn_all2all_intra_node.py:266
↓ 1 callersFunction_triton_func
()
python/triton_dist/benchmark/bench_allgather_gemm.py:152
↓ 1 callersFunction_triton_warmup
()
python/triton_dist/test/nvidia/test_moe_utils.py:126
↓ 1 callersFunction_triton_warmup
()
python/triton_dist/test/nvidia/test_allreduce.py:180
↓ 1 callersMethod_try_resolve_from_ctx
Try to find class in ctx that has this method.
python/little_kernel/core/passes/utils/method_resolver.py:87
↓ 1 callersMethod_try_resolve_from_ctx
Try to find class in ctx that has this method.
python/little_kernel/core/passes/utils/type_inference/method_resolver.py:144
↓ 1 callersMethod_try_resolve_from_special_struct
Try to resolve method return type from special struct registry.
python/little_kernel/core/passes/utils/method_resolver.py:74
↓ 1 callersMethod_try_resolve_from_special_struct
Try to resolve method return type from special struct registry.
python/little_kernel/core/passes/utils/type_inference/method_resolver.py:117
↓ 1 callersMethod_unparse_module
Unparse a Module node, handling IR nodes separately.
python/little_kernel/core/passes/utils/enhanced_unparse.py:341
↓ 1 callersMethod_update_metrics
(self, op_type: str, io_tensors: List[List[torch.Tensor]], extra_params: Dict[str, Any] = {})
python/triton_dist/mega_triton_kernel/models/model_builder.py:135
↓ 1 callersMethod_validate
(self)
python/triton_dist/mega_triton_kernel/core/task_base.py:45
↓ 1 callersMethod_validate_context
Validate context contains `little_kernel.language` or LLType objects.
python/little_kernel/codegen/codegen_base.py:103
↓ 1 callersFunction_verify
()
python/triton_dist/test/nvidia/test_fast_allgather.py:91
↓ 1 callersFunction_verify_and_reorg_tracks
return List[(block_idx, group_idx, task_type, start_time, end_time)] sorted by start_time
python/triton_dist/tools/profiler/viewer.py:70
↓ 1 callersMethod_visit_expr
Visit an expression node using mixin's visit method.
python/little_kernel/codegen/special_struct/translator.py:209
↓ 1 callersMethod_visit_if_body
Helper method to process if statement body and else block. Args: node: ast.If node to process is_else_if: If
python/little_kernel/codegen/visitors/control_flow_codegen.py:188
↓ 1 callersMethod_write_output
Write JSON to disk with proper structure
python/triton_dist/profiler_utils.py:151
↓ 1 callersFunctionact_mul_up_tile_compute
(tile_id, input, output, M, N, ACT_FN, BLOCK_SIZE_M: tl.constexpr, BLOCK_SIZE_N: t
python/triton_dist/mega_triton_kernel/kernels/activation.py:31
↓ 1 callersFunctionadd_continuous
( lhs: torch.Tensor, rhs: torch.Tensor, out: Optional[torch.Tensor], num_ctas=16, num_warp
python/triton_dist/kernels/nvidia/reduce_scatter.py:206
↓ 1 callersMethodadd_input_producer
(self, input_index, src_node, src_out_idx)
python/triton_dist/mega_triton_kernel/core/graph.py:68
↓ 1 callersFunctionadd_license
Add license header to a file that is missing it.
scripts/check_license.py:217
↓ 1 callersFunctionadd_link_to_backends
(external_only, materialization=False)
python/setup.py:892
↓ 1 callersFunctionadd_link_to_distributed
()
python/setup.py:952
↓ 1 callersFunctionadd_link_to_proton
()
python/setup.py:943
↓ 1 callersFunctionadd_link_to_pymxshmem
()
python/setup.py:961
↓ 1 callersMethodaddr_of
(self)
python/little_kernel/core/expr.py:220
↓ 1 callersFunctionag_gemm_inter_node
allgather gemm for inter-node Allgather global matrix A and do matmul with local matrix B, produces local matrix C Args: a (torch.Te
python/triton_dist/kernels/metax/allgather_gemm.py:1232
↓ 1 callersFunctionag_gemm_inter_node_op
allgather gemm for inter-node Allgather global matrix A and do matmul with local matrix B, produces local matrix C Args: a (torch.Te
python/triton_dist/kernels/metax/allgather_gemm.py:905
↓ 1 callersFunctionag_gemm_intra_node_op
(A: torch.Tensor, B: torch.Tensor, C: torch.Tensor, ctx: AllGatherGEMMTensorParallelContext,
python/triton_dist/kernels/amd/allgather_gemm.py:1039
↓ 1 callersFunctionag_gemm_intra_node_op
no-tma allgather gemm for intra-node Allgather global matrix A and do matmul with local matrix B, produces local matrix C Args: a (t
python/triton_dist/kernels/metax/allgather_gemm.py:801
↓ 1 callersFunctionag_gemm_persistent_op
(a, b, c, rank,
tutorials/07-overlapping-allgather-gemm.py:367
↓ 1 callersFunctionag_intra_node_nvlink_small_msg
(symm_data, comm_buf, rank, num_ranks, need_block_level_sync=False)
python/triton_dist/test/nvidia/test_ag_small_msg.py:121
↓ 1 callersFunctionalign_k
(k)
python/little_kernel/design/sm90_bf16_gemm.py:273
↓ 1 callersFunctionalign_n
(n)
python/little_kernel/design/sm90_bf16_gemm.py:277
↓ 1 callersFunctionall_to_all_copy_engine
Args: input: Input tensor of shape (M, K) input_scale: Optional input scales (M,) comm_data_buffers: List of communicati
python/triton_dist/kernels/nvidia/all_to_all_single_gemm.py:39
↓ 1 callersFunctionall_to_all_single_2d
( ctx: AllToAllSingle2DContext, input_tensor: torch.Tensor, output_tensor: torch.Tensor, input
python/triton_dist/kernels/nvidia/all_to_all_single_2d.py:161
↓ 1 callersFunctionallgather
(A: torch.Tensor, ctx)
python/triton_dist/kernels/amd/allgather_gemm.py:1169
↓ 1 callersFunctionallgather
( full_symm: torch.Tensor, shard: torch.Tensor, barrier: torch.Tensor, workgroups_per_rank: in
python/triton_dist/kernels/amd/allgather.py:638
↓ 1 callersFunctionallgather_chunked
( full_symm: torch.Tensor, shard: torch.Tensor, group_barrier: torch.Tensor, grid_barrier: tor
python/triton_dist/kernels/amd/allgather.py:752
↓ 1 callersFunctionallgather_chunked_pull
( full_symm: torch.Tensor, shard: torch.Tensor, group_barrier: torch.Tensor, grid_barrier: tor
python/triton_dist/kernels/amd/allgather.py:657
↓ 1 callersFunctionallgather_chunked_pull_fused
( full_symm: torch.Tensor, shard: torch.Tensor, group_barrier: torch.Tensor, grid_barrier: tor
python/triton_dist/kernels/amd/allgather.py:693
↓ 1 callersFunctionallgather_chunked_pull_packed_fused
( full_symm: torch.Tensor, shard: torch.Tensor, group_barrier: torch.Tensor, grid_barrier: tor
python/triton_dist/kernels/amd/allgather.py:721
↓ 1 callersFunctionallgather_gemm
consume gemm
tutorials/ascend/01-ascend-allgather-gemm.py:174
↓ 1 callersFunctionallgather_strided_chunked_pull_packed_kernel
with the local_ptr
python/triton_dist/kernels/amd/allgather.py:316
↓ 1 callersFunctionallreduce_one_shot_multimem_intra_node_kernel
(pid, num_pid, symm_in_ptr, out_ptr, elems)
python/triton_dist/mega_triton_kernel/kernels/allreduce.py:35
↓ 1 callersFunctionallreduce_op
(ctx: GemmARContext, c, gemm_config: triton.Config, TILE_MAP_LEVEL=0, copy_to_local=True, USE
python/triton_dist/kernels/nvidia/gemm_allreduce.py:812
↓ 1 callersFunctionanalyze_testcases
()
python/triton_dist/test/nvidia/test_all_to_all_vdev_2d_offset.py:444
↓ 1 callersFunctionappend_result_immediately
Write single result immediately to CSV with forced flushing
python/triton_dist/benchmark/bench_tp_mlp.py:77
↓ 1 callersFunctionapply_patch_inductor_triton_compat
()
python/triton_dist/tools/monkey_inductor.py:338
↓ 1 callersFunctionapply_patch_inductor_triton_heuristics
()
python/triton_dist/tools/monkey_inductor.py:331
↓ 1 callersFunctionassert_bitwise_equal
(x: torch.Tensor, y: torch.Tensor, verbose=True)
python/triton_dist/test/utils.py:100
↓ 1 callersFunctionatomic_add
semantic should be one of ["monotonic", "release", "acquire", "acq_rel] scope should be one of ["workgroup", "agent", "system]
python/triton_dist/language/extra/hip/language_extra.py:162
↓ 1 callersFunctionatomic_add
custom atomic_add implementation using extern_elementwise :param scope: one of "gpu", "sys". default to "gpu" :param semantic: one of "releas
python/triton_dist/language/extra/cuda/language_extra.py:753
↓ 1 callersFunctionatomic_uint64_nonfetch
Atomic non-fetch uint64 operation (thread scope). Args: dest: Symmetric address on target PE val: Value to be used in atomic
python/triton_dist/language/extra/hip/libmori_shmem_device.py:942
↓ 1 callersFunctionbarrier_all_block
( barrier_ptrs: ll.ptr[ll.ptr[ll.int32]], rank: ll.int32, num_ranks: ll.int32, )
python/little_kernel/design/flashcomm_barrier.py:43
↓ 1 callersFunctionbarrier_all_block
(barrier_ptrs: ll.ptr[ll.ptr[ll.int32]], rank: ll.int32, num_ranks: ll.int32)
python/little_kernel/design/flashcomm_compute.py:100
↓ 1 callersFunctionbarrier_all_intra_node_atomic_cas_block
NOTE: this function should only be called with atomic support. memory over PCI-e does not support atomic r/w. DON'T use this function on such platfor
python/triton_dist/mega_triton_kernel/kernels/barrier.py:33
↓ 1 callersFunctionbarrier_all_intra_node_non_atomic_block
symm_flags is expected to: 1. of int32 dtype 2. has at least num_ranks * 2 elements 3. of symmetric pointer
python/triton_dist/kernels/nvidia/common_ops.py:184
↓ 1 callersFunctionbarrier_all_ipc_kernel
(rank, num_ranks, comm_buf_base_ptrs)
python/triton_dist/kernels/amd/common_ops.py:138
↓ 1 callersFunctionbarrier_all_ipc_kernel_v2
(rank, num_ranks, comm_buf_base_ptrs)
python/triton_dist/kernels/amd/common_ops.py:123
↓ 1 callersFunctionbench
(kernel, M, N, K, iters=20)
python/little_kernel/benchmark/gemm_sm100/gemm_level1.py:226
↓ 1 callersFunctionbench
(kernel, M, N, K, iters=20)
python/little_kernel/benchmark/gemm_sm100/gemm_level8.py:363
↓ 1 callersFunctionbench
(kernel, M, N, K, iters=20)
python/little_kernel/benchmark/gemm_sm100/gemm_level2.py:232
↓ 1 callersFunctionbench
(kernel, M, N, K, iters=20)
python/little_kernel/benchmark/gemm_sm100/gemm_level7.py:335
↓ 1 callersFunctionbench
(kernel, M, N, K, iters=20)
python/little_kernel/benchmark/gemm_sm100/gemm_level3.py:238
↓ 1 callersFunctionbench
(kernel, M, N, K, iters=20)
python/little_kernel/benchmark/gemm_sm100/gemm_level9.py:409
↓ 1 callersFunctionbench
(kernel, M, N, K, iters=20)
python/little_kernel/benchmark/gemm_sm100/gemm_level4.py:224
↓ 1 callersFunctionbench
(kernel, M, N, K, iters=20)
python/little_kernel/benchmark/gemm_sm100/gemm_level5.py:255
↓ 1 callersFunctionbench
(kernel, M, N, K, iters=20)
python/little_kernel/benchmark/gemm_sm100/gemm_level6.py:345
↓ 1 callersFunctionbench
(kernel, M, N, K, iters=20, label="")
python/little_kernel/benchmark/gemm_sm90/gemm_v2.py:154
↓ 1 callersFunctionbench
(kernel, M, N, K, iters=20)
python/little_kernel/benchmark/gemm_sm90/gemm_v8.py:227
↓ 1 callersFunctionbench
(kernel, M, N, K, iters=20)
python/little_kernel/benchmark/gemm_sm90/gemm_v9.py:255
↓ 1 callersFunctionbench
(kernel, M, N, K, iters=20)
python/little_kernel/benchmark/gemm_sm90/gemm_v6.py:206
↓ 1 callersFunctionbench
(kernel, M, N, K, iters=20)
python/little_kernel/benchmark/gemm_sm90/gemm_v7.py:213
↓ 1 callersFunctionbench
(kernel, M, N, K, iters=20)
python/little_kernel/benchmark/gemm_sm90/gemm_v4.py:164
↓ 1 callersFunctionbench
(kernel, M, N, K, iters=20)
python/little_kernel/benchmark/gemm_sm90/gemm_v5.py:185
↓ 1 callersFunctionbench
(kernel, M, N, K, iters=20, label="")
python/little_kernel/benchmark/gemm_sm90/gemm_v3.py:193
↓ 1 callersFunctionbench
(kernel, M, N, K, iters=20)
python/little_kernel/benchmark/gemm_sm90/gemm_v10.py:263
↓ 1 callersFunctionbench
Benchmark performance.
python/little_kernel/benchmark/gemm_sm90/gemm_v1.py:198
↓ 1 callersFunctionbench_flash_attn
(builder, BATCH, N_CTX, Q_HEAD, KV_HEAD, HEAD_DIM, IS_CAUSAL, dtype=torch.bfloat16, qkv_pack=False)
python/triton_dist/mega_triton_kernel/test/ops/test_flash_attn.py:76
↓ 1 callersFunctionbenchmark
Benchmark the implementation
python/triton_dist/test/nvidia/test_all_to_all_single_gemm.py:396
↓ 1 callersFunctionbenchmark
(SP_GROUP, args)
python/triton_dist/test/nvidia/test_ulysses_sp_infer_qkv_proj_a2a.py:182
↓ 1 callersFunctionbisect_right_with_offset_kernel
index = bisect(sorted_ptr, values, N) off = sorted_ptr[index - 1] : let suppose values[-1] = 0 remainder = values - off It's expecte
python/triton_dist/kernels/nvidia/threadblock_swizzle_ag_moe_triton.py:81
↓ 1 callersFunctionbisect_triton
(sorted_tensor, values_tensor, side="left", aligned=False)
python/triton_dist/test/nvidia/test_common_ops.py:114
↓ 1 callersFunctionblockIdx_x
Block index in the x dimension.
python/little_kernel/language/intrin/simt.py:66
↓ 1 callersFunctionblock_scan_inclusive
(value: ll.int32, warp_sums: ll.ptr[ll.int32])
python/little_kernel/design/flashcomm_compute.py:81
↓ 1 callersFunctionbroadcast_cpu
(tensor: torch.Tensor, src: int, group: torch.distributed.ProcessGroup)
shmem/mxshmem_bind/pymxshmem/python/pymxshmem/__init__.py:17
↓ 1 callersFunctionbroadcast_naive_block
(dst_ptr, src_ptr, nbytes)
python/triton_dist/kernels/nvidia/low_latency_allgather.py:609
↓ 1 callersFunctionbuild_compute_offset_kernel
Build the compute offset kernel using LK's build() API. shared_mem_bytes = NUM_WARPS * (num_experts + 1) * sizeof(int32) The kernel uses
python/little_kernel/design/test_flashcomm_compute.py:88
↓ 1 callersFunctionbuild_dispatch_layout_kernel
Build the dispatch layout kernel.
python/little_kernel/design/test_flashcomm_compute.py:244
↓ 1 callersMethodbuild_extension_cmake
(self, ext)
python/setup.py:664
↓ 1 callersFunctionbuild_kernel
Build a launchable kernel using LittleKernel's Python binding. Parameters ---------- grid : tuple Grid dimensions (num_blocks, 1,
python/little_kernel/design/sm90_bf16_gemm.py:245
↓ 1 callersFunctionbuild_kernel
(kernel_name="basic")
python/little_kernel/design/flashcomm_dispatch.py:390
↓ 1 callersFunctionbuild_kernel
Build the barrier kernel.
python/little_kernel/design/flashcomm_barrier.py:85
↓ 1 callersFunctionbuild_kernel
()
python/little_kernel/design/flashcomm_postprocess.py:183
↓ 1 callersFunctionbuild_kernel
(kernel_name="offset")
python/little_kernel/design/flashcomm_compute.py:343
← previousnext →1,201–1,300 of 5,103, ranked by callers