MCPcopy Create free account

hub / github.com/ByteDance-Seed/Triton-distributed / functions

Functions5,103 in github.com/ByteDance-Seed/Triton-distributed

↓ 4 callersMethodget_sm_activity
(self)
python/triton_dist/mega_triton_kernel/models/model_builder.py:164
↓ 4 callersMethodinfer_constant_type
Infer LLType for ast.Constant nodes (handles primitives and LLType constants)
python/little_kernel/core/passes/utils/type_inference/type_inference_visitors.py:40
↓ 4 callersMethodinit_triton_dist_AR_ctx
(self, max_M: int = 128, ar_method: AllReduceMethod = AllReduceMethod.DoubleTree)
python/triton_dist/models/dense.py:190
↓ 4 callersFunctionis_ascend
Checks if 'npu-smi' is available on the system's PATH.
python/triton_dist/utils.py:89
↓ 4 callersFunctionis_cuda
()
python/triton_dist/test/amd/test_matmul_amd.py:133
↓ 4 callersFunctionis_offline_build
Downstream projects and distributions which bootstrap their own dependencies from scratch and run builds in offline sandboxes may set `TR
python/setup.py:243
↓ 4 callersMethodis_struct
Check if the type is a struct
python/little_kernel/core/type_system.py:59
↓ 4 callersMethodlltype_to_cpp
Convert an LLType instance to a valid C++ type string.
python/little_kernel/codegen/registries/type_converter.py:49
↓ 4 callersFunctionload_autotune_data
(filename: str | Path)
python/triton_dist/tune.py:175
↓ 4 callersMethodmake_add
(self, lhs: torch.Tensor, rhs: torch.Tensor, output: torch.Tensor, layer_id=0)
python/triton_dist/mega_triton_kernel/models/model_builder.py:451
↓ 4 callersFunctionmake_cuda_graph
(mempool, func)
python/triton_dist/test/amd/test_tp_attn.py:64
↓ 4 callersFunctionmake_cuda_graph
(mempool, func)
python/triton_dist/test/amd/test_tp_e2e.py:65
↓ 4 callersFunctionmake_kernel_algo_info_struct_name
(c_kernel_name)
python/triton_dist/tools/compile_aot.py:184
↓ 4 callersMethodmega_dispatch_group_gemm
Fused dispatch + grouped GEMM-1. Returns ``(gemm1_out[M, inter], desc)``.
python/triton_dist/layers/amd/ep_a2a_fused_layer.py:189
↓ 4 callersMethodmega_forwrad
(self, input_ids: torch.LongTensor)
python/triton_dist/mega_triton_kernel/models/dense.py:191
↓ 4 callersMethodmega_group_gemm_combine
Fused grouped GEMM-2 + combine (serial/gather). Returns ``[num_tokens, hidden]``.
python/triton_dist/layers/amd/ep_a2a_fused_layer.py:213
↓ 4 callersFunctionmembar
(scope: tl.constexpr)
python/triton_dist/kernels/amd/low_latency_all_to_all_v2.py:62
↓ 4 callersFunctionmultimem_st_b32
(ptr, val0, _semantic=None)
python/triton_dist/language/extra/cuda/language_extra.py:283
↓ 4 callersFunctionperf_func
Run function with warmup, return last result and avg time in ms.
python/triton_dist/test/amd/test_ep_a2a.py:191
↓ 4 callersFunctionprepare_chunk_indices
(cu_seqlens: torch.IntTensor, chunk_size: int)
python/triton_dist/kernels/nvidia/gdn.py:56
↓ 4 callersFunctionref_paged_attn
( query: torch.Tensor, key_cache: torch.Tensor, value_cache: torch.Tensor, query_lens: List[in
python/triton_dist/test/nvidia/test_decode_attn.py:68
↓ 4 callersMethodregister_un_op
Register a unary operator handler.
python/little_kernel/core/passes/utils/registries/operator_registry.py:63
↓ 4 callersFunctionreshape_2d
(arr_2d: List[List[Tile]], )
python/triton_dist/kernels/nvidia/threadblock_swizzle_ag_moe.py:110
↓ 4 callersMethodresolve_method_return_type
Resolve method return type using multiple strategies. Args: var_name: Name of the variable (e.g., 'scheduler')
python/little_kernel/core/passes/utils/type_inference/method_resolver.py:153
↓ 4 callersFunctionrow_to_chunk_idx
(row, M_PER_CHUNK, M_per_rank, chunks_per_rank)
python/triton_dist/kernels/amd/allgather_gemm.py:478
↓ 4 callersFunctionrun_moe_ag_triton_non_overlap
(x_shard: torch.Tensor, weights: torch.Tensor, chosen_experts: torch.Tensor,
python/triton_dist/kernels/nvidia/allgather_group_gemm.py:901
↓ 4 callersFunctionrun_multimem_ld_reduce
(symm_tensor: torch.Tensor, acc_dtype: torch.dtype, num_grids=4, num_warps=32)
python/triton_dist/test/nvidia/test_multimem_ld_reduce.py:67
↓ 4 callersMethodscan_for_locals
Scans a list of statements for variable definitions and adds them to current scope. This follows Python's semantics: any ass
python/little_kernel/core/passes/utils/scope_manager.py:238
↓ 4 callersFunctionseed_everything
Seed everything for better reproducibility. (some pytorch operation is non-deterministic like the backprop of grid_samples)
python/triton_dist/models/utils.py:75
↓ 4 callersMethodsetUp
(self)
python/little_kernel/tests/unit/test_stmt.py:109
↓ 4 callersMethodset_signal
(self, pp_rank, buffer_id, value, stream=None, num_barriers=1)
python/triton_dist/layers/nvidia/p2p.py:137
↓ 4 callersFunctionshard_local
(tensor: torch.Tensor, world_size: int, dim: int, local_rank: int)
python/triton_dist/mega_triton_kernel/models/layers/tp_attn.py:30
↓ 4 callersFunctionshard_local
(tensor: torch.Tensor, world_size: int, dim: int, local_rank: int)
python/triton_dist/layers/amd/tp_attn.py:38
↓ 4 callersFunctionsort_by_vectors
Sort 2D tensor rows lexicographically (order-agnostic comparison).
python/triton_dist/test/amd/test_ep_a2a.py:156
↓ 4 callersFunctionswizzle_2d
(tile_id, num_pid_m, num_pid_n, GROUP_SIZE_M: tl.constexpr)
python/triton_dist/kernels/nvidia/allgather_group_gemm.py:525
↓ 4 callersFunctionswizzle_2d_by_group_n
if we choose tile first in N within group_size_N, maybe each group with N = 1024, for BLOCK_SIZE_N = 64, then 16 tiles per tiled_m. maybe too muc
python/triton_dist/kernels/nvidia/moe_reduce_ar.py:47
↓ 4 callersFunctionswizzle_2d_by_group_n
if we choose tile first in N within group_size_N, maybe each group with N = 1024, for BLOCK_SIZE_N = 64, then 16 tiles per tiled_m. maybe too muc
python/triton_dist/kernels/nvidia/moe_reduce_rs.py:124
↓ 4 callersFunctionsync_warp
(_semantic=None)
python/triton_dist/kernels/nvidia/ep_all2all_fused.py:41
↓ 4 callersFunctiontile_kernel_topk_reduce_token_intra_node
( pid, num_pid, num_input_tokens_per_rank, # [world_size] scatter_send_buf, #[max_tokens, to
python/triton_dist/kernels/nvidia/ep_all2all_fused.py:485
↓ 4 callersFunctiontorch_a2a
Reference PyTorch all-to-all implementation
python/triton_dist/test/nvidia/test_all_to_all_single_gemm.py:149
↓ 4 callersFunctiontorch_func
()
python/triton_dist/test/metax/test_ag_gemm_inter_node.py:86
↓ 4 callersFunctiontranspose2d
python/triton_dist/kernels/nvidia/threadblock_swizzle_ag_moe.cc:153
↓ 4 callersFunctionwarp_scan_inclusive
(value: ll.int32)
python/little_kernel/design/flashcomm_compute.py:48
↓ 4 callersMethodwrite_main
Write to code buffer without newline (for mixin compatibility).
python/little_kernel/codegen/special_struct/translator.py:92
↓ 3 callersFunctionCUDA_CHECK
(err)
python/triton_dist/test/nvidia/test_sp_decode_attn.py:69
↓ 3 callersFunction__ballot_sync
( mask, predicate, _semantic=None, )
python/triton_dist/language/extra/cuda/language_extra.py:863
↓ 3 callersFunction__tid__
(axis: core.constexpr, _semantic=None)
python/triton_dist/language/extra/maca/language_extra.py:36
↓ 3 callersFunction_a2a
(a2a_tensor)
python/triton_dist/test/nvidia/test_llm_ulysess_gemm_all2all_intra_node.py:158
↓ 3 callersFunction_a2a
(a2a_tensor)
python/triton_dist/test/nvidia/test_llm_ulysess_pre_attn_all2all_intra_node.py:188
↓ 3 callersFunction_a2a
(a2a_tensor)
python/triton_dist/test/nvidia/test_ulysses_sp_dispatch.py:64
↓ 3 callersFunction_barrier_all_impl
(SCOPE_SUFFIX: core.constexpr, _semantic=None)
python/triton_dist/language/extra/hip/librocshmem_device.py:404
↓ 3 callersFunction_barrier_all_impl
(SCOPE_SUFFIX: core.constexpr, _semantic=None)
python/triton_dist/language/extra/maca/libmxshmem_device.py:217
↓ 3 callersFunction_barrier_all_impl
(SCOPE_SUFFIX: core.constexpr, _semantic=None)
python/triton_dist/language/extra/cuda/libnvshmem_device.py:258
↓ 3 callersFunction_barrier_impl
(team, SCOPE_SUFFIX: core.constexpr, _semantic=None)
python/triton_dist/language/extra/cuda/libnvshmem_device.py:229
↓ 3 callersFunction_build_dispatch_metadata
Compute, in torch, everything the fused kernel needs that is cheaper / safer to derive on the host: the cross-rank split exchange, the destination
python/triton_dist/kernels/amd/ep_all2all_fused.py:558
↓ 3 callersFunction_check_buf
(src, dst)
python/triton_dist/kernels/nvidia/ulysses_sp_dispatch.py:417
↓ 3 callersFunction_compute_pid
(tile_id, num_pid_in_group, num_pid_m, GROUP_SIZE_M, NUM_GEMM_SMS)
python/triton_dist/kernels/nvidia/sp_ulysess_qkv_gemm_all2all.py:54
↓ 3 callersFunction_ensure_amdsmi_initialized
()
python/triton_dist/amd_utils.py:54
↓ 3 callersMethod_format_attr_value
Format an attribute value for display.
python/little_kernel/core/passes/utils/enhanced_unparse.py:155
↓ 3 callersFunction_get_bus_bw_gbps_between
(device_id_i: int, device_id_j: int)
python/triton_dist/amd_utils.py:393
↓ 3 callersMethod_get_method_return_type_from_class
Get method return type from class definition by parsing its AST.
python/little_kernel/core/passes/utils/method_resolver.py:41
↓ 3 callersMethod_get_method_return_type_from_class
Get method return type from class definition by parsing its AST.
python/little_kernel/core/passes/utils/type_inference/method_resolver.py:41
↓ 3 callersMethod_get_next_k_group
Get next valid K group.
python/little_kernel/atom/scheduler.py:150
↓ 3 callersMethod_inline_call
Core inlining logic: - is_statement: True if call is standalone (e.g., 'func()'), False if assigned (e.g., 'x = func()') - lh
python/little_kernel/core/passes/inline.py:151
↓ 3 callersFunction_inplace_copy_from_comm_buf_to_out
(src, dst)
python/triton_dist/kernels/nvidia/ulysses_sp_dispatch.py:638
↓ 3 callersFunction_is_all_ranks_bitwise_match
(t: torch.Tensor)
python/triton_dist/test/nvidia/test_multimem_ld_reduce.py:109
↓ 3 callersFunction_is_hip_platform
Checks if 'rocm-smi' is available on the system's PATH.
python/setup.py:67
↓ 3 callersMethod_is_special_struct_constructor
Check if function is a special struct constructor. Returns the struct name if it's a registered special struct constructor,
python/little_kernel/core/passes/special_struct_materialize_pass.py:105
↓ 3 callersFunction_kernel_inner_tile_copy
( src_base_ptr, dst_base_ptr, seq, tile_id_seq, stride_src_seq, stride_src_hd, str
python/triton_dist/kernels/nvidia/ulysses_sp_dispatch.py:280
↓ 3 callersFunction_lltype_to_cpp_simple
Simple LLType to C++ type conversion for simple_builtin. This is a simplified version that doesn't require a full emitter context.
python/little_kernel/language/simple_builtin.py:238
↓ 3 callersFunction_load_v4_impl
(ptr, suffix: core.constexpr, scope: core.constexpr = "", semantic: core.constexpr = "", _se
python/triton_dist/language/extra/cuda/language_extra.py:144
↓ 3 callersFunction_multimem_ld_reduce_128bit
(ptr, acc_dtype: core.constexpr, suffix: core.constexpr, _semantic=None)
python/triton_dist/language/extra/cuda/language_extra.py:413
↓ 3 callersFunction_multimem_ld_reduce_p_128bit
(ptr, mask, suffix: core.constexpr, _semantic=None)
python/triton_dist/language/extra/cuda/language_extra.py:376
↓ 3 callersFunction_multimem_st_v2_impl
(ptr, val0, val1, suffix: core.constexpr, _semantic=None)
python/triton_dist/language/extra/cuda/language_extra.py:288
↓ 3 callersFunction_multimem_st_v4_impl
(ptr, val0, val1, val2, val3, suffix: core.constexpr, _semantic=None)
python/triton_dist/language/extra/cuda/language_extra.py:320
↓ 3 callersFunction_ntid_wrapper
(axis: core.constexpr, _semantic=None)
python/triton_dist/language/extra/hip/language_extra.py:76
↓ 3 callersFunction_ntid_wrapper
(axis: core.constexpr, _semantic=None)
python/triton_dist/language/extra/cuda/language_extra.py:508
↓ 3 callersFunction_random_sleep
()
python/triton_dist/test/nvidia/test_common_ops.py:47
↓ 3 callersMethod_resolve_callable
Resolve a callable from AST node.
python/little_kernel/core/passes/special_struct_materialize_pass.py:165
↓ 3 callersFunction_slice_and_reshape
(src, dst)
python/triton_dist/kernels/nvidia/ulysses_sp_dispatch.py:692
↓ 3 callersFunction_sync_all_impl
(SCOPE_SUFFIX: core.constexpr, _semantic=None)
python/triton_dist/language/extra/cuda/libnvshmem_device.py:287
↓ 3 callersFunction_tid_wrapper
(axis: core.constexpr, _semantic=None)
python/triton_dist/language/extra/hip/language_extra.py:50
↓ 3 callersFunction_tid_wrapper
(axis: core.constexpr, _semantic=None)
python/triton_dist/language/extra/cuda/language_extra.py:462
↓ 3 callersFunctionall_to_all_v_offset_op_v2
(ctx: AllToAllContext, rank_in_row: bool, input: torch.Tensor, output: torch.Tensor,
python/triton_dist/kernels/nvidia/all_to_all_vdev_2d_offset.py:575
↓ 3 callersFunctionapply_rotary_pos_emb
Applies Rotary Position Embedding inplace.
python/triton_dist/mega_triton_kernel/test/torch_impl_utils.py:45
↓ 3 callersMethodbackward
(ctx, *grad_outputs)
python/triton_dist/function/amd/ep_moe_fused.py:81
↓ 3 callersFunctionbarrier_async
(pg: torch.distributed.ProcessGroup)
python/triton_dist/utils.py:942
↓ 3 callersFunctionbenchmark_latency_memory
(func, iters=100, warmup_iters=10, pre_func=None)
python/triton_dist/profiler_utils.py:372
↓ 3 callersFunctionbisect_right_kernel
(sorted_values_ptr, # Pointer to sorted input array (1D) target_values, # Pointer to
python/triton_dist/kernels/nvidia/common_ops.py:323
↓ 3 callersMethodbuild_fwd
(self, hidden_states: torch.Tensor, kv_cache: PagedKVCache)
python/triton_dist/mega_triton_kernel/models/dense.py:167
↓ 3 callersFunctionbuild_rocshmem
()
python/setup.py:490
↓ 3 callersFunctioncalc_gather_index
( scatter_index: torch.Tensor, row_start: int, row_end: int, BLOCK_SIZE: int = 512, )
python/triton_dist/test/amd/test_all_to_all.py:47
↓ 3 callersFunctioncalc_gather_index
( scatter_index: torch.Tensor, row_start: int, row_end: int, BLOCK_SIZE: int = 1024, )
python/triton_dist/test/nvidia/test_all_to_all.py:47
↓ 3 callersFunctioncalc_gather_scatter_index_triton
( chosen_experts: torch.Tensor, nexperts: int, alignment_by_expert: int = 1, )
python/triton_dist/kernels/nvidia/moe_utils.py:308
↓ 3 callersFunctioncalc_gather_scatter_index_v2_triton
( chosen_experts: torch.Tensor, nexperts: int, alignment_by_expert: int = 1, )
python/triton_dist/kernels/nvidia/moe_utils.py:341
↓ 3 callersFunctioncalc_scatter_index_stable
(chosen_experts: torch.Tensor)
python/triton_dist/test/amd/test_all_to_all.py:95
↓ 3 callersFunctioncalc_scatter_index_stable
(chosen_experts: torch.Tensor)
python/triton_dist/test/nvidia/test_all_to_all.py:93
↓ 3 callersFunctioncalc_scatter_index_stable
(chosen_experts: torch.Tensor)
python/triton_dist/test/nvidia/test_ep_a2a.py:138
↓ 3 callersFunctioncdiv
python/triton_dist/kernels/nvidia/threadblock_swizzle_ag_moe.cc:97
↓ 3 callersFunctioncheck_allclose
Check if two tensors are close within a tolerance.
python/triton_dist/test/amd/test_tp_attn.py:50
← previousnext →401–500 of 5,103, ranked by callers