Code
Hub
Workspaces
Following
Trending
Connect
MCP
copy
Create free account
hub
/
github.com/ByteDance-Seed/Triton-distributed
/ functions
Functions
5,103 in github.com/ByteDance-Seed/Triton-distributed
⨍
Functions
5,103
◇
Types & classes
521
↳
Endpoints
149
↓ 1 callers
Function
_build_wgmma_fn_body
()
python/little_kernel/language/intrin/wgmma.py:126
↓ 1 callers
Function
_check
(out: torch.Tensor, ref: torch.Tensor, msg: str = "Triton")
python/triton_dist/test/nvidia/test_ep_a2a.py:48
↓ 1 callers
Function
_check_autotune_version
()
python/triton_dist/tune.py:98
↓ 1 callers
Method
_check_error
Check CUDA error and raise exception if needed.
python/little_kernel/runtime/tma_descriptor.py:125
↓ 1 callers
Method
_check_is_const
Check if annotation is const[...]
python/little_kernel/core/passes/utils/type_inference/type_inference_visitors.py:170
↓ 1 callers
Method
_check_is_template
Check if annotation is template[...]
python/little_kernel/core/passes/utils/type_inference/type_inference_visitors.py:149
↓ 1 callers
Method
_check_is_tuple
Check if annotation is Tuple[...] or tuple[...]
python/little_kernel/core/passes/utils/type_inference/type_inference_visitors.py:191
↓ 1 callers
Function
_check_signature_or_throw
check all scalar arguments or constexpr arguments. no check pointer dtypes argument signature should be: [*]dtype[:annotation] *
python/triton_dist/tools/compile_aot.py:147
↓ 1 callers
Function
_check_swizzled
(swizzled)
python/triton_dist/kernels/nvidia/gemm_rs_threadblock_swizzle.py:259
↓ 1 callers
Function
_check_threadblock_swizzle_with
(M)
python/triton_dist/kernels/nvidia/ag_gemm_threadblock_swizzle.py:323
↓ 1 callers
Function
_check_threadblock_swizzle_with
(M)
python/triton_dist/kernels/nvidia/gemm_rs_threadblock_swizzle.py:266
↓ 1 callers
Function
_check_tiled_m
(tiled_m_global)
python/triton_dist/kernels/nvidia/threadblock_swizzle_ag_moe_triton.py:346
↓ 1 callers
Method
_choose_track
(self, ts_start, ts_end)
python/triton_dist/tools/profiler/viewer.py:105
↓ 1 callers
Method
_chunkify_list
Split list into chunks for parallel processing
python/triton_dist/profiler_utils.py:142
↓ 1 callers
Method
_collect_all_to_rank0
(self)
python/triton_dist/profiler_utils.py:246
↓ 1 callers
Method
_collect_local_vars
Collect local variables defined in the target function
python/little_kernel/core/passes/inline.py:281
↓ 1 callers
Method
_collect_variables
Collect all variable names in the calling context
python/little_kernel/core/passes/inline.py:268
↓ 1 callers
Function
_compare_deps
(deps)
python/triton_dist/tune.py:246
↓ 1 callers
Method
_compile
(self)
python/little_kernel/runtime/kernel.py:77
↓ 1 callers
Function
_compile_kernel
( workspace: Path, signature: str, func: triton.JITFunction, out_name: str, num_warps: int
python/triton_dist/tools/compile_aot.py:205
↓ 1 callers
Function
_compute_pid
(tile_id, num_pid_in_group, num_pid_m, GROUP_SIZE_M, NUM_SMS)
python/triton_dist/test/nvidia/test_gemm_ar.py:107
↓ 1 callers
Method
_compute_step
The pure computation logic for one micro-batch.
python/triton_dist/test/nvidia/test_pp_block.py:141
↓ 1 callers
Function
_contextual_tuning_run
(self: Autotuner, *args, **kwargs)
python/triton_dist/autotuner.py:155
↓ 1 callers
Method
_create_tensor_dtype_ast
Create AST node for ll.Tensor[elem_dtype]
python/little_kernel/core/passes/flatten_empty.py:77
↓ 1 callers
Function
_deps_list_to_dependency
(deps_list, layer_id, task_id)
python/triton_dist/mega_triton_kernel/core/graph.py:51
↓ 1 callers
Function
_do_bench_iterator
(funcs, n_repeat=5, n_warmup=3, quantiles=None, return_mode="mean")
python/triton_dist/autotuner.py:105
↓ 1 callers
Function
_enhanced_dump_recursive
Recursively dump a node with custom attributes.
python/little_kernel/core/passes/utils/enhanced_dump.py:212
↓ 1 callers
Method
_evaluate_expr
Evaluate an AST expression using the function's global context
python/little_kernel/core/passes/flatten_empty.py:66
↓ 1 callers
Function
_extract_asm_info
Extract assembly template, constraints, and operands from an asm() call. Args: asm_call: AST node representing asm() call
python/little_kernel/language/simple_builtin.py:115
↓ 1 callers
Method
_extract_member_vars_from_init
Extract member variables and template parameters from __init__ method.
python/little_kernel/codegen/special_struct/struct_converter.py:135
↓ 1 callers
Method
_extract_struct_args
Extract struct arguments from constructor call. This is a generic implementation that extracts all arguments (positional and
python/little_kernel/core/passes/special_struct_materialize_pass.py:364
↓ 1 callers
Function
_find_asm_call
Find the first asm() call in an AST node. Supports: - asm(...) - ll.asm(...) - ll.asm_volatile(...)
python/little_kernel/language/simple_builtin.py:94
↓ 1 callers
Function
_find_return_variable
Find the return variable name from function body.
python/little_kernel/language/simple_builtin.py:153
↓ 1 callers
Function
_format_attr_value
Format an attribute value for display.
python/little_kernel/core/passes/dump.py:52
↓ 1 callers
Function
_gen_inputs
(max_local_seq_len, iter=0, is_debug=False)
python/triton_dist/test/nvidia/test_llm_ulysess_gemm_all2all_intra_node.py:221
↓ 1 callers
Function
_gen_inputs
(max_local_seq_len)
python/triton_dist/test/nvidia/test_llm_ulysess_all2all_gemm_intra_node.py:189
↓ 1 callers
Function
_gen_inputs
(max_local_seq_len)
python/triton_dist/test/nvidia/test_llm_ulysess_post_attn_all2all_intra_node.py:166
↓ 1 callers
Function
_gen_inputs
(max_local_seq_len, iter=0, is_debug=False)
python/triton_dist/test/nvidia/test_llm_ulysess_pre_attn_all2all_intra_node.py:245
↓ 1 callers
Function
_generate_codegen_func
Generate a codegen function for the simple_builtin. The codegen function will be called during code generation with C++ code strings
python/little_kernel/language/simple_builtin.py:299
↓ 1 callers
Method
_generate_cpp_struct
Generate the complete C++ struct definition.
python/little_kernel/codegen/special_struct/struct_converter.py:312
↓ 1 callers
Function
_generate_eval_return_type
Generate eval_return_type function from function signature. Args: func: Function with type annotations func_ast: Functio
python/little_kernel/language/simple_builtin.py:507
↓ 1 callers
Method
_generate_instance_method_call
Generate code for instance method call.
python/little_kernel/codegen/visitors/call_codegen.py:217
↓ 1 callers
Method
_generate_static_method_call
Generate code for static/class method call.
python/little_kernel/codegen/visitors/call_codegen.py:191
↓ 1 callers
Method
_generate_struct_name
Generate a unique, valid C++ struct name from StructType fields.
python/little_kernel/codegen/codegen_base.py:173
↓ 1 callers
Method
_generate_struct_name
Generate a unique, valid C++ struct name from StructType fields.
python/little_kernel/codegen/registries/type_converter.py:43
↓ 1 callers
Function
_get_active_nvlinks_pynvml
(gpu_index: int)
python/triton_dist/nv_utils.py:122
↓ 1 callers
Function
_get_bus_bw_gbps_a2a_rocm
()
python/triton_dist/amd_utils.py:368
↓ 1 callers
Function
_get_bus_bw_gbps_between_amdsmi
(device_id_i, device_id_j)
python/triton_dist/amd_utils.py:300
↓ 1 callers
Function
_get_bus_bw_gbps_between_rocm
(device_id_i: int, device_id_j: int)
python/triton_dist/amd_utils.py:387
↓ 1 callers
Function
_get_cuda_capability
Return (major, minor) CUDA compute capability, or None if no CUDA GPU.
python/little_kernel/tests/conftest.py:31
↓ 1 callers
Function
_get_current_gpu_clock_rate_in_khz_amdsmi
(device_id: int)
python/triton_dist/amd_utils.py:94
↓ 1 callers
Function
_get_current_gpu_clock_rate_in_khz_rocm
(device_id=None)
python/triton_dist/amd_utils.py:72
↓ 1 callers
Function
_get_default_num_stages
()
python/triton_dist/test/nvidia/test_compile_aot.py:90
↓ 1 callers
Function
_get_default_num_xcds
()
python/triton_dist/kernels/amd/allgather_gemm.py:44
↓ 1 callers
Method
_get_default_value
Get default value string for C++ initialization.
python/little_kernel/codegen/special_struct/struct_converter.py:291
↓ 1 callers
Method
_get_dtype
Convert tensor dtype to CUDA tensor map dtype.
python/little_kernel/runtime/tma_descriptor.py:133
↓ 1 callers
Function
_get_gpu_clock_mhz
Query current SM clock frequency via nvidia-smi.
python/little_kernel/benchmark/utils.py:67
↓ 1 callers
Function
_get_gpu_numa_node
(gpu_index=0)
python/triton_dist/nv_utils.py:104
↓ 1 callers
Function
_get_gpu_performance_mode_amdsmi
(device_id: int)
python/triton_dist/amd_utils.py:452
↓ 1 callers
Function
_get_gpu_performance_mode_nvsmi
(device_id: int)
python/triton_dist/nv_utils.py:382
↓ 1 callers
Function
_get_gpu_performance_mode_pynvml
(device_id: int)
python/triton_dist/nv_utils.py:386
↓ 1 callers
Function
_get_gpu_performance_mode_rocm
(device_id: int)
python/triton_dist/amd_utils.py:447
↓ 1 callers
Function
_get_gpu_uuid
(device_id: int)
python/triton_dist/amd_utils.py:261
↓ 1 callers
Function
_get_gpu_uuid_by_physical_device_id
(device_id: int)
python/triton_dist/amd_utils.py:207
↓ 1 callers
Function
_get_intrin_info
Try to detect if a Call node is calling an intrin/builtin function and return its info. Args: node: The Call AST node ct
python/little_kernel/core/passes/utils/enhanced_dump.py:48
↓ 1 callers
Method
_get_intrin_info
Try to detect if a Call node is calling an intrin/builtin function.
python/little_kernel/core/passes/utils/enhanced_unparse.py:97
↓ 1 callers
Function
_get_max_gpu_clock_rate_in_khz_amdsmi
(device_id)
python/triton_dist/amd_utils.py:112
↓ 1 callers
Function
_get_max_gpu_clock_rate_in_khz_rocm
(device_id)
python/triton_dist/amd_utils.py:122
↓ 1 callers
Method
_get_num_1d_blocks_per_group
Calculate optimal number of 1D blocks per group.
python/little_kernel/atom/scheduler.py:132
↓ 1 callers
Function
_get_numa_node_amdsmi
(device_id: int)
python/triton_dist/amd_utils.py:146
↓ 1 callers
Function
_get_numa_node_pynvml
(gpu_index)
python/triton_dist/nv_utils.py:217
↓ 1 callers
Function
_get_numa_node_rocm
Uses `rocm-smi --showtoponuma --json` and returns {"card0": 0, ...}
python/triton_dist/amd_utils.py:153
↓ 1 callers
Function
_get_nvlink_max_speed_gbps_nvsmi
Returns total NVLink bandwidth in GB/s for specified GPU
python/triton_dist/nv_utils.py:168
↓ 1 callers
Function
_get_nvlink_max_speed_gbps_pynvml
(gpu_index=0)
python/triton_dist/nv_utils.py:147
↓ 1 callers
Function
_get_nvml_gpu_uuid
(device_id: int)
python/triton_dist/nv_utils.py:342
↓ 1 callers
Function
_get_pcie_link_max_speed_gbps_nvsmi
Returns the maximum PCIe link speed in GB/s for specified GPU
python/triton_dist/nv_utils.py:261
↓ 1 callers
Function
_get_pcie_link_max_speed_gbps_pynvml
(gpu_index=0)
python/triton_dist/nv_utils.py:268
↓ 1 callers
Function
_get_peer_tensor
(t, peer)
python/triton_dist/utils.py:278
↓ 1 callers
Function
_get_physical_gpu_uuid_nvsmi
(device_id: int)
python/triton_dist/nv_utils.py:365
↓ 1 callers
Function
_get_physical_gpu_uuid_pynvml
(device_id: int)
python/triton_dist/nv_utils.py:359
↓ 1 callers
Function
_get_physical_gpu_uuid_rocm
(device_id: int)
python/triton_dist/amd_utils.py:243
↓ 1 callers
Method
_get_swizzle_mode
Convert swizzle mode to CUDA enum.
python/little_kernel/runtime/tma_descriptor.py:150
↓ 1 callers
Method
_get_tune_log_path
(self, key, pg: torch.distributed.ProcessGroup | None)
python/triton_dist/tune.py:428
↓ 1 callers
Function
_get_xgmi_max_speed_gbps_amdsmi
()
python/triton_dist/amd_utils.py:404
↓ 1 callers
Function
_get_xgmi_max_speed_gbps_rocm
()
python/triton_dist/amd_utils.py:418
↓ 1 callers
Method
_handle_align_call
Handle `ll.align_memory` calls, extract align_bytes and scope, then update memory analysis.
python/little_kernel/core/passes/mem_analysis.py:63
↓ 1 callers
Method
_handle_builtin_call
Handle builtin function calls.
python/little_kernel/codegen/visitors/call_codegen.py:299
↓ 1 callers
Method
_handle_constant_func_call
Handle function calls where func is a Constant.
python/little_kernel/codegen/visitors/call_codegen.py:355
↓ 1 callers
Method
_handle_empty_call
Handle `ll.empty` calls.
python/little_kernel/core/passes/mem_analysis.py:123
↓ 1 callers
Method
_handle_single_name_assignment
Handle single name assignment (e.g., x = 5).
python/little_kernel/codegen/visitors/statement_codegen.py:309
↓ 1 callers
Method
_handle_special_struct_method_call
Handle special struct method calls.
python/little_kernel/codegen/visitors/call_codegen.py:372
↓ 1 callers
Method
_handle_struct_stub_constructor_assignment
Handle struct stub constructor assignment.
python/little_kernel/codegen/visitors/statement_codegen.py:269
↓ 1 callers
Method
_handle_struct_stub_constructor_call
Handle struct stub constructor calls.
python/little_kernel/codegen/visitors/call_codegen.py:269
↓ 1 callers
Method
_handle_struct_stub_method_call
Handle struct stub method calls.
python/little_kernel/codegen/visitors/call_codegen.py:167
↓ 1 callers
Method
_handle_struct_stub_tuple_unpack
Handle tuple unpacking for struct stub method calls.
python/little_kernel/codegen/visitors/statement_codegen.py:435
↓ 1 callers
Method
_handle_subscript_assignment
Handle subscript assignment (e.g., latency[0] = value).
python/little_kernel/codegen/visitors/statement_codegen.py:369
↓ 1 callers
Method
_handle_tuple_assignment
Handle tuple assignment (e.g., (x, y) = (5, 3.14)).
python/little_kernel/codegen/visitors/statement_codegen.py:381
↓ 1 callers
Method
_handle_tuple_unpacking_from_call
Handle tuple unpacking from function call: a, b = xxx.func() -> a = xxx.func(..., b). First target receives return value, others are
python/little_kernel/codegen/special_struct/translator.py:648
↓ 1 callers
Method
_handle_zeros_call
Handle `ll.zeros` like ll.empty (same shape/dtype/scope, alloc+zero).
python/little_kernel/core/passes/mem_analysis.py:119
↓ 1 callers
Method
_has_duplicate_kwargs
(kwargs1, kwargs2)
python/triton_dist/tune.py:318
← previous
next →
1,001–1,100 of 5,103, ranked by callers