↓ 1 callersFunction_triton_dist_impl(input, weight, seq_lens_cpu, input_scale, weight_scale, bias)
python/triton_dist/test/nvidia/test_llm_ulysess_gemm_all2all_intra_node.py:265
↓ 1 callersFunctionact_mul_up_tile_compute(tile_id, input, output, M, N, ACT_FN, BLOCK_SIZE_M: tl.constexpr,
BLOCK_SIZE_N: t
python/triton_dist/mega_triton_kernel/kernels/activation.py:31
↓ 1 callersFunctionallreduce_op(ctx: GemmARContext, c, gemm_config: triton.Config, TILE_MAP_LEVEL=0, copy_to_local=True,
USE
python/triton_dist/kernels/nvidia/gemm_allreduce.py:812
↓ 1 callersFunctionatomic_add semantic should be one of ["monotonic", "release", "acquire", "acq_rel] scope should be one of ["workgroup", "agent", "system]
python/triton_dist/language/extra/hip/language_extra.py:162
↓ 1 callersFunctionbench_flash_attn(builder, BATCH, N_CTX, Q_HEAD, KV_HEAD, HEAD_DIM, IS_CAUSAL, dtype=torch.bfloat16, qkv_pack=False)
python/triton_dist/mega_triton_kernel/test/ops/test_flash_attn.py:76