Skip to content

fix(perfmodel): roof Triton bf16 x mxfp4 MoE GEMMs at the activation width - #1096

Open
mehdi-saeedi wants to merge 3 commits into
AMD-AGI:mainfrom
mehdi-saeedi:fix/perfmodel/triton-mxfp4-moe-precision
Open

mehdi-saeedi wants to merge 3 commits into
AMD-AGI:mainfrom
mehdi-saeedi:fix/perfmodel/triton-mxfp4-moe-precision

Conversation

@mehdi-saeedi

Copy link
Copy Markdown
Contributor

Problem

moe_triton_unfused_up and moe_triton_unfused_down take the compute precision from the weight dtype. For a bf16 x mxfp4 MoE GEMM (e.g. Triton matmul_ogs) this selects matrix_fp4, but the math runs at the activation width: the weights are dequantized and multiplied in 16-bit. The kernel is then compared against the wrong peak.

Fix

Use the wider of the input and weight dtypes (_wider_compute_precision). If the wider spelling has no simulation equivalent, fall back to the weight dtype, which is the previous behaviour.

Reference CSVs updated

Per CONTRIBUTING.md, these reference outputs change:

  • tests/traces/inference/vllm_decode_full/perf_csvs/MoE_unfused_fwd.csv
  • tests/traces/inference/vllm_decode_full/perf_csvs/unified_perf_summary.csv
  • tests/traces/inference/vllm_prefilldecode_piecewise/perf_csvs/MoE_unfused_fwd.csv
  • tests/traces/inference/vllm_prefilldecode_piecewise/perf_csvs/unified_perf_summary.csv

The only change is the Compute Spec column, matrix_fp4 -> matrix_bf16, in the 8 Triton MoE rows. These fixtures run without a GPU arch file, so no roofline columns change, and no other columns or rows change.

Tests

  • New test_moe_triton_unfused_roofs_at_activation_width, including an fp8 case that covers the fallback.

…width

moe_triton_unfused_up/down took their compute precision from the weight
dtype, so vLLM's _matmul_ogs_*_bf16xbf16xmxfp4 kernels were roofed at the
FP4 matrix peak. These kernels convert the MXFP4 weights and multiply at
the 16-bit activation width, so prefill instances reported several times
100% of roofline on GPUs whose FP4 peak is low or emulated.

Use the wider of the activation and weight dtypes, falling back to the
weight dtype when the wider spelling has no simulation equivalent, and
update the two vLLM GPT-OSS reference reports, where only Compute Spec
changes for the two Triton MoE rows (matrix_fp4 -> matrix_bf16).

Co-authored-by: Cursor <cursoragent@cursor.com>
@codecov-commenter

codecov-commenter commented Oct 7, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

Kernel names with no fp4/fp8 marker leave weight_dtype unset; check that
moe_triton_unfused_up/down still report no compute precision there.

Co-authored-by: Cursor <cursoragent@cursor.com>

@tsrikris tsrikris left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants