Repository navigation
fix(perfmodel): roof Triton bf16 x mxfp4 MoE GEMMs at the activation width - #1096
Open
mehdi-saeedi wants to merge 3 commits into
Open
mehdi-saeedi wants to merge 3 commits into
mehdi-saeedi wants to merge 3 commits into
Conversation
…width moe_triton_unfused_up/down took their compute precision from the weight dtype, so vLLM's _matmul_ogs_*_bf16xbf16xmxfp4 kernels were roofed at the FP4 matrix peak. These kernels convert the MXFP4 weights and multiply at the 16-bit activation width, so prefill instances reported several times 100% of roofline on GPUs whose FP4 peak is low or emulated. Use the wider of the activation and weight dtypes, falling back to the weight dtype when the wider spelling has no simulation equivalent, and update the two vLLM GPT-OSS reference reports, where only Compute Spec changes for the two Triton MoE rows (matrix_fp4 -> matrix_bf16). Co-authored-by: Cursor <cursoragent@cursor.com>
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
Kernel names with no fp4/fp8 marker leave weight_dtype unset; check that moe_triton_unfused_up/down still report no compute precision there. Co-authored-by: Cursor <cursoragent@cursor.com>
17 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
moe_triton_unfused_upandmoe_triton_unfused_downtake the compute precision from the weight dtype. For a bf16 x mxfp4 MoE GEMM (e.g. Tritonmatmul_ogs) this selectsmatrix_fp4, but the math runs at the activation width: the weights are dequantized and multiplied in 16-bit. The kernel is then compared against the wrong peak.Fix
Use the wider of the input and weight dtypes (
_wider_compute_precision). If the wider spelling has no simulation equivalent, fall back to the weight dtype, which is the previous behaviour.Reference CSVs updated
Per
CONTRIBUTING.md, these reference outputs change:tests/traces/inference/vllm_decode_full/perf_csvs/MoE_unfused_fwd.csvtests/traces/inference/vllm_decode_full/perf_csvs/unified_perf_summary.csvtests/traces/inference/vllm_prefilldecode_piecewise/perf_csvs/MoE_unfused_fwd.csvtests/traces/inference/vllm_prefilldecode_piecewise/perf_csvs/unified_perf_summary.csvThe only change is the
Compute Speccolumn,matrix_fp4->matrix_bf16, in the 8 Triton MoE rows. These fixtures run without a GPU arch file, so no roofline columns change, and no other columns or rows change.Tests
test_moe_triton_unfused_roofs_at_activation_width, including an fp8 case that covers the fallback.