Repository navigation
fix(perfmodel): sweep tile shapes in the Triton MXFP4/MXFP6 microbenchmark - #1098
mehdi-saeedi wants to merge 3 commits into
Conversation
…hmark The tl.dot_scaled MXFP4/MXFP6 GEMM benchmark launched one fixed tile (128x128x256, 8 warps), which suits gfx950 but is a poor fit for RDNA4 (gfx1201), where smaller BLOCK_K tiles are much faster on the same shapes. The reported peak therefore depended on whether that one tile happened to fit the target. Try a small set of tiles per shape (MX_TILE_CONFIGS, starting with the original tile so no target can regress) and keep the fastest; tiles that do not divide the shape or fail to compile are skipped. Each entry also carries num_stages: None keeps Triton's backend default (the original behaviour), and a single-stage tile is included because dropping software pipelining helps further on RDNA4. Also import the modules in the test helpers _import_microbench, _import_microbench_rocprof and _import_fp4fp6_helpers, which returned undefined names, so every torch-enabled test using them failed with a NameError (CI without torch skips them). Co-authored-by: Cursor <cursoragent@cursor.com>
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
|
|
||
| def _import_fp4fp6_helpers(): | ||
| _require_torch() | ||
| from TraceLens.PerfModel.benchmarking import fp4fp6_helpers as fp |
There was a problem hiding this comment.
Can we gather all imports up top?
There was a problem hiding this comment.
Done in 40738fb. fp4fp6_helpers, microbench and microbench_rocprof are now imported once at the top of the module. The import sits behind an ImportError guard, because those modules import torch and CI runs without it; the _import_* helpers keep their torch skip and just return the module. Same results either way: 40 passed with torch, and 12 passed / 28 skipped without it, as before.
| # entry is the original fixed tile, which suits gfx950; on RDNA4 the smaller | ||
| # BLOCK_K tiles are much faster, and single-stage (no software pipelining) | ||
| # helps further. | ||
| MX_TILE_CONFIGS: Tuple[Tuple[int, int, int, int, Optional[int]], ...] = ( |
There was a problem hiding this comment.
Can we gather env vars up top as well?
There was a problem hiding this comment.
Done in 40738fb: I read this as the module-level settings, since the file reads no environment variables. MX_TILE_CONFIGS now sits at the top beside MX_BLOCK. Happy to move anything else if you meant something different.
|
@mohbasit to sign off |
…the top - The test module imports fp4fp6_helpers, microbench and microbench_rocprof once at the top, behind an ImportError guard because they import torch and CI runs without it; the _import_* helpers keep their torch skip and return the module. - MX_TILE_CONFIGS moves up beside MX_BLOCK with the module's other constants. Co-authored-by: Cursor <cursoragent@cursor.com>
Problem
The Triton MXFP4/MXFP6 GEMM microbenchmark launches a single fixed tile (128x128x256, 8 warps). That tile suits gfx950 but greatly understates other targets such as RDNA4 (gfx1201), so the reported peak depends on whether that one tile happens to fit.
Fix
MX_TILE_CONFIGS, and keep the fastest result. The first entry is the original tile, so no target can regress.(BLOCK_M, BLOCK_N, BLOCK_K, num_warps, num_stages).num_stages=Nonekeeps Triton's backend default (the original behaviour). One single-stage tile is included because dropping software pipelining helps further on RDNA4.Also
The test helpers
_import_microbench,_import_microbench_rocprofand_import_fp4fp6_helpersreturned undefined names, so every torch-enabled test that used them failed withNameError. CI runs without torch and skips these tests, which hid the bug. The helpers now import and return the modules.Tests
This touches
fp4fp6_helpers.py, as does the companion PR for MX peaks on targets without native FP4/FP6; the two merge cleanly in either order.