Skip to content

fix(perfmodel): sweep tile shapes in the Triton MXFP4/MXFP6 microbenchmark - #1098

Open
mehdi-saeedi wants to merge 3 commits into
AMD-AGI:mainfrom
mehdi-saeedi:fix/perfmodel/mx-microbench-tile-sweep
Open

mehdi-saeedi wants to merge 3 commits into
AMD-AGI:mainfrom
mehdi-saeedi:fix/perfmodel/mx-microbench-tile-sweep

Conversation

@mehdi-saeedi

Copy link
Copy Markdown
Contributor

Problem

The Triton MXFP4/MXFP6 GEMM microbenchmark launches a single fixed tile (128x128x256, 8 warps). That tile suits gfx950 but greatly understates other targets such as RDNA4 (gfx1201), so the reported peak depends on whether that one tile happens to fit.

Fix

  • Sweep a short tile list, MX_TILE_CONFIGS, and keep the fastest result. The first entry is the original tile, so no target can regress.
  • Each entry is (BLOCK_M, BLOCK_N, BLOCK_K, num_warps, num_stages). num_stages=None keeps Triton's backend default (the original behaviour). One single-stage tile is included because dropping software pipelining helps further on RDNA4.
  • Tiles that don't divide the shape (the kernel has no bounds masks) or that fail to compile or launch are skipped. If every applicable tile fails, the last error is raised.

Also

The test helpers _import_microbench, _import_microbench_rocprof and _import_fp4fp6_helpers returned undefined names, so every torch-enabled test that used them failed with NameError. CI runs without torch and skips these tests, which hid the bug. The helpers now import and return the modules.

Tests

  • The original tile stays first.
  • The fastest tile is kept and non-dividing tiles are skipped.
  • An error is raised when every tile fails.

This touches fp4fp6_helpers.py, as does the companion PR for MX peaks on targets without native FP4/FP6; the two merge cleanly in either order.

…hmark

The tl.dot_scaled MXFP4/MXFP6 GEMM benchmark launched one fixed tile
(128x128x256, 8 warps), which suits gfx950 but is a poor fit for RDNA4
(gfx1201), where smaller BLOCK_K tiles are much faster on the same
shapes. The reported peak therefore depended on whether that one tile
happened to fit the target.

Try a small set of tiles per shape (MX_TILE_CONFIGS, starting with the
original tile so no target can regress) and keep the fastest; tiles that
do not divide the shape or fail to compile are skipped. Each entry also
carries num_stages: None keeps Triton's backend default (the original
behaviour), and a single-stage tile is included because dropping
software pipelining helps further on RDNA4.

Also import the modules in the test helpers _import_microbench,
_import_microbench_rocprof and _import_fp4fp6_helpers, which returned
undefined names, so every torch-enabled test using them failed with a
NameError (CI without torch skips them).

Co-authored-by: Cursor <cursoragent@cursor.com>
@codecov-commenter

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

Comment thread tests/test_perfmodel_benchmarking.py Outdated

def _import_fp4fp6_helpers():
_require_torch()
from TraceLens.PerfModel.benchmarking import fp4fp6_helpers as fp

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we gather all imports up top?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done in 40738fb. fp4fp6_helpers, microbench and microbench_rocprof are now imported once at the top of the module. The import sits behind an ImportError guard, because those modules import torch and CI runs without it; the _import_* helpers keep their torch skip and just return the module. Same results either way: 40 passed with torch, and 12 passed / 28 skipped without it, as before.

# entry is the original fixed tile, which suits gfx950; on RDNA4 the smaller
# BLOCK_K tiles are much faster, and single-stage (no software pipelining)
# helps further.
MX_TILE_CONFIGS: Tuple[Tuple[int, int, int, int, Optional[int]], ...] = (

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we gather env vars up top as well?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done in 40738fb: I read this as the module-level settings, since the file reads no environment variables. MX_TILE_CONFIGS now sits at the top beside MX_BLOCK. Happy to move anything else if you meant something different.

@tsrikris

tsrikris commented Oct 9, 2026

Copy link
Copy Markdown
Contributor

@mohbasit to sign off

tsrikris and others added 2 commits October 9, 2026 12:17
…the top

- The test module imports fp4fp6_helpers, microbench and
  microbench_rocprof once at the top, behind an ImportError guard
  because they import torch and CI runs without it; the _import_*
  helpers keep their torch skip and return the module.
- MX_TILE_CONFIGS moves up beside MX_BLOCK with the module's other
  constants.

Co-authored-by: Cursor <cursoragent@cursor.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants