[Bench] Group benchmark rows by workload label and keep manifest order - #54
Merged
Merged
Conversation
The data pages sorted an op's rows by case id, so a size sweep read t1, t128, t2048, t32, and ops with several tensor templates numbered the key in one order and the table in another. Rows now keep the snapshot's order, which is the manifest's, and the table follows the key cluster by cluster. Each manifest label is one row group: the label once, spanning its dtype rows, and a dtype column beside it. The key lists each label once with the values that vary and the dtypes it ran at, so the W codes are gone. A row no manifest describes takes its id, trailing dtype names split off, as its label.
…cleanly A key row gave only the symbol values, so reading a shape meant substituting them into the template by hand. Each label now prints its shapes with the symbols substituted, and the symbol values are no longer listed beside them. A cluster with a single label moved every value above the label, since each counted as shared; values now go above only when two or more labels share them. A wrapped key row started its continuation with a middot, and tensors and scalars broke wherever the width ran out. The middot of an entry that opens a line is now clipped, the scalars wrap as one part after the tensors, and a cluster with a label over 32 characters sets each label on its own line.
Checked against TileOPs fc6f03c59, several pages described interfaces or numbers that main no longer has. Writing a Spec and Adding an Op now use main's workload labels and roofline formulas. A kernel implements `forward`, which the base class's `__call__` runs, and the smoke, full and nightly marks match what the PR and nightly jobs run. torch.compile: only a target-served call keys its kernel by device, and an operator name may also end in `_without_<output>`. Adding a Backend: the memo table is dropped when a failed call revokes the target decision, not bounded, and a null-defaulted param arrives as the value the op settled on. Timing: copies are collected and reported as `uncounted_copy_ms` unless the case passes `count_copies=True`, and a lost-records phase is measured at most three times in all. The home page no longer says kernels are auto-tuned on first use, since tuning is opt-in. The memory-bound roofline takes H200's HBM bandwidth from the current profile, 4.50 TB/s rather than 4.07, which moves the ridge to 12.72 flop/byte and silu to 10% of the compute ceiling; the figure is redrawn to those numbers. The API index states the one exception to family order, Top-k.
…ught in line with main Adding a Backend described an older tileops-backend-example: `kernels.py` and a `pending.py` registered under a misspelt key, `_detect`, four error paths, and two `requires_cuda_runtime` tests with 22 or 20 passing. The example now keeps `TARGET` and `detect` in `target.py`, one module per op under `ops/`, and a `BUILDERS` table that `__init__.py` registers in a loop; it has three error paths and 24 tests that pass with or without a visible GPU. The walkthrough, file table, test count and adaptation steps follow it. The sentences added in the previous commit are reworded where they read as fragments.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problems
FusedMoeSharedExpertFwdlistedt1, t128, t2048, t32, t4096, t512, t64.BmmFp8Fwd,GroupedQueryAttentionPrefillPagedWithKVCacheFwd,FusedTopKFwd,GemmFwd) numbered the table in one order and the key in another, soW4sat aboveW2in the key.Wrow, so a reader had to look up the code in the key to find which label and dtype a number belonged to.B=[2048],normalized_shape=[4096]), so reading a shape meant substituting them into[*B, *normalized_shape]by hand.LayerNormFwdllama-13b-prefill,dit-xl-2).__call__main replaced withforward, wrong test-mark roles, a device-keyed memo for every kernel, a bounded backend memo table, copies the timer supposedly never sees, and H200 bandwidth 4.07 TB/s where the profile now gives 4.50.kernels.py, apending.pyunder a misspelt key,_detect, four error paths, and tworequires_cuda_runtimetests with 22 or 20 passing.Changes
dtypecolumn (fp8e4m3/bf16for a case with two dtype indices). TheWcodes are gone.out_dtype,cache_dtype) is not repeated there.x: [2048, 4096]); the template stays once above the labels, and the symbol values are no longer listed.target.py, one module per op underops/, aBUILDERStable, three error paths, 24 tests passing with and without a visible GPU. Worked examples the manual still lacks (type families, ADTs, generators withrequires, composition, multi-casedtype_cases) are left to a separate PR.Before and after,
FusedMoeSharedExpertFwdon the 2026-09-27 snapshot: