Skip to content

[FEAT] Add the TTIR reader - #495

Open
mark14wu wants to merge 1 commit into
ir-mode-corefrom
split/ttir-reader
Open

mark14wu wants to merge 1 commit into
ir-mode-corefrom
split/ttir-reader

Conversation

@mark14wu

@mark14wu mark14wu commented Oct 4, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Adds the TTIR reader that IR clients build on: tilelens.ir.ttir_reader.parse_ttir reads one kernel specialization's TTIR into an AccessGraph — the function arguments, every global load, store and atomic as an element offset from a pointer argument with its mask and branch path, the loop and its loop-carried pointers, and the integer width obligations under which the unbounded-integer reading is exact. What it cannot represent it refuses with an UnsupportedTTIR of a typed kind (indirect-address, control-flow, call, inline-asm, ...).

Part of the stack #494 → #479 → #495 → #480 → #482.

How the TTIR is read

tilelens/ir/_mlir_walk.py parses the text with Triton 3.8's MLIR bindings and takes from them the op tree, values, locs and the integer and enum attributes the reader needs (arith.cmpi predicate, program-id axes, make_range bounds, expand_dims axis, integer constants). A line scan of the same text supplies only what the bindings cannot: op line numbers, the locs of ops without results (e.g. tt.store), scf.for's unsigned keyword and a generic-form tt.load's operand segments. Every printed loc of an op with results must equal the loc the bindings report, so a misread line is a refusal, never a silent misread. Inputs that would abort the process inside Triton (block-pointer types, integer constants wider than 64 bits) never reach the bindings.

Tests

The tests compile their TTIR at test time instead of reading committed goldens: tests/unit/ir/ttir_corpus.py compiles the kernels in ttir_kernels.py / reader_kernels.py up to the TTIR stage only (no GPU, no Triton cache, about 0.3 s for the whole corpus) and holds the nine hand-written texts inline. Triton 3.8, CPU only: 564 passed at this branch's tip.

Validation of the bindings-first walk against the previous text-based one on 7,152 distinct TTIR texts (566k ops): identical op trees, values and every attribute the reader reads; the reader's result is identical on all but one hand-written generic-form text, which is now read instead of refused.

Left out for now (kept on ir-mode-tests-archive): reader fields only a race detector reads (atomic details, pid axes, the loop's induction variable, DataDep.keep, ...), and hardening/meta tests (mutation, fuzz, independent extractor, process/thread/fork, determinism, smoke tests with no unique coverage).

tilelens.ir.ttir_reader.parse_ttir reads one kernel specialization's TTIR
into an AccessGraph: the function arguments, every global load, store and
atomic as an element offset from a pointer argument with its mask and
branch path, the loop and its loop-carried pointers, and the integer width
obligations under which the unbounded-integer reading is exact. What it
cannot represent it refuses with an UnsupportedTTIR of a typed kind
(indirect-address, control-flow, call, inline-asm, ...). ParseCache's
default reader is this one from now on.

The structure comes from tilelens.ir._mlir_walk. Triton 3.8's MLIR
bindings parse the text and give the op tree, the values, their locs, and
the integer and enum attributes the reader reads (the cmpi predicate, the
program-id axes, make_range's bounds, expand_dims' axis, integer
constants). A line scan of the same text gives only what the bindings
cannot: each op's line and a region op's closing line, the locs of ops
without results (through the #loc alias table), scf.for's `unsigned`
keyword and a generic-form tt.load's operand segments. Every printed loc
of an op with results must equal the loc the bindings report for it, so a
misread line is a reader-misalignment refusal, never a silent misread.
Two inputs would abort the process inside Triton, so they never reach the
bindings: a block-pointer type is refused before parsing, and an integer
constant wider than 64 bits reads as non-integer data (a test runs that
case in a child process). On any other Triton release the walk refuses.

The tests compile their TTIR at test time instead of reading committed
goldens. tests/unit/ir/ttir_corpus.py compiles the kernels of
ttir_kernels.py and reader_kernels.py up to the TTIR stage only (no GPU,
no Triton cache, about 0.3 s for the whole corpus) and holds the nine
hand-written texts inline; text(name) keeps the former golden names
(ttir/<name>.ttir, reader_ttir/<name>.ttir). The corpus also holds the
four kernels the compiled sanitizer's tests read: golden_pid_branch,
golden_grid_stride, golden_cas and golden_gather (sm80).

Left out for now; they stay on ir-mode-tests-archive:
- reader fields only a race detector reads: AtomicInfo (and the atomics'
  rmw_op, sem and scope), AccessEvent.atomic/atomic_val/atomic_cmp/
  elem_float/is_read/is_write, FuncArg.elem_float, AccessGraph.pid_axes,
  LoopInfo.induction_var, DataDep.keep, observed_indices, and ttg.barrier's
  address space. Observed stays;
- hardening and meta tests: mutation, fuzz and independent-extractor
  tests, process, thread, fork and memory tests, determinism and frozen
  graph tests, and the end-to-end smoke tests, which cover no line the
  other tests leave uncovered.
@github-actions

github-actions Bot commented Oct 4, 2026

Copy link
Copy Markdown

Performance Benchmark

Benchmark main (min) PR (min) Change Samples
gemm 0.097s 0.100s +2.5% 20 / 20
gemm_oob 0.108s 0.109s +0.5% 20 / 20
indirect_load 0.020s 0.019s -0.2% 20 / 20
nested_loop 0.203s 0.202s -0.4% 20 / 20
block_pointer_loop_advance 0.114s 0.114s +0.5% 20 / 20
liger_jsd 0.136s 0.136s +0.3% 20 / 20
flaggems_layernorm 0.351s 0.351s -0.1% 20 / 20
swiglu 0.164s 0.163s -0.8% 20 / 20
cross_entropy 0.930s 0.932s +0.2% 20 / 20
fused_linear_jsd 0.203s 0.207s +1.7% 20 / 20
Total 2.325s 2.332s +0.3% N/A

Iterations: 1 warmup + 20 measured
Samples are shown as main / PR; long pytest benchmarks may use fewer samples.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant