Skip to content

[FEAT] Add the host-compiled IR lifecycle for clients - #479

Open
mark14wu wants to merge 1 commit into
split/trace-fixesfrom
ir-mode-core
Open

mark14wu wants to merge 1 commit into
split/trace-fixesfrom
ir-mode-core

Conversation

@mark14wu

@mark14wu mark14wu commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Lets a TileLens client analyze the kernels Triton compiles for a launch instead of interpreting it, without a GPU: a client declares NEEDS_INTERPRETER = False and receives, for every traced launch, the TTIR compiled on the host for each config. IR mode supports Triton 3.8 only.

Part of the stack #494 → #479 → #495 → #480 → #482. It sits on #494 (trace fixes). The TTIR reader comes in #495 and the first real client, the compiled sanitizer, in #480.

What's included

  • Core lifecycle (tilelens/core/client.py, trace.py)
    • LaunchCall / LaunchEvent and the begin_launch, abort_launch, before_launch, after_launch and compile_failed hooks; each begun launch ends in exactly one of finalize or abort_launch. A concurrent launch of one trace from another thread is refused.
    • Client declarations NEEDS_INTERPRETER, IR_STAGES, LAUNCH and ir_target; interpreter hooks reach interpreting clients only.
    • ClientManager.ir_capture routes every JITFunction.run call of the traced kernel through a host compile and delivers one event per (specialization, binding) per launch. A failing compile is compile_failed data and never fails the launch; a call that does not bind raises as it does untraced.
    • Every (pruned) autotune config is compiled through a separate IR runner chain. An IR-only launch never runs the real kernel; a mixed trace compiles for the IR clients, then interprets for the eager ones. Autotuner copies drop the call's tensors when the call returns or raises.
  • Host compile (tilelens/core/host_compile.py): binds and packs a call with the JIT's own binder and compiles it through the backend's TTIR passes for the target (cuda:89 by default, TILELENS_IR_TARGET, Client.ir_target), with target queries answered by that target and every device query refused. On a Triton other than 3.8 every IR launch is reported unsupported and the program continues.
  • tilelens.ir: IRClient (a finalize template over a per-launch artifact log), LaunchBinding / TensorFacts, a content-addressed ParseCache and the IRVerdict records, which tilelens.save() round-trips.
  • CI installs triton==3.8.0 and triton_kernels at v3.8.0, with UV_NO_SYNC=1 so uv run keeps that pin.

Not supported yet (refused instead of half-done)

  • Running the real kernel after the host compile: an IR client must declare LAUNCH = "skip".
  • Stages other than "ttir" in IR_STAGES.
  • Several IR targets in one trace.
  • Whole-pipeline compiles (TRITON_KERNEL_OVERRIDE, USE_IR_LOC, ir_override, custom pipelines).

Testing

Triton 3.8, CPU only: 373 passed at this branch's tip (this layer's lifecycle, host-compile, capture and client tests plus #494's), with no new failure in the eager test files.

The earlier, larger version of this PR (conformance suite, golden corpus, multi-release support) is preserved on branch ir-mode-tests-archive.

@github-actions

github-actions Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Performance Benchmark

Benchmark main (min) PR (min) Change Samples
gemm 0.106s 0.105s -0.4% 20 / 20
gemm_oob 0.117s 0.118s +0.9% 20 / 20
indirect_load 0.022s 0.022s -0.9% 20 / 20
nested_loop 0.236s 0.235s -0.3% 20 / 20
block_pointer_loop_advance 0.127s 0.126s -0.6% 20 / 20
liger_jsd 0.139s 0.139s -0.3% 20 / 20
flaggems_layernorm 0.397s 0.397s -0.0% 20 / 20
swiglu 0.169s 0.171s +0.9% 20 / 20
cross_entropy 0.973s 0.973s -0.0% 20 / 20
fused_linear_jsd 0.211s 0.212s +0.3% 20 / 20
Total 2.498s 2.498s +0.0% N/A

Iterations: 1 warmup + 20 measured
Samples are shown as main / PR; long pytest benchmarks may use fewer samples.

A client can now declare NEEDS_INTERPRETER = False and receive, for every
traced launch, the kernels Triton compiles for each config instead of the
interpreted run. The compile runs on the host for a configured GPU target,
so nothing needs a GPU.

Core lifecycle (tilelens/core/client.py, trace.py):
- LaunchCall / LaunchEvent; the begin_launch, abort_launch, before_launch,
  after_launch and compile_failed hooks; each begun launch ends in exactly
  one of finalize or abort_launch. begin_launch now creates the per-launch
  Launch; a concurrent launch of one trace from another thread is refused,
  and so is a second traced launch capturing the same JITFunction.
- Client declarations NEEDS_INTERPRETER, IR_STAGES, LAUNCH and ir_target.
  Interpreter hooks reach interpreting clients only.
- ClientManager.ir_capture routes every JITFunction.run call of the traced
  kernel through a host compile and delivers one event per (specialization,
  binding fingerprint) per launch; a failing compile is compile_failed data
  and never fails the launch; a call that does not bind raises as untraced.
- TritonTrace compiles every (pruned) config through a separate IR runner
  chain. An IR-only launch never launches the real kernel; a mixed trace
  compiles for the IR clients, then interprets for the eager ones.
- Each runner chain's Autotuner copies drop the call's arguments (nargs)
  and restore_value clones (restore_copies) when the call returns or
  raises, so no caller tensor outlives a launch that raised.

Host compile (tilelens/core/host_compile.py): binds and packs a call with
the JIT's own binder and compiles it through the backend's TTIR passes for
the target, with the front end's target queries answered by that target
and every device query refused. IR targets are "cuda:89" by default
(TILELENS_IR_TARGET, Client.ir_target); IR mode runs on Triton 3.8 only,
checked in one place (config.ir_triton_unsupported).

tilelens.ir: IRClient (a finalize template over a per-launch ArtifactLog),
LaunchBinding/TensorFacts, a content-addressed ParseCache and the IRVerdict
records, which saved traces round-trip. ParseCache's default reader is the
TTIR reader, which this change does not add: until it exists, a lookup
with the default reader is an error outcome. The tests use stub readers.

Not supported yet, refused instead of half-done:
- running the real kernel after the host compile: an IR client must declare
  LAUNCH = "skip";
- stages other than "ttir": IR_STAGES naming another stage is refused when
  the client is added;
- several IR targets in one trace: refused when the launch compiles;
- whole-pipeline compiles: TRITON_KERNEL_OVERRIDE, USE_IR_LOC, an
  ir_override option and a custom pipeline make the host compile raise
  HostCompileUnavailable.

CI installs triton==3.8.0 and triton_kernels at v3.8.0, with UV_NO_SYNC so
`uv run` keeps that pin.
@mark14wu
mark14wu removed this pull request from stack #481 October 4, 2026 23:48
@mark14wu
mark14wu changed the base branch from main to split/trace-fixes October 4, 2026 23:48
@mark14wu
mark14wu added this pull request to stack #496 October 4, 2026 23:48
@mark14wu mark14wu changed the title [FEAT] Add a host-compiled IR layer for clients [FEAT] Add the host-compiled IR lifecycle for clients Oct 4, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant