Skip to content

[FIX] Keep every client a trace is given - #487

Open
mark14wu wants to merge 5 commits into
ir-mode-loweringfrom
claude/duplicate-client-instance-loss-ir
Open

mark14wu wants to merge 5 commits into
ir-mode-loweringfrom
claude/duplicate-client-instance-loss-ir

Conversation

@mark14wu

@mark14wu mark14wu commented Oct 3, 2026

Copy link
Copy Markdown
Collaborator

Stacked on #482 (ir-mode-lowering); the reproduction needs Sanitizer(compile=True) from #480. The core change is in tilelens/core/client.py, which is identical from #479 to #482, so it can also be folded lower in the stack.

Problem

ClientManager.add_clients silently dropped a client when one of its class was already in the trace, and a client of another class with the same NAME silently replaced the first. So two compiled sanitizers for different targets kept only the first:

sm80 = Sanitizer(compile=True, target="cuda:80", abort_on_error=False)
sm90 = Sanitizer(compile=True, target="cuda:90", abort_on_error=False)
traced = tilelens.trace(sm90)(tilelens.trace(sm80)(kernel))
traced[(4,)](x, 100, BLOCK=32)
# before: only cuda:80 is compiled; sm90.last_status is None, no error
# after:  both targets are compiled and checked, each with its own verdict

The same loss hit eager clients, e.g. Sanitizer(abort_on_error=False) stacked on an existing Sanitizer was dropped without a word.

Change

A client passed to a trace is now in the trace afterwards, or add_clients raises; nothing is dropped silently:

  • Adding the same object again changes nothing.
  • IR clients are all kept, so a trace can hold one instance of a client class per target. The core already compiled once per distinct target and delivered each client only its own target's events; the dedup rule was what kept a second instance out.
  • A second interpreting client of one NAME raises ValueError (with the "trace the kernel twice" advice the launch-conflict error gives), since one interpreted run serves one client per name.
  • A client named by string (trace("sanitizer")) is a no-op when the trace already has a client of that name: no settings of the caller's are lost, and the documented stacked-decorator case in test_trace_decorator_add_clients keeps working. The string is resolved to its class before anything is constructed.

API notes:

  • ClientManager.clients is now a list in trace order (it was a NAME-keyed dict); get_client(name) returns the first client of that NAME.
  • A same-NAME IR client no longer replaces the existing one; it is added beside it, so a skip/run conflict between them is now raised.
  • README: the compiled-sanitizer target section shows stacking one sanitizer per target.

Tests

  • New: test_sanitizers_for_two_targets_share_one_trace (end to end, cuda:80 ok / cuda:90 violations in one launch), test_instances_of_one_ir_client_class_keep_their_own_targets, test_add_clients_keeps_every_ir_client_instance, test_add_clients_refuses_a_second_interpreting_client_of_one_name, plus a refusal check in test_trace_decorator_add_clients. All of these fail on the base branch for the reason above.
  • Updated tests that read clients as a dict, and test_launch_conflict_check_sees_every_ir_client (formerly asserted the replacement).

Commands (Triton 3.6, CPU only):

CUDA_VISIBLE_DEVICES="" python -m pytest tests/unit/test_ir_lifecycle.py tests/unit/sanitizer_compiled/test_client.py tests/unit/test_wrapper.py tests/end_to_end/test_core.py::test_trace_decorator_add_clients tests/end_to_end/test_compiled_sanitizer.py
CUDA_VISIBLE_DEVICES="" python -m pytest tests/ --ignore=tests/end_to_end/test_gluon.py

Full suite: 1607 passed, 14 failed. The 14 fail identically on the base branch and are environment-related here: Gluon gfx1250.cluster import (6), tile-* executables not on PATH (5), and the three known test_sanitizer.py failures. test_gluon.py was ignored because it fails to collect for the same Gluon import reason.

Not in this PR

  • Two compiled-sanitizer verdicts in Launch.records both carry client="compiled_sanitizer" and no target; they are told apart by trace order or by each instance's last_verdict.
  • Separate, pre-existing: when two different interpreting clients share a trace, patch_op re-wraps the original op for each client, so only the last client's op callbacks fire.

Lets a client analyze a kernel's compiled IR instead of interpreting it,
without a GPU.

- core: clients declare NEEDS_INTERPRETER, IR_STAGES and LAUNCH; IR
  clients get launch events from a capture around jit_fn.run (every
  autotune/heuristics config, deduplicated by binding), take no part in
  op/loop patching or the pre_run vote, and conflicting launch
  preferences are refused at registration. Each launch gets its own
  Launch, client finalize is isolated, and runner chains are rebuilt
  instead of mutating the user's Autotuner/Heuristics.
- host compile: TTIR through the JIT's own binder and specialization,
  compiled on the host for a configurable target (default cuda:89,
  TILELENS_IR_TARGET); no driver or device access on the IR path.
- tilelens.ir: a TTIR reader that walks the MLIR bindings and reads the
  attributes they cannot expose from the aligned text; the AccessGraph
  term model with bit widths and width obligations; capture, launch
  binding, IRClient base and IRVerdict records (saved by tilelens.save).
- Tested on Triton 3.6 and 3.8; other releases are refused unless
  TILELENS_IR_ALLOW_UNTESTED_TRITON is set, and IR tests skip there.
- Tests: reader conformance suite (static footprint vs Triton's
  interpreter), golden TTIR per release, lifecycle and host-compile
  tests, and tools/ir_bulk_conformance.py.
The golden regeneration test took the generator's path out of the locs
but not Triton's: a golden whose kernel calls into Triton's own sources
(tl.cdiv, tl.zeros, ...) names the directory Triton is installed at, so
it never regenerated byte for byte on another machine (CI failed on
golden_matmul_tma_s1_sm90 under Triton 3.8). Take both paths out before
comparing.
Sanitizer(compile=True), or tile-sanitizer --compile, checks a launch
statically on the host-compiled TTIR instead of interpreting it. The
kernel is not launched, so no GPU is needed and outputs are not written.

- Every autotune config is checked with Z3 against the launch's scalar
  arguments, grid and tensor layouts; a launch is ok only if every config
  is proven in bounds.
- Findings: out-of-bounds (against the view's element footprint, gaps
  of strided views included), integer-overflow on address, mask, branch
  and loop values, and division-by-zero, each with a witness.
- Anything that cannot be modeled is reported as unsupported with a
  typed refusal; a solver unknown or timeout never counts as a proof.
  A kernel that fails to compile for the target is unsupported and the
  program continues; a call that does not bind raises as untraced.
- Each launch records an IRVerdict with per-config verdicts in
  Launch.records.
The compiled sanitizer's evaluator becomes tilelens.ir.lowering, the one
reading of the TTIR reader's term algebra as Z3 terms, so the compiled
race detector can sit on the same semantics instead of a second copy.

- lowering.py is mechanism only: operator semantics (truncating division,
  unsigned ops and predicates as their signed twins under the width
  obligations, IntCast as its operand, i1 coercion, loop terms) live
  there; every leaf (scalar arguments, pids, grid, lanes, the loop
  iteration, observations, unmodeled values) comes from the client's
  TermLeaves, which also raises the client's own refusals.
- It never creates a Z3 context: constants use the leaves' context and
  every other term its operands', so a client keeps all of its terms in
  one context (per check for the sanitizer).
- The walk is iterative and memoized by term identity, so terms deeper
  than the recursion limit lower.
- The sanitizer keeps its policy (obligation findings, refusals, solving,
  timeouts). Its leaves reach the lowering through a weak proxy and a
  refused loop keeps a never-raised copy of its refusal, so a check's Z3
  context is freed when the check returns instead of by the cyclic GC on
  another thread (concurrent checks could hang or crash).
- Its walks stop at the induction variable, as before, so a loop bound
  read only through a Select arm keeps that arm's guard.

No verdict, finding or witness changes: the sanitizer's results match
the previous evaluator on the golden texts and the differential corpus
on Triton 3.6 and 3.8. Hand-built graphs outside the reader's
invariants (e.g. an iter arg naming another loop) now raise instead of
being misread.
ClientManager.add_clients dropped a client silently when one of its
class was already in the trace, and a client of another class with the
same NAME replaced the first one. So two Sanitizer(compile=True)
instances for different targets left only the first: the second never
compiled, its last_status stayed None, and nothing said so.

A client passed to a trace is now in it afterwards, or add_clients
raises:

- adding the same object again changes nothing;
- IR clients are all kept, so a trace can hold one instance of a class
  per target, each compiled for its own target with its own verdict;
- a second interpreting client of one NAME raises ValueError, since one
  interpreted run serves one client per name;
- a client named by string ("sanitizer") is a no-op when the trace
  already has one of that name, as no settings of the caller's are lost.

ClientManager.clients is now a list in trace order, and get_client
returns the first client of a NAME.
@mark14wu
mark14wu force-pushed the ir-mode-lowering branch 2 times, most recently from 8d30a3c to b6ab11c Compare October 4, 2026 23:47

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant