Skip to content

[COMPAT] Require Triton 3.8 - #485

Merged
mark14wu merged 1 commit into
mainfrom
claude/triton-upgrade-3-8-cc420f
Oct 3, 2026
Merged

mark14wu merged 1 commit into
mainfrom
claude/triton-upgrade-3-8-cc420f

Conversation

@mark14wu

@mark14wu mark14wu commented Oct 3, 2026

Copy link
Copy Markdown
Collaborator

Summary

Raises the minimum Triton version from 3.6.0 to 3.8.0 and removes the compatibility code that only existed for older releases.

The old triton>=3.6.0 floor was already inaccurate: on Triton 3.6, import tilelens.core.simulation.gluon fails because triton.experimental.gluon.language.amd.gfx1250 has no cluster there. CI has been resolving Triton 3.8.0 with torch 2.14.1 (torch 2.14 pins triton~=3.8.0), so 3.8 is what is actually tested.

Changes

  • pyproject.toml: triton>=3.6.0 → triton>=3.8.0.
  • Gluon simulator and frontend: replace the guarded imports of amd.gfx1250 (and its async_copy / cluster / mbarrier / tdm) and nvidia.blackwell.clc with plain imports, and remove the None handling that depended on them (get_partitioned_shared_layout, _GLUON_BUILTIN_MODULES, _GLUON_BUILTIN_CLASSES, _GLUON_NON_SEMANTIC_BUILTINS, _existing_ops).
  • Gluon simulator: import _mxfp_value_handle_to_float32 and _unpack_e2m1 from triton.runtime.interpreter and delete the local fallback copies (Triton 3.6 did not have these helpers; 3.8 does).
  • Triton frontend: snapshot tl.core.tensor_descriptor_base without a hasattr guard. Update the comments that cited <3.6 support. TileLens keeps its own snapshot scope because it covers both tl and tl.core for nested and generated kernels, while Triton's scope only records the modules visible from fn.
  • Tests: remove skips that never fire on 3.8: CDNA4 async copy, TMA im2col (×3), fp4-padded tensor memory, and the tensor_descriptor_base patch-scope test.

On Triton 3.8.0, each removed fallback's primary path exists. The Ampere async_copy_global_to_local skip in test_gluon.py is unchanged: that name does not exist in 3.6 or 3.8, so it is not a version shim.

Testing

  • Checked on Triton 3.8.0 (torch 2.14.1) that every removed import or attribute fallback resolves to its primary path, and that tilelens.core.frontend.gluon and tilelens.core.simulation.gluon import cleanly.
  • pre-commit passes on the changed files (ruff, ruff-format, mypy, codespell).
  • pytest --collect-only on tests/end_to_end/test_gluon.py and tests/unit/test_patch_scope.py works.
  • I did not run the full test suite locally; this PR's CI covers it on Triton 3.8.0.

Follow-ups

Raise the Triton floor from 3.6.0 to 3.8.0. Main no longer imported its
Gluon simulator on 3.6 (gfx1250 `cluster` is missing there), so the old
floor was already inaccurate; CI has been testing 3.8.0 (torch 2.14.1).

With 3.8 as the minimum, drop the fallbacks that only served older
releases: the gfx1250 and Blackwell `clc` import guards and the None
handling behind them, the local copies of `_mxfp_value_handle_to_float32`
and `_unpack_e2m1`, the `tensor_descriptor_base` guard, and the test
skips for CDNA4 async copy, TMA im2col, and fp4-padded tensor memory.
@github-actions

github-actions Bot commented Oct 3, 2026

Copy link
Copy Markdown

Performance Benchmark

Benchmark main (min) PR (min) Change Samples
gemm 0.105s 0.105s +0.0% 20 / 20
gemm_oob 0.117s 0.116s -0.2% 20 / 20
indirect_load 0.022s 0.022s -0.9% 20 / 20
nested_loop 0.237s 0.235s -0.9% 20 / 20
block_pointer_loop_advance 0.125s 0.126s +1.0% 20 / 20
liger_jsd 0.139s 0.139s +0.3% 20 / 20
flaggems_layernorm 0.396s 0.397s +0.2% 20 / 20
swiglu 0.169s 0.171s +0.9% 20 / 20
cross_entropy 0.977s 0.980s +0.3% 20 / 20
fused_linear_jsd 0.211s 0.210s -0.3% 20 / 20
Total 2.496s 2.500s +0.2% N/A

Iterations: 1 warmup + 20 measured
Samples are shown as main / PR; long pytest benchmarks may use fewer samples.

@mark14wu
mark14wu merged commit 2731482 into main Oct 3, 2026
4 checks passed
@mark14wu
mark14wu deleted the claude/triton-upgrade-3-8-cc420f branch October 3, 2026 03:04
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants