Repository navigation
[COMPAT] Require Triton 3.8 - #485
Merged
Merged
Conversation
Raise the Triton floor from 3.6.0 to 3.8.0. Main no longer imported its Gluon simulator on 3.6 (gfx1250 `cluster` is missing there), so the old floor was already inaccurate; CI has been testing 3.8.0 (torch 2.14.1). With 3.8 as the minimum, drop the fallbacks that only served older releases: the gfx1250 and Blackwell `clc` import guards and the None handling behind them, the local copies of `_mxfp_value_handle_to_float32` and `_unpack_e2m1`, the `tensor_descriptor_base` guard, and the test skips for CDNA4 async copy, TMA im2col, and fp4-padded tensor memory.
Performance Benchmark
Iterations: 1 warmup + 20 measured |
Jokeren
approved these changes
Oct 3, 2026
This was referenced Oct 3, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Raises the minimum Triton version from 3.6.0 to 3.8.0 and removes the compatibility code that only existed for older releases.
The old
triton>=3.6.0floor was already inaccurate: on Triton 3.6,import tilelens.core.simulation.gluonfails becausetriton.experimental.gluon.language.amd.gfx1250has noclusterthere. CI has been resolving Triton 3.8.0 with torch 2.14.1 (torch 2.14 pinstriton~=3.8.0), so 3.8 is what is actually tested.Changes
pyproject.toml:triton>=3.6.0→triton>=3.8.0.amd.gfx1250(and itsasync_copy/cluster/mbarrier/tdm) andnvidia.blackwell.clcwith plain imports, and remove theNonehandling that depended on them (get_partitioned_shared_layout,_GLUON_BUILTIN_MODULES,_GLUON_BUILTIN_CLASSES,_GLUON_NON_SEMANTIC_BUILTINS,_existing_ops)._mxfp_value_handle_to_float32and_unpack_e2m1fromtriton.runtime.interpreterand delete the local fallback copies (Triton 3.6 did not have these helpers; 3.8 does).tl.core.tensor_descriptor_basewithout ahasattrguard. Update the comments that cited<3.6support. TileLens keeps its own snapshot scope because it covers bothtlandtl.corefor nested and generated kernels, while Triton's scope only records the modules visible fromfn.tensor_descriptor_basepatch-scope test.On Triton 3.8.0, each removed fallback's primary path exists. The Ampere
async_copy_global_to_localskip intest_gluon.pyis unchanged: that name does not exist in 3.6 or 3.8, so it is not a version shim.Testing
tilelens.core.frontend.gluonandtilelens.core.simulation.gluonimport cleanly.pre-commitpasses on the changed files (ruff, ruff-format, mypy, codespell).pytest --collect-onlyontests/end_to_end/test_gluon.pyandtests/unit/test_patch_scope.pyworks.Follow-ups