Repository navigation
SPIRVIntrinsics: sub-group votes and collectives, unchecked shuffle lanes - #526
Merged
Merged
Conversation
vchuravy
added a commit
to JuliaGPU/KernelAbstractions.jl
that referenced
this pull request
Oct 4, 2026
Replace the local workarounds with the votes and unchecked shuffle lanes from JuliaGPU/OpenCL.jl#526, taken from its branch until it is released: through `[sources]`, and explicitly where that doesn't apply (Julia 1.10 on CI, and the Buildkite jobs, whose OpenCL job developed SPIRVIntrinsics from OpenCL.jl's ka-0.10 branch). Assisted-by: Claude Code (Opus 5.5)
This was referenced Oct 4, 2026
Open
vchuravy
added a commit
to JuliaGPU/KernelAbstractions.jl
that referenced
this pull request
Oct 4, 2026
Implement `KI.sub_group_reduce` and `KI.sub_group_scan` with the collectives of `cl_khr_subgroups` from SPIRVIntrinsics (JuliaGPU/OpenCL.jl#526): for `+` on 32- and 64-bit integers and floats, and `min`/`max` on integers. Floats keep the fallback for `min` and `max`, as OpenCL treats NaN and the sign of zero differently. Test the operators and types that backends may implement natively, including a NaN, and that POCL uses the native reduction. Assisted-by: Claude Code (Opus 5.5)
vchuravy
added a commit
to JuliaGPU/KernelAbstractions.jl
that referenced
this pull request
Oct 4, 2026
The wrong results of the native `cl_khr_subgroups` collectives had two causes: SPIRVIntrinsics declared them without `convergent`, so LLVM duplicated the calls into divergent branches (fixed in JuliaGPU/OpenCL.jl#526), and PoCL 7.2 gives a peeled work-item its own copy of a collective's scratch memory after a branch with an early exit, as bounds checks emit (fixed on PoCL's main branch, backport to 7.2 in pocl/pocl#2373). With both fixed, i.e. with `POCL_WORK_GROUP_METHOD=cbs` for now, the native reductions and scans pass the tests. Keep them behind `NATIVE_COLLECTIVES` until `pocl_standalone_jll` includes the PoCL fix. Assisted-by: Claude Code (Opus 5.5)
This was referenced Oct 4, 2026
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #526 +/- ##
=======================================
Coverage 86.29% 86.29%
=======================================
Files 19 19
Lines 1642 1642
=======================================
Hits 1417 1417
Misses 225 225 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
vchuravy
commented
Oct 4, 2026
The shuffles were declared through `@builtin_ccall` without the `convergent` attribute, so LLVM was free to make them control-dependent on additional values. For `r = sub_group_shuffle(i <= m ? x[i] : 0, 1)`, jump threading duplicated the call into both arms of the branch, so the work-items of a sub-group executed different calls, which is undefined behavior (PoCL returned zero to all of them). Add a `convergent = true` option to `@builtin_ccall`, and use it for the shuffles and the barriers, replacing the barriers' own `convergent_call!`. Also pass the lane of `sub_group_shuffle` and the mask of `sub_group_shuffle_xor` modulo `UInt32`: an out-of-range lane, e.g. of a shuffle-up written as `sub_group_shuffle(x, lane - offset)`, now gives an undefined value as with the OpenCL C built-ins, instead of throwing an `InexactError`. Assisted-by: Claude Code (Opus 5.5)
Add `sub_group_any` and `sub_group_all` (`cl_khr_subgroups`), and
`sub_group_ballot` (`cl_khr_subgroup_ballot`), which returns OpenCL's `uint4`
mask as an `NTuple{4, VecElement{UInt32}}`.
These call the OpenCL C built-ins, which the SPIR-V back-end lowers to
`OpGroupAny`, `OpGroupAll` and `OpGroupNonUniformBallot`, since
`@builtin_ccall` can't mangle the `bool` arguments of the SPIR-V wrapper
built-ins. spirv2clc doesn't implement these instructions, so the tests are
skipped there.
Assisted-by: Claude Code (Opus 5.5)
Add the collectives of `cl_khr_subgroups`, which the SPIR-V back-end lowers to `OpGroup*` instructions, for 32- and 64-bit integers and floats, and Float16: `sub_group_reduce_*`, `sub_group_scan_inclusive_*` and `sub_group_scan_exclusive_*` for `add`, `min` and `max`, and `sub_group_broadcast` (with a 1-based lane, like `sub_group_shuffle`). Like the shuffles, they are called as `convergent`. The test of a bounds-checked reduction of values from a divergent branch needs pocl_jll 7.2.1+1, which includes pocl/pocl#2239. Assisted-by: Claude Code (Opus 5.5)
maleadt
force-pushed
the
vc/subgroup-votes
branch
from
October 4, 2026 18:57
fa10418 to
2e42134
Compare
Member
|
LGTM. Reworked a bit, e.g., applying the |
maleadt
approved these changes
Oct 5, 2026
vchuravy
added a commit
to JuliaGPU/KernelAbstractions.jl
that referenced
this pull request
Oct 6, 2026
Replace the local workarounds with the votes and unchecked shuffle lanes from JuliaGPU/OpenCL.jl#526, taken from its branch until it is released: through `[sources]`, and explicitly where that doesn't apply (Julia 1.10 on CI, and the Buildkite jobs, whose OpenCL job developed SPIRVIntrinsics from OpenCL.jl's ka-0.10 branch). Assisted-by: Claude Code (Opus 5.5)
vchuravy
added a commit
to JuliaGPU/KernelAbstractions.jl
that referenced
this pull request
Oct 6, 2026
Implement `KI.sub_group_reduce` and `KI.sub_group_scan` with the collectives of `cl_khr_subgroups` from SPIRVIntrinsics (JuliaGPU/OpenCL.jl#526): for `+` on 32- and 64-bit integers and floats, and `min`/`max` on integers. Floats keep the fallback for `min` and `max`, as OpenCL treats NaN and the sign of zero differently. Test the operators and types that backends may implement natively, including a NaN, and that POCL uses the native reduction. Assisted-by: Claude Code (Opus 5.5)
vchuravy
added a commit
to JuliaGPU/KernelAbstractions.jl
that referenced
this pull request
Oct 6, 2026
The wrong results of the native `cl_khr_subgroups` collectives had two causes: SPIRVIntrinsics declared them without `convergent`, so LLVM duplicated the calls into divergent branches (fixed in JuliaGPU/OpenCL.jl#526), and PoCL 7.2 gives a peeled work-item its own copy of a collective's scratch memory after a branch with an early exit, as bounds checks emit (fixed on PoCL's main branch, backport to 7.2 in pocl/pocl#2373). With both fixed, i.e. with `POCL_WORK_GROUP_METHOD=cbs` for now, the native reductions and scans pass the tests. Keep them behind `NATIVE_COLLECTIVES` until `pocl_standalone_jll` includes the PoCL fix. Assisted-by: Claude Code (Opus 5.5)
vchuravy
added a commit
to JuliaGPU/KernelAbstractions.jl
that referenced
this pull request
Oct 10, 2026
Replace the local workarounds with the votes and unchecked shuffle lanes from JuliaGPU/OpenCL.jl#526, taken from its branch until it is released: through `[sources]`, and explicitly where that doesn't apply (Julia 1.10 on CI, and the Buildkite jobs, whose OpenCL job developed SPIRVIntrinsics from OpenCL.jl's ka-0.10 branch). Assisted-by: Claude Code (Opus 5.5)
vchuravy
added a commit
to JuliaGPU/KernelAbstractions.jl
that referenced
this pull request
Oct 10, 2026
Implement `KI.sub_group_reduce` and `KI.sub_group_scan` with the collectives of `cl_khr_subgroups` from SPIRVIntrinsics (JuliaGPU/OpenCL.jl#526): for `+` on 32- and 64-bit integers and floats, and `min`/`max` on integers. Floats keep the fallback for `min` and `max`, as OpenCL treats NaN and the sign of zero differently. Test the operators and types that backends may implement natively, including a NaN, and that POCL uses the native reduction. Assisted-by: Claude Code (Opus 5.5)
vchuravy
added a commit
to JuliaGPU/KernelAbstractions.jl
that referenced
this pull request
Oct 10, 2026
The wrong results of the native `cl_khr_subgroups` collectives had two causes: SPIRVIntrinsics declared them without `convergent`, so LLVM duplicated the calls into divergent branches (fixed in JuliaGPU/OpenCL.jl#526), and PoCL 7.2 gives a peeled work-item its own copy of a collective's scratch memory after a branch with an early exit, as bounds checks emit (fixed on PoCL's main branch, backport to 7.2 in pocl/pocl#2373). With both fixed, i.e. with `POCL_WORK_GROUP_METHOD=cbs` for now, the native reductions and scans pass the tests. Keep them behind `NATIVE_COLLECTIVES` until `pocl_standalone_jll` includes the PoCL fix. Assisted-by: Claude Code (Opus 5.5)
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds the sub-group votes and collectives that KernelInterface's sub-group operations need on SPIR-V back-ends (JuliaGPU/KernelAbstractions.jl#831), and fixes a miscompile of the existing sub-group shuffles.
Convergent builtins
The shuffles were declared without the
convergentattribute, so LLVM was free to make them control-dependent on additional values. Forjump threading duplicated the shuffle into both arms of the branch. The work-items of a sub-group then executed different calls, which is undefined behavior; on PoCL every work-item got 0.
@builtin_ccallnow takes aconvergent = trueoption, used for the shuffles, the new votes and collectives, and the barriers (which had their ownconvergent_call!helper before).Shuffle lanes
sub_group_shuffleconverted the lane withUInt32(i - 1), which throws anInexactError(and leaves an error branch in the kernel) for an out-of-range lane, e.g. in a shuffle-up written assub_group_shuffle(x, lane - offset). The lane is now passed moduloUInt32, and so is the mask ofsub_group_shuffle_xor. An out-of-range lane gives an undefined value, as with the OpenCL C built-ins.Votes
sub_group_any(::Bool)andsub_group_all(::Bool)(cl_khr_subgroups), andsub_group_ballot(::Bool)::NTuple{4, VecElement{UInt32}}(cl_khr_subgroup_ballot, OpenCL'suint4mask). These call the OpenCL C built-ins, which the SPIR-V back-end lowers, because@builtin_ccallcan't mangle theboolarguments of__spirv_GroupNonUniformAny/All/Ballot.Collectives
The
cl_khr_subgroupscollectives forInt32,UInt32,Int64,UInt64,Float16,Float32andFloat64:sub_group_reduce_{add,min,max}sub_group_scan_inclusive_{add,min,max}sub_group_scan_exclusive_{add,min,max}sub_group_broadcast(x, lane), with a 1-based lane likesub_group_shuffleThe non-uniform arithmetic of
cl_khr_subgroup_non_uniform_arithmetic(mul,and,or,xor, …) and the char/short types ofcl_khr_subgroup_extended_typesaren't included, because PoCL doesn't support them.Tests
convergent.