Skip to content

SPIRVIntrinsics: sub-group votes and collectives, unchecked shuffle lanes - #526

Merged
maleadt merged 3 commits into
mainfrom
vc/subgroup-votes
Oct 5, 2026
Merged

maleadt merged 3 commits into
mainfrom
vc/subgroup-votes

Conversation

@vchuravy

@vchuravy vchuravy commented Oct 4, 2026 •

Copy link
Copy Markdown
Member

Adds the sub-group votes and collectives that KernelInterface's sub-group operations need on SPIR-V back-ends (JuliaGPU/KernelAbstractions.jl#831), and fixes a miscompile of the existing sub-group shuffles.

Convergent builtins

The shuffles were declared without the convergent attribute, so LLVM was free to make them control-dependent on additional values. For

x = i <= m ? in[i] : 0
r = sub_group_shuffle(x, 1)

jump threading duplicated the shuffle into both arms of the branch. The work-items of a sub-group then executed different calls, which is undefined behavior; on PoCL every work-item got 0.

@builtin_ccall now takes a convergent = true option, used for the shuffles, the new votes and collectives, and the barriers (which had their own convergent_call! helper before).

Shuffle lanes

sub_group_shuffle converted the lane with UInt32(i - 1), which throws an InexactError (and leaves an error branch in the kernel) for an out-of-range lane, e.g. in a shuffle-up written as sub_group_shuffle(x, lane - offset). The lane is now passed modulo UInt32, and so is the mask of sub_group_shuffle_xor. An out-of-range lane gives an undefined value, as with the OpenCL C built-ins.

Votes

sub_group_any(::Bool) and sub_group_all(::Bool) (cl_khr_subgroups), and sub_group_ballot(::Bool)::NTuple{4, VecElement{UInt32}} (cl_khr_subgroup_ballot, OpenCL's uint4 mask). These call the OpenCL C built-ins, which the SPIR-V back-end lowers, because @builtin_ccall can't mangle the bool arguments of __spirv_GroupNonUniformAny/All/Ballot.

Collectives

The cl_khr_subgroups collectives for Int32, UInt32, Int64, UInt64, Float16, Float32 and Float64:

  • sub_group_reduce_{add,min,max}
  • sub_group_scan_inclusive_{add,min,max}
  • sub_group_scan_exclusive_{add,min,max}
  • sub_group_broadcast(x, lane), with a 1-based lane like sub_group_shuffle

The non-uniform arithmetic of cl_khr_subgroup_non_uniform_arithmetic (mul, and, or, xor, …) and the char/short types of cl_khr_subgroup_extended_types aren't included, because PoCL doesn't support them.

Tests

vchuravy added a commit to JuliaGPU/KernelAbstractions.jl that referenced this pull request Oct 4, 2026
Replace the local workarounds with the votes and unchecked shuffle lanes from
JuliaGPU/OpenCL.jl#526, taken from its branch until it is released: through
`[sources]`, and explicitly where that doesn't apply (Julia 1.10 on CI, and the
Buildkite jobs, whose OpenCL job developed SPIRVIntrinsics from OpenCL.jl's
ka-0.10 branch).

Assisted-by: Claude Code (Opus 5.5)
@vchuravy vchuravy changed the title SPIRVIntrinsics: sub-group votes, unchecked shuffle lanes SPIRVIntrinsics: sub-group votes and collectives, unchecked shuffle lanes Oct 4, 2026
vchuravy added a commit to JuliaGPU/KernelAbstractions.jl that referenced this pull request Oct 4, 2026
Implement `KI.sub_group_reduce` and `KI.sub_group_scan` with the collectives of
`cl_khr_subgroups` from SPIRVIntrinsics (JuliaGPU/OpenCL.jl#526): for `+` on
32- and 64-bit integers and floats, and `min`/`max` on integers. Floats keep the
fallback for `min` and `max`, as OpenCL treats NaN and the sign of zero
differently.

Test the operators and types that backends may implement natively, including a
NaN, and that POCL uses the native reduction.

Assisted-by: Claude Code (Opus 5.5)
vchuravy added a commit to JuliaGPU/KernelAbstractions.jl that referenced this pull request Oct 4, 2026
The wrong results of the native `cl_khr_subgroups` collectives had two causes:
SPIRVIntrinsics declared them without `convergent`, so LLVM duplicated the
calls into divergent branches (fixed in JuliaGPU/OpenCL.jl#526), and PoCL 7.2
gives a peeled work-item its own copy of a collective's scratch memory after a
branch with an early exit, as bounds checks emit (fixed on PoCL's main branch,
backport to 7.2 in pocl/pocl#2373).

With both fixed, i.e. with `POCL_WORK_GROUP_METHOD=cbs` for now, the native
reductions and scans pass the tests. Keep them behind `NATIVE_COLLECTIVES`
until `pocl_standalone_jll` includes the PoCL fix.

Assisted-by: Claude Code (Opus 5.5)
@codecov

codecov Bot commented Oct 4, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 86.29%. Comparing base (b6b2429) to head (2e42134).

Additional details and impacted files
@@           Coverage Diff           @@
##             main     #526   +/-   ##
=======================================
  Coverage   86.29%   86.29%           
=======================================
  Files          19       19           
  Lines        1642     1642           
=======================================
  Hits         1417     1417           
  Misses        225      225           

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

Comment thread test/intrinsics.jl Outdated
The shuffles were declared through `@builtin_ccall` without the `convergent`
attribute, so LLVM was free to make them control-dependent on additional
values. For `r = sub_group_shuffle(i <= m ? x[i] : 0, 1)`, jump threading
duplicated the call into both arms of the branch, so the work-items of a
sub-group executed different calls, which is undefined behavior (PoCL returned
zero to all of them).

Add a `convergent = true` option to `@builtin_ccall`, and use it for the
shuffles and the barriers, replacing the barriers' own `convergent_call!`.

Also pass the lane of `sub_group_shuffle` and the mask of
`sub_group_shuffle_xor` modulo `UInt32`: an out-of-range lane, e.g. of a
shuffle-up written as `sub_group_shuffle(x, lane - offset)`, now gives an
undefined value as with the OpenCL C built-ins, instead of throwing an
`InexactError`.

Assisted-by: Claude Code (Opus 5.5)
Add `sub_group_any` and `sub_group_all` (`cl_khr_subgroups`), and
`sub_group_ballot` (`cl_khr_subgroup_ballot`), which returns OpenCL's `uint4`
mask as an `NTuple{4, VecElement{UInt32}}`.

These call the OpenCL C built-ins, which the SPIR-V back-end lowers to
`OpGroupAny`, `OpGroupAll` and `OpGroupNonUniformBallot`, since
`@builtin_ccall` can't mangle the `bool` arguments of the SPIR-V wrapper
built-ins. spirv2clc doesn't implement these instructions, so the tests are
skipped there.

Assisted-by: Claude Code (Opus 5.5)
Add the collectives of `cl_khr_subgroups`, which the SPIR-V back-end lowers to
`OpGroup*` instructions, for 32- and 64-bit integers and floats, and Float16:
`sub_group_reduce_*`, `sub_group_scan_inclusive_*` and
`sub_group_scan_exclusive_*` for `add`, `min` and `max`, and
`sub_group_broadcast` (with a 1-based lane, like `sub_group_shuffle`). Like the
shuffles, they are called as `convergent`.

The test of a bounds-checked reduction of values from a divergent branch needs
pocl_jll 7.2.1+1, which includes pocl/pocl#2239.

Assisted-by: Claude Code (Opus 5.5)
@maleadt
maleadt force-pushed the vc/subgroup-votes branch from fa10418 to 2e42134 Compare October 4, 2026 18:57
@maleadt

maleadt commented Oct 4, 2026

Copy link
Copy Markdown
Member

LGTM. Reworked a bit, e.g., applying the convergent fix to other intrinsics and testing using Float16, but kept the fundamentals.

@christiangnrd christiangnrd left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@maleadt
maleadt merged commit 9e8999d into main Oct 5, 2026
16 checks passed
@maleadt
maleadt deleted the vc/subgroup-votes branch October 5, 2026 05:14
vchuravy added a commit to JuliaGPU/KernelAbstractions.jl that referenced this pull request Oct 6, 2026
Replace the local workarounds with the votes and unchecked shuffle lanes from
JuliaGPU/OpenCL.jl#526, taken from its branch until it is released: through
`[sources]`, and explicitly where that doesn't apply (Julia 1.10 on CI, and the
Buildkite jobs, whose OpenCL job developed SPIRVIntrinsics from OpenCL.jl's
ka-0.10 branch).

Assisted-by: Claude Code (Opus 5.5)
vchuravy added a commit to JuliaGPU/KernelAbstractions.jl that referenced this pull request Oct 6, 2026
Implement `KI.sub_group_reduce` and `KI.sub_group_scan` with the collectives of
`cl_khr_subgroups` from SPIRVIntrinsics (JuliaGPU/OpenCL.jl#526): for `+` on
32- and 64-bit integers and floats, and `min`/`max` on integers. Floats keep the
fallback for `min` and `max`, as OpenCL treats NaN and the sign of zero
differently.

Test the operators and types that backends may implement natively, including a
NaN, and that POCL uses the native reduction.

Assisted-by: Claude Code (Opus 5.5)
vchuravy added a commit to JuliaGPU/KernelAbstractions.jl that referenced this pull request Oct 6, 2026
The wrong results of the native `cl_khr_subgroups` collectives had two causes:
SPIRVIntrinsics declared them without `convergent`, so LLVM duplicated the
calls into divergent branches (fixed in JuliaGPU/OpenCL.jl#526), and PoCL 7.2
gives a peeled work-item its own copy of a collective's scratch memory after a
branch with an early exit, as bounds checks emit (fixed on PoCL's main branch,
backport to 7.2 in pocl/pocl#2373).

With both fixed, i.e. with `POCL_WORK_GROUP_METHOD=cbs` for now, the native
reductions and scans pass the tests. Keep them behind `NATIVE_COLLECTIVES`
until `pocl_standalone_jll` includes the PoCL fix.

Assisted-by: Claude Code (Opus 5.5)
vchuravy added a commit to JuliaGPU/KernelAbstractions.jl that referenced this pull request Oct 10, 2026
Replace the local workarounds with the votes and unchecked shuffle lanes from
JuliaGPU/OpenCL.jl#526, taken from its branch until it is released: through
`[sources]`, and explicitly where that doesn't apply (Julia 1.10 on CI, and the
Buildkite jobs, whose OpenCL job developed SPIRVIntrinsics from OpenCL.jl's
ka-0.10 branch).

Assisted-by: Claude Code (Opus 5.5)
vchuravy added a commit to JuliaGPU/KernelAbstractions.jl that referenced this pull request Oct 10, 2026
Implement `KI.sub_group_reduce` and `KI.sub_group_scan` with the collectives of
`cl_khr_subgroups` from SPIRVIntrinsics (JuliaGPU/OpenCL.jl#526): for `+` on
32- and 64-bit integers and floats, and `min`/`max` on integers. Floats keep the
fallback for `min` and `max`, as OpenCL treats NaN and the sign of zero
differently.

Test the operators and types that backends may implement natively, including a
NaN, and that POCL uses the native reduction.

Assisted-by: Claude Code (Opus 5.5)
vchuravy added a commit to JuliaGPU/KernelAbstractions.jl that referenced this pull request Oct 10, 2026
The wrong results of the native `cl_khr_subgroups` collectives had two causes:
SPIRVIntrinsics declared them without `convergent`, so LLVM duplicated the
calls into divergent branches (fixed in JuliaGPU/OpenCL.jl#526), and PoCL 7.2
gives a peeled work-item its own copy of a collective's scratch memory after a
branch with an early exit, as bounds checks emit (fixed on PoCL's main branch,
backport to 7.2 in pocl/pocl#2373).

With both fixed, i.e. with `POCL_WORK_GROUP_METHOD=cbs` for now, the native
reductions and scans pass the tests. Keep them behind `NATIVE_COLLECTIVES`
until `pocl_standalone_jll` includes the PoCL fix.

Assisted-by: Claude Code (Opus 5.5)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants