Repository navigation
Conversation
Contributor
Benchmark ResultsShow table
Benchmark PlotsA plot of the benchmark results have been uploaded as an artifact to the workflow run for this PR. |
Add the sub-group operations that e.g. Molly.jl's CUDA kernels use, so that they can be written portably: - shuffles `shfl` (from a given lane), `shfl_up` and `shfl_xor`, next to `shfl_down`. Backends implement them for primitive types; a fallback shuffles `isbits` structs and tuples field by field, and `supports_shuffle` checks their fields. - votes `sub_group_any`, `sub_group_all` and `sub_group_ballot` (a `UInt64` mask, for sub-groups of at most 64 work-items), required with sub-group support. - `get_max_sub_group_size` is now required to be a constant of the generated code. Implement them for POCL; its sub-group width is folded into the IR before optimization. Assisted-by: Claude Code (Opus 5.5)
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## vc/ki-subgroup-ops #830 +/- ##
======================================================
+ Coverage 74.15% 75.70% +1.55%
======================================================
Files 24 25 +1
Lines 2275 2437 +162
======================================================
+ Hits 1687 1845 +158
- Misses 588 592 +4 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
vchuravy
force-pushed
the
vc/groupreduce
branch
from
October 3, 2026 22:49
507c83a to
df037a3
Compare
On Julia 1.10, inference gives up on the recursive call of `shfl_fields` through the shuffle of a nested field (e.g. a tuple in a struct), leaving a dynamic invocation in the kernel. Generate the shuffles of all primitive fields directly instead. Assisted-by: Claude Code (Opus 5.5)
vchuravy
force-pushed
the
vc/groupreduce
branch
from
October 4, 2026 06:44
df037a3 to
25c9cb5
Compare
vchuravy
added this pull request to stack #834
October 4, 2026 06:46
Replace the local workarounds with the votes and unchecked shuffle lanes from JuliaGPU/OpenCL.jl#526, taken from its branch until it is released: through `[sources]`, and explicitly where that doesn't apply (Julia 1.10 on CI, and the Buildkite jobs, whose OpenCL job developed SPIRVIntrinsics from OpenCL.jl's ka-0.10 branch). Assisted-by: Claude Code (Opus 5.5)
vchuravy
force-pushed
the
vc/groupreduce
branch
from
October 4, 2026 07:13
25c9cb5 to
930dddb
Compare
…ce and scan Fill the gaps that a survey of the packages using warp operations (KomaMRI, KernelIntrinsics/KernelForge, AcceleratedKernels#93, ParallelStencil, ClimaCore, ...) showed: - Primitive types that a backend doesn't support natively (e.g. `Bool`, `Char`, or 64-bit types on Metal) are shuffled as `UInt32` words, which backends now have to support. Structs keep being shuffled field by field. - Shuffles within segments of `width` lanes (`shfl(val, lane, width)` etc.), with CUDA's semantics, built on `shfl`. - `sub_group_match_any(val)`, the mask of the lanes with the same value, with a fallback built on `shfl` and `sub_group_ballot`. - `sub_group_reduce(op, val)` and `sub_group_scan(op, val)` with fallbacks built on the shuffles, which backends can implement with native operations. - Document how partial sub-groups behave. The new tests are in a function of their own: as part of `interface_testsuite`, compiling the host code crashed LLVM. Assisted-by: Claude Code (Opus 5.5)
vchuravy
force-pushed
the
vc/groupreduce
branch
from
October 4, 2026 08:34
930dddb to
5ff905d
Compare
Implement `KI.sub_group_reduce` and `KI.sub_group_scan` with the collectives of `cl_khr_subgroups` from SPIRVIntrinsics (JuliaGPU/OpenCL.jl#526): for `+` on 32- and 64-bit integers and floats, and `min`/`max` on integers. Floats keep the fallback for `min` and `max`, as OpenCL treats NaN and the sign of zero differently. Test the operators and types that backends may implement natively, including a NaN, and that POCL uses the native reduction. Assisted-by: Claude Code (Opus 5.5)
PoCL's `cl_khr_subgroups` reductions and scans lose the values of work-items that computed them in a divergent branch (PoCL 7.2; Intel's OpenCL runtime is fine), so `@groupreduce` in a `@kernel`, whose padding work-items are masked, returned garbage. Use KernelInterface's fallbacks again, and test reductions and scans of values from a divergent branch. Assisted-by: Claude Code (Opus 5.5)
vchuravy
force-pushed
the
vc/groupreduce
branch
from
October 4, 2026 08:59
5ff905d to
e78926a
Compare
Like the shuffles, the votes, `sub_group_match_any`, `sub_group_reduce` and `sub_group_scan` exchange values, not memory; communicating through memory within a sub-group needs `sub_group_barrier`. Assisted-by: Claude Code (Opus 5.5)
vchuravy
force-pushed
the
vc/groupreduce
branch
from
October 4, 2026 10:19
e78926a to
8ab1401
Compare
The wrong results of the native `cl_khr_subgroups` collectives had two causes: SPIRVIntrinsics declared them without `convergent`, so LLVM duplicated the calls into divergent branches (fixed in JuliaGPU/OpenCL.jl#526), and PoCL 7.2 gives a peeled work-item its own copy of a collective's scratch memory after a branch with an early exit, as bounds checks emit (fixed on PoCL's main branch, backport to 7.2 in pocl/pocl#2373). With both fixed, i.e. with `POCL_WORK_GROUP_METHOD=cbs` for now, the native reductions and scans pass the tests. Keep them behind `NATIVE_COLLECTIVES` until `pocl_standalone_jll` includes the PoCL fix. Assisted-by: Claude Code (Opus 5.5)
vchuravy
force-pushed
the
vc/groupreduce
branch
from
October 4, 2026 10:34
8ab1401 to
81e4230
Compare
Assisted-by: Claude Code (Opus 5.5)
vchuravy
force-pushed
the
vc/groupreduce
branch
from
October 4, 2026 10:43
81e4230 to
06fdf57
Compare
`sub_group_any`, `sub_group_all`, `sub_group_ballot` and `sub_group_match_any` with a `width`, like the shuffles: the votes of segments of `width` lanes, with masks that have a bit per lane of the segment. Fallbacks use the ballot and match of the whole sub-group. With the shuffles with a width, this allows e.g. tiles of 32 lanes on sub-groups of 64. Assisted-by: Claude Code (Opus 5.5)
vchuravy
force-pushed
the
vc/groupreduce
branch
from
October 4, 2026 10:59
06fdf57 to
280c4e1
Compare
… the width - If a work-group is 1-D or its x extent is a multiple of the sub-group width, sub-groups are formed from consecutive work-items, x fastest. This holds on CUDA, AMD, PoCL, rusticl and Intel's CPU OpenCL runtime (which forms sub-groups per row otherwise), and lets kernels with e.g. (32, 8) work-groups rely on the layout. - `shfl_down` and `shfl_up` return the work-item's own value where the source lane is past the sub-group width, like CUDA's shuffles. That's free on CUDA and Metal, and a select on AMD and SPIR-V (POCL), and makes them consistent with the shuffles with a `width`. `shfl_xor` requires a mask below the width. Test both. Assisted-by: Claude Code (Opus 5.5)
Porting KomaMRI.jl showed the fallback of `sub_group_reduce` to cost ~40% of a reduction-heavy kernel on CUDA: its loop ran to the run-time sub-group size, so it didn't unroll, and every lane was masked. `sub_group_reduce` is now an ordered butterfly with `shfl_xor` over the constant sub-group width (combining the lower block first, so it still only needs associativity), which gives every work-item of a full sub-group the result without a broadcast. Partial sub-groups skip the blocks without work-items and broadcast the result of the first lane. All sub-groups run the same shuffles, which PoCL needs; document that requirement of the POCL backend. `sub_group_scan` loops to the constant width as well. Test reductions and scans in a work-group of several sub-groups, the last one partial. Assisted-by: Claude Code (Opus 5.5)
vchuravy
force-pushed
the
vc/groupreduce
branch
from
October 4, 2026 12:00
280c4e1 to
a20329e
Compare
The new fallbacks of `sub_group_reduce` and `sub_group_scan` loop over the constant sub-group width. PoCL 7.2 miscompiles the unrolled shuffles after a branch with an early exit, as bounds-checked `@kernel`s have (the bug fixed by pocl/pocl#2239), so `@groupreduce` with sub-groups gave wrong results. Override them for POCL with a loop bounded by the smaller of the width and the work-group size, which doesn't unroll and is the same for all sub-groups of a work-group. Assisted-by: Claude Code (Opus 5.5)
vchuravy
force-pushed
the
vc/groupreduce
branch
from
October 4, 2026 12:21
a20329e to
5f5c328
Compare
…ork-items Assisted-by: Claude Code (Opus 5.5)
vchuravy
force-pushed
the
vc/groupreduce
branch
from
October 4, 2026 12:31
5f5c328 to
a424d5d
Compare
….2.1+1 pocl_standalone_jll 7.2.1+1 includes the WorkitemLoops fix (pocl/pocl#2239, JuliaPackaging/Yggdrasil#15001), so implement `KI.sub_group_reduce` and `KI.sub_group_scan` with the native `cl_khr_subgroups` collectives where they have Julia's semantics, and only keep the workaround for the fallbacks with older builds. The version is checked at precompile time, since the compat bound can't distinguish builds. Note: PoCL's kernel cache doesn't distinguish 7.2.1+0 and +1, so binaries miscompiled by the former can be reused by the latter until the cache is cleared (pocl/pocl#2374, JuliaPackaging/Yggdrasil#15002). Assisted-by: Claude Code (Opus 5.5)
Rework of #559 on top of KernelInterface: - `@groupreduce(op, val, neutral[, groupsize]; subgroups=false)` reduces over the workgroup and returns the result on every work-item. It uses a local-memory tree by default, or a two-level reduction based on `KI.shfl_down` with `subgroups=true` (gated on the host by `KI.supports_shuffle`). The local memory is sized by the static workgroup size or an explicit upper bound. - `@subgroupreduce(op, val, neutral)` reduces over the sub-group with shuffles; the result is defined on the first lane. - Both are collectives in `@kernel`: the split treats them like `@synchronize`, and padding work-items contribute `neutral` without evaluating `val`, so ndranges that are not a multiple of the workgroup size work. Co-authored-by: Anton Smirnov <tonysmn97@gmail.com> Assisted-by: Claude Code (Opus 5.5)
- `@groupscan(op, val, neutral[, groupsize]; inclusive = true)` scans over the workgroup in the order of the local linear index, with a Hillis-Steele scan in double-buffered local memory, sized like `@groupreduce`. - `@subgroupscan(op, val, neutral; inclusive = true)` scans over the lanes of a sub-group with `KI.shfl_up`. Both only need `op` to be associative, and are collectives in `@kernel` like the reductions: padding work-items contribute `neutral`. Assisted-by: Claude Code (Opus 5.5)
`@subgroupreduce` and `@subgroupscan`, and the sub-group stage of `@groupreduce`, now use `KI.sub_group_reduce` and `KI.sub_group_scan`, which backends can implement with native operations. `@subgroupreduce` returns the result on every work-item of the sub-group. Test reductions of (value, index) pairs, and don't assume that sub-groups are formed from consecutive work-items. Assisted-by: Claude Code (Opus 5.5)
KernelAbstractions' collectives are executed by the padding work-items of a partial workgroup, but direct calls of KernelInterface's sub-group functions aren't: kernels that use them need `unsafe_indices=true`. Assisted-by: Claude Code (Opus 5.5)
The type of the value passed to `@groupreduce`, `@subgroupreduce`,
`@groupscan` or `@subgroupscan` may differ between the work-items: in a
`@kernel` the padding work-items contribute `neutral` instead of `val`, so
`@groupreduce(+, x[i]::Float32, 0.0)` reduces a `Union{Float32, Float64}`,
and an accumulator that only some work-items add a `Float64` to is a `Union`
as well. Julia union-splits the call of the collective with such an argument
into one call per type, so the work-items of a workgroup executed different
copies of its barriers and shuffles. On PoCL this silently gave wrong
results (0 for the reduction of a padded workgroup, NaN for the energy of a
Float32 system with a Float64 Coulomb constant in Molly).
Convert the value to the type of `neutral` at the call site instead, so that
only the conversion is union-split, and test mixed types for all four
collectives.
Assisted-by: Claude Code (Opus 5.5)
vchuravy
force-pushed
the
vc/groupreduce
branch
from
October 4, 2026 16:35
8333aea to
9b20320
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Supersedes #559 by @pxl-th (credited as co-author), reworked on top of KernelInterface.
Stacked on #831:
@subgroupscanneedsKI.shfl_up(and the struct shuffles) from there, so this PR targets that branch for now. Merge #831 first; this one then retargets tomain.API
@groupreduce(op, val, neutral[, groupsize]; subgroups = false)reducesvalover the workgroup and returns the result on every work-item.subgroups = true: each sub-group reduces withKI.shfl_down, then the sub-group results are combined (constant number of barriers). Gate it on the host withKI.supports_shuffle(backend, T)and pass it in as a constant (e.g.::Val{S}).groupsizeas a compile-time upper bound for dynamic workgroup sizes.@subgroupreduce(op, val, neutral)reduces over the sub-group withKI.sub_group_reduce(backends may use native reductions); the result is defined on every lane. It only needsopto be associative.@groupscan(op, val, neutral[, groupsize]; inclusive = true)scans in the order of@index(Local, Linear), inclusive or (withinclusive = false) exclusive. It uses a Hillis-Steele scan in double-buffered local memory, sized like@groupreduce. There is no sub-group variant: how work-items form sub-groups is unspecified in KernelInterface, so a scan built from sub-group scans wouldn't follow the local index order.@subgroupscan(op, val, neutral; inclusive = true)scans over the lanes of a sub-group withKI.sub_group_scan.opto be associative. They are collectives like the reductions, so padding work-items contributeneutral. A use case is stream compaction: offsets from an exclusive@groupscanof the predicates (cf. Molly.jl's neighbor finder).Differences to #559
shfl_down/supports_warp_reductionare gone: they'reKI.shfl_down/KI.supports_shufflenow, and the sub-group width comes from KI instead of a hardcoded 32.@kernelsplit, like@synchronize. Padding work-items take part, contributeneutraland don't evaluateval. This fixes the uninitialized-local-memory issue raised in Implement groupreduce API #559, and is whyneutralis required.y[i] = @groupreduce(...)) is a macro-expansion error, since it would run on padding work-items.Values of different types
The macros convert
valto the type ofneutralat the call site. Otherwise, avalwhose type differs between work-items makes Julia union-split the call, and the work-items run different copies of its barriers and shuffles. Two cases trigger this:neutralhaving a different type thanval(e.g.@groupreduce(+, x[i]::Float32, 0.0));On POCL this silently returned wrong results. It was found through a NaN in the Molly.jl port.
Tests
test/groupreduce.jlruns as part of the backend testsuite and covers:+/maxon Int32, Int64 and Float32unsafe_indices@groupscan(inclusive/exclusive): partial and non-power-of-two workgroups, a dynamic workgroup size with a bound, Cartesian workgroups with padding in the middle, reuse in a loop@subgroupscan, checked against the lane order the kernel reports@groupreduceof (value, index) pairs (argmin)@subgroupreduceon every lane, without assuming that sub-groups are consecutive work-items, and the macro errorsThe full test suite passes locally on POCL.
🤖 Generated with Claude Code