Repository navigation
Conversation
Benchmark ResultsShow table
Benchmark PlotsA plot of the benchmark results have been uploaded as an artifact to the workflow run for this PR. |
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## main #831 +/- ##
==========================================
- Coverage 79.06% 74.15% -4.92%
==========================================
Files 24 24
Lines 2040 2275 +235
==========================================
+ Hits 1613 1687 +74
- Misses 427 588 +161 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
| # the sub-group width is fixed, so make `get_max_sub_group_size` a constant, as | ||
| # KernelInterface requires (this runs before optimization) | ||
| gvs = LLVM.globals(mod) | ||
| if haskey(gvs, "__spirv_BuiltInSubgroupMaxSize") | ||
| gv = gvs["__spirv_BuiltInSubgroupMaxSize"] | ||
| for use in collect(LLVM.uses(gv)) | ||
| load = LLVM.user(use) | ||
| load isa LLVM.LoadInst || continue | ||
| LLVM.replace_uses!(load, ConstantInt(LLVM.value_type(load), sg_size)) | ||
| LLVM.erase!(load) | ||
| end | ||
| isempty(LLVM.uses(gv)) && LLVM.erase!(gv) | ||
| end |
There was a problem hiding this comment.
This is perhaps a bit sketchy, and would need to be replicated for OpenCL/oneAPI
There was a problem hiding this comment.
Eh, I'm not sure I like this. Isn't the driver responsible for specializing on this? It looks like other drivers do this, so this might be a PoCL deficiency if anything.
There was a problem hiding this comment.
Alternative: pocl/pocl#2375
So I'd prefer if we leave this out for now, and I'll work with upstream to get this polished and backported.
There was a problem hiding this comment.
Claude:
PoCL does specialize: its workgroup pass replaces _pocl_sub_group_size with the intel_reqd_sub_group_size value, so the final kernel has a constant either way. It does so late, though, after it has formed the barrier regions (on the CPU device every shuffle is a work-group barrier) and only right before its final O3. So the shuffle loops in the KI reduce/scan fallbacks stay as barrier loops around work-item loops and are never unrolled. Without a constant on the Julia side, the testsuite passes but those fallbacks are ~1.2–1.7x slower. The other backends get their constant on the Julia side too: AMDGPU.jl folds llvm.amdgcn.wavefrontsize in finish_module!, and the CUDA/Metal KI PRs hardcode 32, since NVPTX on LLVM 18 doesn't fold %warpsize in the middle end.
There was a problem hiding this comment.
In addition, my PR also folds get_local_size exposing further optimizations.
KernelInterface (JuliaGPU/KernelAbstractions.jl#831) now shuffles primitive types that a back-end doesn't support natively as `UInt32` words, so the Metal-specific split into halves isn't needed anymore. Assisted-by: Claude Code (Opus 5.5)
Add the sub-group operations that e.g. Molly.jl's CUDA kernels use, so that they can be written portably: - shuffles `shfl` (from a given lane), `shfl_up` and `shfl_xor`, next to `shfl_down`. Backends implement them for primitive types; a fallback shuffles `isbits` structs and tuples field by field, and `supports_shuffle` checks their fields. - votes `sub_group_any`, `sub_group_all` and `sub_group_ballot` (a `UInt64` mask, for sub-groups of at most 64 work-items), required with sub-group support. - `get_max_sub_group_size` is now required to be a constant of the generated code. Implement them for POCL; its sub-group width is folded into the IR before optimization. Assisted-by: Claude Code (Opus 5.5)
On Julia 1.10, inference gives up on the recursive call of `shfl_fields` through the shuffle of a nested field (e.g. a tuple in a struct), leaving a dynamic invocation in the kernel. Generate the shuffles of all primitive fields directly instead. Assisted-by: Claude Code (Opus 5.5)
Replace the local workarounds with the votes and unchecked shuffle lanes from JuliaGPU/OpenCL.jl#526, taken from its branch until it is released: through `[sources]`, and explicitly where that doesn't apply (Julia 1.10 on CI, and the Buildkite jobs, whose OpenCL job developed SPIRVIntrinsics from OpenCL.jl's ka-0.10 branch). Assisted-by: Claude Code (Opus 5.5)
…ce and scan Fill the gaps that a survey of the packages using warp operations (KomaMRI, KernelIntrinsics/KernelForge, AcceleratedKernels#93, ParallelStencil, ClimaCore, ...) showed: - Primitive types that a backend doesn't support natively (e.g. `Bool`, `Char`, or 64-bit types on Metal) are shuffled as `UInt32` words, which backends now have to support. Structs keep being shuffled field by field. - Shuffles within segments of `width` lanes (`shfl(val, lane, width)` etc.), with CUDA's semantics, built on `shfl`. - `sub_group_match_any(val)`, the mask of the lanes with the same value, with a fallback built on `shfl` and `sub_group_ballot`. - `sub_group_reduce(op, val)` and `sub_group_scan(op, val)` with fallbacks built on the shuffles, which backends can implement with native operations. - Document how partial sub-groups behave. The new tests are in a function of their own: as part of `interface_testsuite`, compiling the host code crashed LLVM. Assisted-by: Claude Code (Opus 5.5)
Implement `KI.sub_group_reduce` and `KI.sub_group_scan` with the collectives of `cl_khr_subgroups` from SPIRVIntrinsics (JuliaGPU/OpenCL.jl#526): for `+` on 32- and 64-bit integers and floats, and `min`/`max` on integers. Floats keep the fallback for `min` and `max`, as OpenCL treats NaN and the sign of zero differently. Test the operators and types that backends may implement natively, including a NaN, and that POCL uses the native reduction. Assisted-by: Claude Code (Opus 5.5)
PoCL's `cl_khr_subgroups` reductions and scans lose the values of work-items that computed them in a divergent branch (PoCL 7.2; Intel's OpenCL runtime is fine), so `@groupreduce` in a `@kernel`, whose padding work-items are masked, returned garbage. Use KernelInterface's fallbacks again, and test reductions and scans of values from a divergent branch. Assisted-by: Claude Code (Opus 5.5)
Like the shuffles, the votes, `sub_group_match_any`, `sub_group_reduce` and `sub_group_scan` exchange values, not memory; communicating through memory within a sub-group needs `sub_group_barrier`. Assisted-by: Claude Code (Opus 5.5)
The wrong results of the native `cl_khr_subgroups` collectives had two causes: SPIRVIntrinsics declared them without `convergent`, so LLVM duplicated the calls into divergent branches (fixed in JuliaGPU/OpenCL.jl#526), and PoCL 7.2 gives a peeled work-item its own copy of a collective's scratch memory after a branch with an early exit, as bounds checks emit (fixed on PoCL's main branch, backport to 7.2 in pocl/pocl#2373). With both fixed, i.e. with `POCL_WORK_GROUP_METHOD=cbs` for now, the native reductions and scans pass the tests. Keep them behind `NATIVE_COLLECTIVES` until `pocl_standalone_jll` includes the PoCL fix. Assisted-by: Claude Code (Opus 5.5)
Assisted-by: Claude Code (Opus 5.5)
`sub_group_any`, `sub_group_all`, `sub_group_ballot` and `sub_group_match_any` with a `width`, like the shuffles: the votes of segments of `width` lanes, with masks that have a bit per lane of the segment. Fallbacks use the ballot and match of the whole sub-group. With the shuffles with a width, this allows e.g. tiles of 32 lanes on sub-groups of 64. Assisted-by: Claude Code (Opus 5.5)
… the width - If a work-group is 1-D or its x extent is a multiple of the sub-group width, sub-groups are formed from consecutive work-items, x fastest. This holds on CUDA, AMD, PoCL, rusticl and Intel's CPU OpenCL runtime (which forms sub-groups per row otherwise), and lets kernels with e.g. (32, 8) work-groups rely on the layout. - `shfl_down` and `shfl_up` return the work-item's own value where the source lane is past the sub-group width, like CUDA's shuffles. That's free on CUDA and Metal, and a select on AMD and SPIR-V (POCL), and makes them consistent with the shuffles with a `width`. `shfl_xor` requires a mask below the width. Test both. Assisted-by: Claude Code (Opus 5.5)
Porting KomaMRI.jl showed the fallback of `sub_group_reduce` to cost ~40% of a reduction-heavy kernel on CUDA: its loop ran to the run-time sub-group size, so it didn't unroll, and every lane was masked. `sub_group_reduce` is now an ordered butterfly with `shfl_xor` over the constant sub-group width (combining the lower block first, so it still only needs associativity), which gives every work-item of a full sub-group the result without a broadcast. Partial sub-groups skip the blocks without work-items and broadcast the result of the first lane. All sub-groups run the same shuffles, which PoCL needs; document that requirement of the POCL backend. `sub_group_scan` loops to the constant width as well. Test reductions and scans in a work-group of several sub-groups, the last one partial. Assisted-by: Claude Code (Opus 5.5)
The new fallbacks of `sub_group_reduce` and `sub_group_scan` loop over the constant sub-group width. PoCL 7.2 miscompiles the unrolled shuffles after a branch with an early exit, as bounds-checked `@kernel`s have (the bug fixed by pocl/pocl#2239), so `@groupreduce` with sub-groups gave wrong results. Override them for POCL with a loop bounded by the smaller of the width and the work-group size, which doesn't unroll and is the same for all sub-groups of a work-group. Assisted-by: Claude Code (Opus 5.5)
…ork-items Assisted-by: Claude Code (Opus 5.5)
….2.1+1 pocl_standalone_jll 7.2.1+1 includes the WorkitemLoops fix (pocl/pocl#2239, JuliaPackaging/Yggdrasil#15001), so implement `KI.sub_group_reduce` and `KI.sub_group_scan` with the native `cl_khr_subgroups` collectives where they have Julia's semantics, and only keep the workaround for the fallbacks with older builds. The version is checked at precompile time, since the compat bound can't distinguish builds. Note: PoCL's kernel cache doesn't distinguish 7.2.1+0 and +1, so binaries miscompiled by the former can be reused by the latter until the cache is cleared (pocl/pocl#2374, JuliaPackaging/Yggdrasil#15002). Assisted-by: Claude Code (Opus 5.5)
- Require the released SPIRVIntrinsics 1.3 instead of the deleted `vc/subgroup-votes` branch of OpenCL.jl, in `[sources]` and in CI. The OpenCL.jl job installs 1.3 rather than the ka-0.10 copy (1.1.4). - The fallback of `KI.sub_group_match_any` takes a step per distinct value, which differs between the sub-groups of a work-group. PoCL needs them to take the same steps, so override it with a loop to a bound that is the same for the whole work-group. Test it with several sub-groups. - Merge `NATIVE_COLLECTIVES` into `POCL_REPLICA_FIX`. - `supports_shuffle` of a type without fields follows `UInt32` instead of being `true` on every backend. - Fix the sizes of primitive types in the `shfl` docstring and the `supports_shuffle` entry of the backend table. Assisted-by: Claude Code
pocl_standalone_jll 7.2.2, which main requires now, includes the fix for PoCL 7.2's peeling of the first work-item (pocl/pocl#2239), so the native reductions and scans are always used, and the loops to a uniform bound that replaced them without the fix can't be reached anymore. Assisted-by: Claude Code
Primitive types that a backend doesn't shuffle natively were always split into UInt32 words, so a `Ptr` or another 64-bit primitive type took two shuffles even on backends with native 64-bit shuffles. Shuffle them as a `UInt64` instead, which a backend without native 64-bit shuffles splits into two words as before, and 16-byte values as two `UInt64` words. Check the generated code on PoCL. Assisted-by: Claude Code
c3e0b7a to
64ac668
Compare
|
Rebased onto main and addressed the review:
The full test suite passes locally on Julia 1.10 and 1.12. Still open from the review: Float16/Float64 on PoCL devices without fp16/fp64 are reported as unsupported for shuffles, even though the #830 is based on this branch and needs a rebase. |
| # not the ka-0.10 copy of SPIRVIntrinsics: KernelAbstractions needs the | ||
| # sub-group votes and collectives of SPIRVIntrinsics 1.3 | ||
| Pkg.add(name="SPIRVIntrinsics", version="1.3")' || exit 3 | ||
|
|
There was a problem hiding this comment.
SPIRVIntrinsics 1.3 has been released
| # Shuffles exchange values between the work-items of a sub-group. Backends implement them for | ||
| # the primitive types for which `supports_shuffle` returns `true`. The fallbacks below shuffle | ||
| # other primitive types as unsigned words, and other `isbits` types field by field. | ||
|
|
There was a problem hiding this comment.
This is in the docstring
| # Shuffles exchange values between the work-items of a sub-group. Backends implement them for | |
| # the primitive types for which `supports_shuffle` returns `true`. The fallbacks below shuffle | |
| # other primitive types as unsigned words, and other `isbits` types field by field. |
| [`get_sub_group_local_id`](@ref) equal to `get_sub_group_local_id() + offset`, for an `offset` | ||
| of at least 0. When that lane is past the sub-group width, i.e. | ||
| `get_sub_group_local_id() + offset > get_max_sub_group_size()`, the result is `val` of the | ||
| work-item itself, like CUDA's `shfl_down_sync`. When the lane is within the width but has no |
There was a problem hiding this comment.
Is there a performance impact to ensuring this on backends that don't make that promise, and if so, is it worth ensuring?
| for n in unique((sg_size, max(sg_size - 3, 1))) | ||
| @testset "sub_group_reduce and sub_group_scan, $n work-items" begin |
There was a problem hiding this comment.
I find this makes the output much easier to parse when things fail
| for n in unique((sg_size, max(sg_size - 3, 1))) | |
| @testset "sub_group_reduce and sub_group_scan, $n work-items" begin | |
| @testset "sub_group_reduce and sub_group_scan" begin | |
| @testset "$n work-items" for n in unique((sg_size, max(sg_size - 3, 1))) |
| return | ||
| end | ||
|
|
||
| function subgroup_communication_testsuite(backend::KI.Backend, AT, sg_size) |
There was a problem hiding this comment.
Needs to actually be added to the testsuite in testsuite.jl for backends to test them.
There was a problem hiding this comment.
Interface.jl should probably get split up I made this when the interface was much smaller but that's out of scope for this PR
| @device_override KI.shfl_xor(val::T, mask::Integer) where {T <: ShuffleTypes} = | ||
| sub_group_shuffle_xor(val, mask) | ||
|
|
||
| @device_override KI.sub_group_any(pred::Bool) = SPIRVIntrinsics.sub_group_any(pred) | ||
|
|
||
| @device_override KI.sub_group_all(pred::Bool) = SPIRVIntrinsics.sub_group_all(pred) |
There was a problem hiding this comment.
Any reason why only sub_group_shuffle_xor isn't qualified?
| # Native reductions and scans of `cl_khr_subgroups`, for `+` on 32- and 64-bit integers and | ||
| # floats, and `min`/`max` on integers (OpenCL's `min` and `max` treat NaN and the sign of zero | ||
| # differently from Julia's). They need the fix of PoCL 7.2's peeling of the first work-item | ||
| # (pocl/pocl#2239), which `pocl_standalone_jll` includes since 7.2.1+1. |
There was a problem hiding this comment.
Do we have thorough tests for the reductions and scans behaviour for the commonly implemented operators? Julia sometimes treats edge cases differently and I assume we want to keep Julia's semantics so we might have to stick to the fallback even when backends havenative implementations
Adds the sub-group primitives that Molly.jl's CUDA kernels use (
ext/MollyCUDAExt.jl), so that they can be written portably on top of KernelInterface. This is a companion to #830 (@groupreduce/@subgroupreduce).New device functions
shfl(val, lane)supports_shuffle(backend, T)shfl_up(val, offset)supports_shuffle(backend, T)shfl_xor(val, mask)supports_shuffle(backend, T)min/max/|) of tile maskssub_group_any(pred)/sub_group_all(pred)supports_subgroupsvote_any_sync)sub_group_ballot(pred)::UInt64supports_subgroups, width ≤ 64count_ones/trailing_zeroson the resultisbitsstructs and tuples field by field (e.g.SVector, Unitful quantities). For such types,supports_shufflechecks their fields. Molly currently gets this by overriding CUDA.jl's internalshfl_recurse.supports_shuffleand the backend implementation notes now cover all four shuffles.Warp size
get_max_sub_group_size()is now required to be a compile-time constant of the generated code. Loops over lanes and shuffle butterflies get specialized for it, which every backend can provide since kernels are already compiled for a fixed width (sub_group_size(backend)). For type-level decisions (e.g. aUInt32vsUInt64mask) the docs point to passingsub_group_size(backend)from the host.POCL
sub_group_shuffle/sub_group_shuffle_xor. Out-of-range lanes give an unspecified value instead of anInexactError.sub_group_any/sub_group_all/sub_group_ballot(cl_khr_subgroups,cl_khr_subgroup_ballot).finish_module!replaces loads of__spirv_BuiltInSubgroupMaxSizewith the width the kernel is compiled for (intel_reqd_sub_group_size), before optimization.Depends on JuliaGPU/OpenCL.jl#526, which adds the votes and unchecked shuffle lanes to SPIRVIntrinsics. Until it's released, this PR takes SPIRVIntrinsics from that branch:
[sources]inProject.toml;[sources];ka-0.10branch.The branch has the LLVM 10 upgrade that
ka-0.10lacks. Before merging, these should be replaced by a compat bound on the release; all of them are markedTODO.Built on the above (fallbacks, backends may override)
Added after a survey of packages that use warp operations: KomaMRI, KernelIntrinsics/KernelForge, AcceleratedKernels#93, ParallelStencil, ClimaCore, IntervalMDP, …
Bool,Char, 64-bit types on Metal, …)UInt32words. Backends now have to supportUInt32natively.shfl(val, lane, width),shfl_down/up/xor(val, x, width)widthlanes with CUDA's semantics (reads from outside the segment return the own value), built onshflsub_group_match_any(val)::UInt64===), found group by group withshfl+sub_group_ballot. CUDA could usematch.any.syncsub_group_any/all(pred, width),sub_group_ballot(pred, width),sub_group_match_any(val, width)widthlanes (masks have a bit per lane of the segment), built on the ballot/match of the whole sub-group. Together with the shuffles with awidth, this lets e.g. 32-lane tiles run on 64-wide sub-groupssub_group_reduce(op, val)shfl_down, then broadcast withshfl. Backends can dispatch ontypeof(op)for native reductions (Metalsimd_sum, SPIR-VGroupNonUniformIAdd, CUDAredux.sync)sub_group_scan(op, val)shfl_upThe docs now also say how partial sub-groups behave: lanes without a work-item give unspecified shuffle values, and the votes, match, reduce and scan only take the existing work-items into account.
The new tests are in small functions of their own,
subgroup_communication_testsuiteand helpers, with concrete loops. An earlier version iterated over heterogeneous tuples of functions and passed the resulting union-of-singletons value throughKI.@launch, whoseGC.@preservethen hit JuliaLang/julia#63482: on 1.12 and 1.13, codegen emits a nullgc_preserve_beginoperand, and LLVM segfaults. That is fixed on master by #63483, but the backport to 1.12/1.13 is still pending.Guarantees added after porting Molly, KomaMRI and ParallelStencil
Follows from a cross-backend investigation with experiments on CUDA, AMD (wave32 and wave64), PoCL, rusticl and Intel's CPU OpenCL runtime.
W, sub-groups are formed from consecutive work-items, x fastest: sub-group(lin - 1) ÷ W + 1, lane(lin - 1) % W + 1.(32, 8)work-groups (ParallelStencil's default) rely on the layout.shfl_down/shfl_uppast the sub-group width return the work-item's own value, as CUDA's do. That makes them consistent with the shuffles with awidth.shfl_xorrequiresmask < W.sub_group_reducefallback. It is now an orderedshfl_xorbutterfly over the constant width, so it unrolls and needs no broadcast in full sub-groups. KomaMRI measured the previous fallback at ~40% of a reduction-heavy kernel on CUDA. All sub-groups run the same shuffles.sub_group_scanalso loops over the constant width.POCLBackend.pocl_standalone_jllincludes the PoCL fix ([pocl] Backport upstream PR #2239 JuliaPackaging/Yggdrasil#15001), POCL overrides the reduce/scan fallbacks with a loop it compiles correctly in bounds-checked kernels.For backend packages
KernelInterface 0.4 is still unreleased, so this extends its contract. CUDA.jl, AMDGPU.jl, oneAPI.jl, Metal.jl and OpenCL.jl will need to:
shfl,shfl_up,shfl_xor,sub_group_any,sub_group_allandsub_group_ballot;shfl_downoverrides to the primitive types, so that struct shuffles reach the fallback;get_max_sub_group_size(e.g.32 % Ton CUDA).One open question:
sub_group_ballotis required for widths ≤ 64. If a backend can't support ballot, it could get its own capability query instead.Tests
shfl,shfl_uplanes and ashfl_xorbutterfly all-reduce for every supported typesupports_shufflefor structsany/all/ballotwith several predicate patternsget_max_sub_group_sizereturns the width🤖 Generated with Claude Code