Skip to content

KernelInterface: sub-group shuffles, votes and a constant width - #3331

Draft
vchuravy wants to merge 8 commits into
ka-0.10from
vc/ki-subgroup-ops
Draft

vchuravy wants to merge 8 commits into
ka-0.10from
vc/ki-subgroup-ops

Conversation

@vchuravy

@vchuravy vchuravy commented Oct 4, 2026 •

Copy link
Copy Markdown
Member

Implements the sub-group communication of KernelInterface's updated contract from JuliaGPU/KernelAbstractions.jl#831 in the CUDA back-end (CUDACore/src/CUDAKernels.jl).

What's implemented

  • Shuffles KI.shfl, KI.shfl_down, KI.shfl_up, KI.shfl_xor, all with the full warp mask, like the existing shfl_down override.
    • The overrides, and KI.supports_shuffle, now only match const ShuffleTypes = Union{Bool, Base.BitInteger, Base.IEEEFloat}. These are the primitive types CUDA.jl's warp shuffles handle, splitting the wider ones into 32-bit shuffles. Other isbits types, Complex included, go to KI's generic fallback, which shuffles them field by field. Before this PR, shfl_down and supports_shuffle were generic where {T} methods.
    • Lanes and offsets are truncated with % UInt32 rather than converted, so an out-of-range value returns an unspecified value instead of throwing an InexactError. shfl_sync takes a 1-based lane and subtracts 1, so KI.shfl wraps the lane into 1:32 first, which is the same wrapping PTX applies to the 0-based lane.
  • Votes KI.sub_group_any, KI.sub_group_all, KI.sub_group_ballot. Ballot widens the UInt32 from vote_ballot_sync to UInt64. Bit i - 1 is set for lane i, since laneid() is 1-based.
  • Constant width: the warp size is now the constant 32i32 (WARP_SIZE) instead of a read of %WARP_SZ. KI.get_max_sub_group_size(T) returns 32 % T, a compile-time constant of the generated code as the contract requires. The other sub-group queries (get_num_sub_groups, get_sub_group_id, the partial-warp size) use the same constant. Every NVIDIA GPU has a warp size of 32, and CUDA.jl's warp intrinsics already hard-code it (ws = Int32(32) in warp.jl). The host-side KI.sub_group_size(::CUDABackend) still queries warpsize(device()), which returns 32.
KernelInterface CUDA.jl
shfl(val, lane) shfl_sync(FULL_MASK, val, lane′) with lane′ = ((lane - 1) % UInt32 & 0x1f) + 1
shfl_down(val, offset) shfl_down_sync(FULL_MASK, val, offset′), offset′ = offset < 32 ? offset % UInt32 : 0
shfl_up(val, offset) shfl_up_sync(FULL_MASK, val, offset′)
shfl(val, lane, width), shfl_down/up(val, offset, width) shfl_sync/shfl_down_sync/shfl_up_sync(FULL_MASK, val, lane′/offset′, width)
shfl_xor(val, mask, width) shfl_xor_sync(FULL_MASK, val, mask < width ? mask : 0, width)
sub_group_reduce(op, val::Union{Int32,UInt32}), op in + min max & | ⊻ redux.sync with active_mask() on sm_80+, KI's fallback otherwise
shfl_xor(val, mask) shfl_xor_sync(FULL_MASK, val, mask % UInt32)
sub_group_any(pred) vote_any_sync(FULL_MASK, pred)
sub_group_all(pred) vote_all_sync(FULL_MASK, pred)
sub_group_ballot(pred) UInt64(vote_ballot_sync(FULL_MASK, pred))
get_max_sub_group_size(T) 32i32 % T
supports_shuffle(::CUDABackend, ::Type{<:ShuffleTypes}) true (composites go to KI's fallback)

Update for the contract changes in JuliaGPU/KernelAbstractions.jl#831

  • Shuffles past the warp: KI now requires shfl_down/shfl_up to return the thread's own value where the source lane is past the sub-group width, like shfl.sync. But shfl.sync only uses the low 5 bits of the offset (checked on the GPU: shfl_down_sync with offset 33 shifts by 1). So offsets of 32 or more become 0. This is folded for constant offsets, and a setp/selp otherwise. shfl_xor is unchanged, since KI requires mask < 32.
  • Shuffles with a width: native overrides, a single shfl.sync with the segment mask in c (e.g. shfl.sync.down.b32 %r, %v, 2, 4127, -1), instead of KI's fallback (shfl.idx plus integer ops). PTX takes the lane modulo width and returns the own value past the segment end. For shfl_xor, PTX does read from a lower segment, while KI wants the own value there. So a mask >= width becomes 0.
  • Cheaper sub-group queries: get_sub_group_id/get_sub_group_size/get_num_sub_groups were not inlined (separate calls in the PTX) and used signed division. They are now @inline and use UInt32 math with shifts and a Horner-form linear thread index. get_sub_group_local_id is laneid() and get_max_sub_group_size is a constant (both unchanged).
  • redux.sync: KI.sub_group_reduce for +, min, max, &, |, ⊻ on Int32/UInt32 uses llvm.nvvm.redux.sync.* with active_mask() when compiled for sm_80+ (compute_capability() folds). Otherwise it invokes KI's fallback. There's no native scan.

Tested on the Quadro RTX 4000 (sm_75) against JuliaGPU/KernelAbstractions.jl#831 at 26c51b4: core/kernelinterface + core/kernelabstractions: 5363 pass, 17 broken, SUCCESS. That includes the new layout tests and the own-value checks. A manual check compared shfl_down/shfl_up (offsets up to typemax(Int32)) and all four width variants (widths 1 to 32) against KI's reference semantics, with no mismatches. I also checked partial last warps at linear ids ≥ 256. redux.sync is only compile-checked (CUDA.code_ptx(...; arch=v"8.0") gives activemask.b32 + redux.sync.add.s32). There is no sm_80+ GPU here to run it, and on sm_75 the test suite exercises the fallback branch.

[TEMP] commit

The second commit, [TEMP] Get KernelAbstractions and KernelInterface from the vc/ki-subgroup-ops branch, changes the [sources] revs in CUDACore, CUDATools, lib/cusparse and test, and the Buildkite clone for Julia 1.10/1.11, from KA's main to vc/ki-subgroup-ops. This makes CI test against #831. Drop it, or squash it into the existing [TEMP] commit, once #831 is merged.

The CUDA back-end doesn't need the SPIRVIntrinsics source that KA's branch uses for POCL. Without it, the test environment resolves the registered SPIRVIntrinsics v1.2.0, and KA precompiles and loads fine.

Heads-up, a problem ka-0.10 already has: KA's main and #831 both require LLVM = "10", but ka-0.10 still has LLVM = "9.6" in CUDACore and "9.3.1" in CUDATools. main has since moved to LLVM.jl 10 (#3323). So ka-0.10 doesn't resolve against either KA branch, and CI on this PR will fail to resolve until ka-0.10 is rebased onto main. This PR doesn't fix that.

Local testing

I tested on a Quadro RTX 4000 (sm_75) with Julia 1.12.7. I used a throwaway local branch where ka-0.10, plus these commits, was rebased onto main; the only conflicts were the LLVM and version compat entries. It resolved KA/KI 0.10.0-dev/0.4.0-dev from #vc/ki-subgroup-ops (7f09a6c), LLVM v10.0.0 and GPUCompiler v2.11.1.

julia --project=test test/runtests.jl core/kernelabstractions core/kernelinterface

Result: Overall | 3612 pass, 17 broken, 3629 total, SUCCESS. The 17 broken tests are existing @test_brokens. This includes KI's testsuite from #831: the new shuffles (including structs), the votes and the constant width.

I also ran these checks by hand:

  • KI.shfl with lanes 0, -5, 33, typemax(Int), shfl_down(Int8, typemax(Int)) and shfl_up(UInt16, 40) don't throw.
  • The PTX for KI.get_max_sub_group_size() doesn't read %WARP_SZ.
  • KI.shfl on ComplexF64 goes through the KI fallback and gives the correct result.
  • supports_shuffle returns true for ComplexF64 and false for Char.

🤖 Generated with Claude Code

maleadt and others added 7 commits September 30, 2026 19:33
KernelAbstractions 0.10 builds on KernelInterface, which defines what a back
end provides: memory and device management, compiling and launching kernels,
and the device-side intrinsics. KernelAbstractions then launches `@kernel`
kernels itself on any KernelInterface back end, which replaces CUDA's copy of
that launch path: partitioning the ndrange, building the kernel's context, and
tuning the workgroup size.

`CUDABackend` now implements KernelInterface in CUDACore, which depends on it
instead of on KernelAbstractions. What KernelAbstractions still needs from a
back end moves to an extension: the `MArray` behind `@private`, the Adapt rule
for `@Const`, and the `maxthreads` hint for a static workgroup size.
`prefer_blocks` now applies in `KI.launch_configuration`, which receives the
number of work-items.

`KI.launch` passes the kernel arguments on as a tuple, so kernels with many
arguments stay cheap to launch, as #3309 made them for the old launch path. It
rejects `threads` and `blocks`, which would override the launch geometry that
KernelInterface validated. `KI.copyto!` accepts dense arrays and contiguous
views of them, as the old `KA.copyto!` did.

KernelAbstractions converts the arguments twice, to determine the types to
compile for and again when launching, where CUDA's launch converted them once.
`cudaconvert` has to be pure, so this only shows with conversions that have
side effects, as in two tests that count them.

Co-authored-by: Christian Guinard <28689358+christiangnrd@users.noreply.github.com>
Sub-groups are warps: implement the sub-group queries, `sub_group_barrier` and
`shfl_down`, and report their support to KernelInterface. KernelInterface
leaves unspecified how work-items are grouped into sub-groups; CUDA forms warps
from consecutive linear thread indices, so the last warp of a block can be
partial, which a test checks.

Co-authored-by: Christian Guinard <28689358+christiangnrd@users.noreply.github.com>
Implement `KI.record_event` and `KI.wait_event` with a `CuEvent` recorded on,
and waited for by, the task's stream. `KernelAbstractions.@spawn` uses them to
order a new task's work after the work its parent had queued, without
synchronizing the parent; the default is a full synchronization.

KernelInterface's testsuite checks them once `record_event` returns an event.
`KI.versioninfo(CUDABackend())` prints `CUDA.versioninfo()`.
… main branch

Neither is registered yet. Julia 1.10 and 1.11 don't pick up the test
project's [sources], so Buildkite develops both explicitly there, and the
GPU-less and Enzyme jobs move to Julia 1.12.

Drop this commit once both are registered.
Implement the sub-group communication of KernelInterface's updated
contract: `shfl`, `shfl_up` and `shfl_xor` next to `shfl_down`, and the
votes `sub_group_any`, `sub_group_all` and `sub_group_ballot`, using the
warp intrinsics with the full mask.

The shuffles, and `supports_shuffle`, are restricted to the primitive
types CUDA's warp shuffles support, so that KernelInterface shuffles
structs and tuples (including `Complex`) field by field. Lanes and
offsets are truncated rather than converted, so that out-of-range values
give an unspecified value instead of throwing.

The warp size is now the constant 32 rather than a read of `%WARP_SZ`,
so that `get_max_sub_group_size` is a compile-time constant of the
generated code, as KernelInterface requires.

Assisted-by: Claude Code (Opus 5.5)
…roup-ops branch

Test against JuliaGPU/KernelAbstractions.jl#831. Drop this commit (going
back to KernelAbstractions' main branch) once that is merged.

Assisted-by: Claude Code (Opus 5.5)
…oup queries, redux

- `shfl.sync` only uses the low 5 bits of the offset, so `shfl_down` and
  `shfl_up` with an offset of 32 or more read from lane `offset % 32`
  rather than returning the thread's own value, as KernelInterface now
  requires. Use an offset of 0 for those (folded for constant offsets).
- Implement the shuffles with a `width` as a single `shfl.sync`, rather
  than KernelInterface's fallback (a `shfl.sync` from a computed lane).
  `shfl_xor` with mask bits at or above the width returns the own value,
  as KernelInterface requires (PTX would read from a lower segment).
- Compute the sub-group queries with unsigned 32-bit integers (shifts
  rather than signed divisions, a Horner-form linear thread index), and
  inline them; before, `get_sub_group_id` and `get_sub_group_size` were
  calls.
- `sub_group_reduce` of `+`, `min`, `max`, `&`, `|` and `⊻` on
  `Int32`/`UInt32` uses `redux.sync` on sm_80 and later, KernelInterface's
  fallback otherwise.

Assisted-by: Claude Code (Opus 5.5)

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants