Repository navigation
Add foreach_index, a parallel loop over indices without a kernel - #827
Merged
Merged
Conversation
A kernel launched without a workgroup size partitions its ndrange with the ndrange itself as a preliminary workgroup size, which is not a valid workgroup size when the ndrange is a range or a CartesianIndices. Use its extents instead, as `Int`s also for a range of unsigned integers, whose length the tuning would otherwise not accept.
A kernel whose body is a loop over the indices of an array needs none of the kernel language beyond the index itself, and writing it out is boilerplate that downstream packages repeat. Add `foreach_index(f, itr, backend = get_backend(itr); workgroupsize)`, which launches one work item per index of `eachindex(itr)` and calls `f` with the index a `for i in eachindex(itr)` loop would produce: a linear index for an `IndexLinear` array, a `CartesianIndex` otherwise. The index space is carried by the `ndrange`, so the kernel takes no argument besides the function, and an index space that does not start at 1 needs no special handling. The function is inlined into the kernel, as an out-of-line call to a closure can spill its captures to local memory. Ported from AcceleratedKernels.jl, without its CPU scheduling keywords (`max_tasks`, `min_elems`, `prefer_threads`): those pick a host-threaded loop, which 0.10 no longer has, and the workgroup size is left to the backend rather than defaulting to 256. Spelled with an underscore to keep the name distinct from the AcceleratedKernels function it does not behave identically to, so that importing both packages does not clash. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DbbtGLjxXpt7kZj6h2yVdK
`foreach_index(f, itr, backend)` called `f` with `eachindex(itr)`, which is what one wants for an array, but for an index space rebases it: `foreach_index(f, 5:10, backend)` called `f(1:6)`, and the interior `CartesianIndices((2:n-1, 2:m-1))` of an array got rebased to start at 1. Split the two: `foreach_index(f, A)` loops over `eachindex(A)` on the backend of `A`, and `foreach_index(f, backend, indices)` over the given range or `CartesianIndices`, as an `ndrange` does. Other index spaces, including stepped ranges, are rejected also when they're empty.
A loop body usually indexes more than one array with the same index. `foreach_index(f, A, Bs...)` loops over `eachindex(A, Bs...)`, as a `for` loop would, so that the index is valid for all of them, and Cartesian if any of them needs it. The arrays must share a backend.
A zero-dimensional `CartesianIndices` has a single index, but as an `ndrange` it cannot be combined with a workgroup size, which has at least one dimension. Launch it as a single work item instead, so that a `workgroupsize` works for every index space.
A variable that is reassigned after a closure captures it is stored in a `Core.Box`, which a kernel cannot access, and compiling the kernel fails with an error about non-isbits kernel arguments that does not mention the variable. Check for boxed captures before launching, which is free for other closures, and name the variable in the error.
Metal has no 64-bit atomic addition.
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #827 +/- ##
===========================================
- Coverage 78.82% 67.67% -11.16%
===========================================
Files 24 25 +1
Lines 2021 2521 +500
===========================================
+ Hits 1593 1706 +113
- Misses 428 815 +387 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
Contributor
Benchmark ResultsShow table
Benchmark PlotsA plot of the benchmark results have been uploaded as an artifact to the workflow run for this PR. |
maleadt
marked this pull request as ready for review
October 3, 2026 18:38
vchuravy
force-pushed
the
tb/foreach_index
branch
from
October 4, 2026 10:06
18e2cc2 to
99223b5
Compare
vchuravy
reviewed
Oct 5, 2026
`foreach_index(f, y, 1:n)` asked every array for its backend, and a range has none, so it failed with an error asking to implement `get_backend` for it, although a range can be indexed on any backend. Like AcceleratedKernels, let ranges, `CartesianIndices` and `LinearIndices`, and Base's views, reshapes and permutations of them, run on the backend of the other arrays. Unlike it, do not fall back to the host if there are no other arrays: ask for the backend instead.
maleadt
force-pushed
the
tb/foreach_index
branch
from
October 5, 2026 15:50
99223b5 to
7f08f18
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Supersedes #779. That PR is based on #772, which may not land, while it only needs the range
ndranges that are already onmain(#771). This branch keeps its commit, rebases it ontomain,and builds on it, mainly to change the semantics for index spaces; see
Changes compared to #779 below.
Many kernels are just a loop body over the indices of some arrays.
foreach_indexlets you writethose as an ordinary function, without the kernel language:
fgets the index afor i in eachindex(y, x)loop would give: a linear index if all arrayssupport it, a
CartesianIndexotherwise. It runs on the backend of the arrays, which must all bethe same.
To loop over something other than the indices of an array, pass a backend and an index space. The
indices are used as they are, offsets included, the same way an
ndrangeworks:Like every other launch it's asynchronous and returns
nothing. Iterations run in no particularorder. Bounds checks stay on unless the body uses
@inbounds.Why
The function comes from AcceleratedKernels.jl, where
foreachindexis the primitive undermap,reduce,sort,accumulateand the others. Moving it into KA gives users the convenient formwithout pulling in AK, and lets AK use KA's version for its GPU path instead of maintaining its
own kernel.
I checked that last point by swapping AK's GPU kernel for
KA.foreach_index(f, backend, indices; workgroupsize=block_size)and running AK's test suite onCUDA: everything passes. Two bodies had to add
@inbounds, since they relied on AK's kernel beinginbounds=true. AK keeps its own publicforeachindex, with its host-threads path and schedulingkeywords. It can only depend on this once it requires KA 0.10.
Changes compared to #779
An index-space form with an explicit backend. #779 ran over
eachindex(itr)withget_backend(itr)as default. For arrays that's right, but for index spaces it isn't:foreach_index(f, 5:10, backend)calledf(1)throughf(6), and a stencil interiorCartesianIndices((2:n-1, 2:m-1))got rebased to start at 1. AK needs the given indices for itsown call sites (e.g. a
CartesianIndicesover offset axes inaccumulate, andaxes(x, d)). Thetwo meanings are now two methods:
foreach_index(f, A, Bs...): the indices of these arrays, on their backend.foreach_index(f, backend, indices): these indices, on this backend.A range or
CartesianIndiceshas no backend, soforeach_index(f, 5:10)is an error, not asilent loop over
1:6. Only unit ranges of integers andCartesianIndicesof those are accepted.Anything else, including stepped ranges, throws an
ArgumentError, also when it's empty.Most frameworks I looked at that accept a begin and end pass the actual indices: Kokkos'
RangePolicy/MDRangePolicy, RAJA segments, Taichi'sndrange, JACC. SYCL, CUB and Warp onlysupport
[0, N).Several arrays. A loop body almost always indexes more than one array, so the array form takes
several, like
eachindex(A, B...). That makes the index valid for all of them, and Cartesian ifany of them needs it.
Zero-dimensional index spaces run as a single work item. An explicit
workgroupsizefailed anassertion otherwise, and AK always passes one.
Boxed captures get a useful error. Capturing a variable that is reassigned (a common mistake
for people new to GPU programming) used to fail with
KernelError: passing non-bitstype argumentand a page of kernel types. Now it says which variable it is and how to fix it. The check
constant-folds away for other closures, also on Julia 1.10.
Smaller things.
I[1]instead ofI.I[1]. A docstring that leads with the asynchronous,unordered semantics, and notes that on the CPU backend each new closure means a kernel compile.
The tests use 32-bit atomics, because Metal has no 64-bit atomic add.
I didn't take the review suggestion to use
@index(Global, Linear)in the linear kernel: that'sthe position in the ndrange, which loses the offset of a range that doesn't start at 1.
A launch fix
The first commit fixes a bug on
mainthat #779 ran into. A kernel launched with a range orCartesianIndicesas itsndrange(#771) and no workgroup size threwMethodError: normalize_workgroupsize(::OneTo): the launch uses the ndrange as a preliminaryworkgroup size before tuning. The #771 tests always passed a workgroup size, so they missed this.
The launch now uses the extents, as
Intalso for unsigned ranges.Testing
The
foreach_indexand offset tests pass on CPU (PoCL), CUDA, AMDGPU (a gfx1036 iGPU), OpenCL(PoCL and NVIDIA), Metal (M1) and oneAPI (Iris Xe). The full suite passes on CPU, and KA's full
testsuite on CUDA. The GPU back-ends were tested with their
ka-0.10branches merged with theirLLVM.jl 10 ports.
Performance matches a hand-written
@kerneleverywhere, since it is one. Median time per launchplus synchronization, for
y[i] = 2x[i] + 1:@kernel@kernel; broadcast 6.7 ms@kernelOn strided views, launching over the
CartesianIndicesdirectly beats AK's linear launch with anindex decode on Metal and AMDGPU, and matches it elsewhere. The one place AK's kernel came out
ahead is large 1-D arrays on the AMD iGPU (6.15 vs 6.42 ms), which comes from its fixed workgroup
size of 256 versus KA's tuned one. A hand-written
@kernelsees the same difference, and AKpasses its own workgroup size anyway.
On the CPU backend it also matches a
@kernel. Each new closure costs about 150 ms to compile, andbelow roughly a million elements
Threads.@threadsis faster (6.5 vs 20 µs for 64k elements),which is why the docstring points to threaded loops for one-off or small work.
Open questions
foreach_indexavoids clashing with AK's exportedforeachindex, which hasdifferent semantics (synchronous host threads, scheduling keywords).
foreachindexwould readmore like
eachindex.