Skip to content

Add foreach_index, a parallel loop over indices without a kernel - #827

Merged
maleadt merged 8 commits into
mainfrom
tb/foreach_index
Oct 5, 2026
Merged

maleadt merged 8 commits into
mainfrom
tb/foreach_index

Conversation

@maleadt

@maleadt maleadt commented Oct 3, 2026

Copy link
Copy Markdown
Member

Supersedes #779. That PR is based on #772, which may not land, while it only needs the range
ndranges that are already on main (#771). This branch keeps its commit, rebases it onto main,
and builds on it, mainly to change the semantics for index spaces; see
Changes compared to #779 below.

Many kernels are just a loop body over the indices of some arrays. foreach_index lets you write
those as an ordinary function, without the kernel language:

function scale!(y, x)
    foreach_index(y, x) do i
        @inbounds y[i] = 2 * x[i] + 1
    end
    return y
end

f gets the index a for i in eachindex(y, x) loop would give: a linear index if all arrays
support it, a CartesianIndex otherwise. It runs on the backend of the arrays, which must all be
the same.

To loop over something other than the indices of an array, pass a backend and an index space. The
indices are used as they are, offsets included, the same way an ndrange works:

foreach_index(get_backend(A), CartesianIndices((2:n-1, 2:m-1))) do I
    # the interior of A
end

Like every other launch it's asynchronous and returns nothing. Iterations run in no particular
order. Bounds checks stay on unless the body uses @inbounds.

Why

The function comes from AcceleratedKernels.jl, where foreachindex is the primitive under map,
reduce, sort, accumulate and the others. Moving it into KA gives users the convenient form
without pulling in AK, and lets AK use KA's version for its GPU path instead of maintaining its
own kernel.

I checked that last point by swapping AK's GPU kernel for
KA.foreach_index(f, backend, indices; workgroupsize=block_size) and running AK's test suite on
CUDA: everything passes. Two bodies had to add @inbounds, since they relied on AK's kernel being
inbounds=true. AK keeps its own public foreachindex, with its host-threads path and scheduling
keywords. It can only depend on this once it requires KA 0.10.

Changes compared to #779

An index-space form with an explicit backend. #779 ran over eachindex(itr) with
get_backend(itr) as default. For arrays that's right, but for index spaces it isn't:
foreach_index(f, 5:10, backend) called f(1) through f(6), and a stencil interior
CartesianIndices((2:n-1, 2:m-1)) got rebased to start at 1. AK needs the given indices for its
own call sites (e.g. a CartesianIndices over offset axes in accumulate, and axes(x, d)). The
two meanings are now two methods:

  • foreach_index(f, A, Bs...): the indices of these arrays, on their backend.
  • foreach_index(f, backend, indices): these indices, on this backend.

A range or CartesianIndices has no backend, so foreach_index(f, 5:10) is an error, not a
silent loop over 1:6. Only unit ranges of integers and CartesianIndices of those are accepted.
Anything else, including stepped ranges, throws an ArgumentError, also when it's empty.

Most frameworks I looked at that accept a begin and end pass the actual indices: Kokkos'
RangePolicy/MDRangePolicy, RAJA segments, Taichi's ndrange, JACC. SYCL, CUB and Warp only
support [0, N).

Several arrays. A loop body almost always indexes more than one array, so the array form takes
several, like eachindex(A, B...). That makes the index valid for all of them, and Cartesian if
any of them needs it.

Zero-dimensional index spaces run as a single work item. An explicit workgroupsize failed an
assertion otherwise, and AK always passes one.

Boxed captures get a useful error. Capturing a variable that is reassigned (a common mistake
for people new to GPU programming) used to fail with KernelError: passing non-bitstype argument
and a page of kernel types. Now it says which variable it is and how to fix it. The check
constant-folds away for other closures, also on Julia 1.10.

Smaller things. I[1] instead of I.I[1]. A docstring that leads with the asynchronous,
unordered semantics, and notes that on the CPU backend each new closure means a kernel compile.
The tests use 32-bit atomics, because Metal has no 64-bit atomic add.

I didn't take the review suggestion to use @index(Global, Linear) in the linear kernel: that's
the position in the ndrange, which loses the offset of a range that doesn't start at 1.

A launch fix

The first commit fixes a bug on main that #779 ran into. A kernel launched with a range or
CartesianIndices as its ndrange (#771) and no workgroup size threw
MethodError: normalize_workgroupsize(::OneTo): the launch uses the ndrange as a preliminary
workgroup size before tuning. The #771 tests always passed a workgroup size, so they missed this.
The launch now uses the extents, as Int also for unsigned ranges.

Testing

The foreach_index and offset tests pass on CPU (PoCL), CUDA, AMDGPU (a gfx1036 iGPU), OpenCL
(PoCL and NVIDIA), Metal (M1) and oneAPI (Iris Xe). The full suite passes on CPU, and KA's full
testsuite on CUDA. The GPU back-ends were tested with their ka-0.10 branches merged with their
LLVM.jl 10 ports.

Performance matches a hand-written @kernel everywhere, since it is one. Median time per launch
plus synchronization, for y[i] = 2x[i] + 1:

2^24 elements strided 2-D view
CUDA (RTX 5080) 165 µs, same as broadcast and @kernel 530 µs (4096²), same as broadcast
AMDGPU (iGPU) 6.4 ms, same as @kernel; broadcast 6.7 ms 2.4 ms (2048²); broadcast 3.3 ms
Metal (M1) 2.38 ms, same as broadcast and @kernel 1.74 ms (2048²); AK's kernel 2.67 ms
oneAPI (Iris Xe) 7.2 ms, all variants within noise 4.9 ms (2048²), all within noise

On strided views, launching over the CartesianIndices directly beats AK's linear launch with an
index decode on Metal and AMDGPU, and matches it elsewhere. The one place AK's kernel came out
ahead is large 1-D arrays on the AMD iGPU (6.15 vs 6.42 ms), which comes from its fixed workgroup
size of 256 versus KA's tuned one. A hand-written @kernel sees the same difference, and AK
passes its own workgroup size anyway.

On the CPU backend it also matches a @kernel. Each new closure costs about 150 ms to compile, and
below roughly a million elements Threads.@threads is faster (6.5 vs 20 µs for 64k elements),
which is why the docstring points to threaded loops for one-off or small work.

Open questions

  • The name. foreach_index avoids clashing with AK's exported foreachindex, which has
    different semantics (synchronous host threads, scheduling keywords). foreachindex would read
    more like eachindex.

maleadt and others added 7 commits October 3, 2026 17:45
A kernel launched without a workgroup size partitions its ndrange with the ndrange itself as a
preliminary workgroup size, which is not a valid workgroup size when the ndrange is a range or a
CartesianIndices. Use its extents instead, as `Int`s also for a range of unsigned integers,
whose length the tuning would otherwise not accept.
A kernel whose body is a loop over the indices of an array needs none of the
kernel language beyond the index itself, and writing it out is boilerplate that
downstream packages repeat.

Add `foreach_index(f, itr, backend = get_backend(itr); workgroupsize)`, which
launches one work item per index of `eachindex(itr)` and calls `f` with the
index a `for i in eachindex(itr)` loop would produce: a linear index for an
`IndexLinear` array, a `CartesianIndex` otherwise. The index space is carried
by the `ndrange`, so the kernel takes no argument besides the function, and an
index space that does not start at 1 needs no special handling. The function is
inlined into the kernel, as an out-of-line call to a closure can spill its
captures to local memory.

Ported from AcceleratedKernels.jl, without its CPU scheduling keywords
(`max_tasks`, `min_elems`, `prefer_threads`): those pick a host-threaded loop,
which 0.10 no longer has, and the workgroup size is left to the backend rather
than defaulting to 256. Spelled with an underscore to keep the name distinct
from the AcceleratedKernels function it does not behave identically to, so that
importing both packages does not clash.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DbbtGLjxXpt7kZj6h2yVdK
`foreach_index(f, itr, backend)` called `f` with `eachindex(itr)`, which is what one wants for
an array, but for an index space rebases it: `foreach_index(f, 5:10, backend)` called `f(1:6)`,
and the interior `CartesianIndices((2:n-1, 2:m-1))` of an array got rebased to start at 1.

Split the two: `foreach_index(f, A)` loops over `eachindex(A)` on the backend of `A`, and
`foreach_index(f, backend, indices)` over the given range or `CartesianIndices`, as an `ndrange`
does. Other index spaces, including stepped ranges, are rejected also when they're empty.
A loop body usually indexes more than one array with the same index. `foreach_index(f, A, Bs...)`
loops over `eachindex(A, Bs...)`, as a `for` loop would, so that the index is valid for all of
them, and Cartesian if any of them needs it. The arrays must share a backend.
A zero-dimensional `CartesianIndices` has a single index, but as an `ndrange` it cannot be
combined with a workgroup size, which has at least one dimension. Launch it as a single work item
instead, so that a `workgroupsize` works for every index space.
A variable that is reassigned after a closure captures it is stored in a `Core.Box`, which a
kernel cannot access, and compiling the kernel fails with an error about non-isbits kernel
arguments that does not mention the variable. Check for boxed captures before launching, which
is free for other closures, and name the variable in the error.
Metal has no 64-bit atomic addition.
@codecov

codecov Bot commented Oct 3, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 67.67%. Comparing base (1213466) to head (7f08f18).
⚠️ Report is 17 commits behind head on main.

Additional details and impacted files
@@             Coverage Diff             @@
##             main     #827       +/-   ##
===========================================
- Coverage   78.82%   67.67%   -11.16%     
===========================================
  Files          24       25        +1     
  Lines        2021     2521      +500     
===========================================
+ Hits         1593     1706      +113     
- Misses        428      815      +387     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@github-actions

github-actions Bot commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

Benchmark Results

Show table
main 99223b5... main / 99223b5...
const/@Const/Float32/262144 0.261 ± 0.012 ms 0.261 ± 0.015 ms 0.996 ± 0.073
const/@Const/Float32/65536 0.0997 ± 0.011 ms 0.0969 ± 0.0076 ms 1.03 ± 0.14
const/@Const/Float64/262144 0.436 ± 0.017 ms 0.444 ± 0.017 ms 0.982 ± 0.053
const/@Const/Float64/65536 0.2 ± 0.012 ms 0.213 ± 0.011 ms 0.94 ± 0.076
const/unmarked/Float32/262144 0.476 ± 0.019 ms 0.471 ± 0.018 ms 1.01 ± 0.056
const/unmarked/Float32/65536 0.153 ± 0.01 ms 0.162 ± 0.011 ms 0.945 ± 0.091
const/unmarked/Float64/262144 0.985 ± 0.13 ms 0.888 ± 0.087 ms 1.11 ± 0.18
const/unmarked/Float64/65536 0.268 ± 0.017 ms 0.248 ± 0.014 ms 1.08 ± 0.091
launch/3D static workgroup, dynamic ndrange 15.1 ± 3 μs 13.6 ± 2.7 μs 1.11 ± 0.31
launch/3D static workgroup, static ndrange 15.5 ± 0.96 μs 13.1 ± 2.4 μs 1.19 ± 0.23
launch/dynamic workgroup, dynamic ndrange 13.8 ± 3.5 μs 13 ± 1.3 μs 1.06 ± 0.29
launch/dynamic workgroup, dynamic ndrange, workgroupsize given 16.2 ± 3.6 μs 15.2 ± 2.7 μs 1.07 ± 0.3
launch/static workgroup, dynamic ndrange 15.7 ± 3.4 μs 12.8 ± 0.73 μs 1.23 ± 0.28
launch/static workgroup, static ndrange 14.9 ± 3.8 μs 14.7 ± 3 μs 1.01 ± 0.33
partition/dynamic workgroup, dynamic ndrange 0.0454 ± 0.0098 μs 0.0382 ± 0.0089 μs 1.19 ± 0.38
partition/static workgroup, dynamic ndrange 0.0476 ± 0.0094 μs 0.0398 ± 0.0078 μs 1.2 ± 0.33
partition/static workgroup, static ndrange 1.13 ± 0.004 ns 0.966 ± 0.16 ns 1.17 ± 0.19
saxpy/default/Float16/1024 15.9 ± 1.5 μs 17 ± 1.6 μs 0.936 ± 0.12
saxpy/default/Float16/1048576 0.453 ± 0.053 ms 0.426 ± 0.046 ms 1.06 ± 0.17
saxpy/default/Float16/16384 0.0567 ± 0.0044 ms 0.0529 ± 0.012 ms 1.07 ± 0.26
saxpy/default/Float16/2048 15.9 ± 2.4 μs 14.5 ± 1.7 μs 1.09 ± 0.21
saxpy/default/Float16/256 0.0509 ± 0.036 ms 15.6 ± 1.5 μs 3.27 ± 2.3
saxpy/default/Float16/262144 0.148 ± 0.025 ms 0.143 ± 0.026 ms 1.03 ± 0.26
saxpy/default/Float16/32768 0.0713 ± 0.016 ms 0.0615 ± 0.0086 ms 1.16 ± 0.31
saxpy/default/Float16/4096 14.8 ± 2.1 μs 0.0509 ± 0.035 ms 0.291 ± 0.21
saxpy/default/Float16/512 14.3 ± 1.5 μs 16.1 ± 1.7 μs 0.887 ± 0.13
saxpy/default/Float16/64 17.9 ± 1.9 μs 14.6 ± 35 μs 1.22 ± 2.9
saxpy/default/Float16/65536 0.0712 ± 0.011 ms 0.0718 ± 0.023 ms 0.992 ± 0.35
saxpy/default/Float32/1024 16 ± 2.2 μs 16.1 ± 1.3 μs 0.99 ± 0.16
saxpy/default/Float32/1048576 0.66 ± 0.047 ms 0.56 ± 0.1 ms 1.18 ± 0.23
saxpy/default/Float32/16384 0.0542 ± 0.043 ms 0.0414 ± 0.042 ms 1.31 ± 1.7
saxpy/default/Float32/2048 16.5 ± 2.4 μs 13.8 ± 1.6 μs 1.19 ± 0.22
saxpy/default/Float32/256 16.9 ± 2.4 μs 13.9 ± 1.5 μs 1.21 ± 0.22
saxpy/default/Float32/262144 0.211 ± 0.032 ms 0.193 ± 0.025 ms 1.09 ± 0.22
saxpy/default/Float32/32768 0.0632 ± 0.0061 ms 0.0646 ± 0.0065 ms 0.978 ± 0.14
saxpy/default/Float32/4096 15.2 ± 1.1 μs 14.1 ± 2.4 μs 1.08 ± 0.2
saxpy/default/Float32/512 15.3 ± 2.7 μs 15.2 ± 2.7 μs 1.01 ± 0.26
saxpy/default/Float32/64 14.4 ± 1.9 μs 16.4 ± 2.6 μs 0.877 ± 0.18
saxpy/default/Float32/65536 0.0803 ± 0.019 ms 0.0791 ± 0.018 ms 1.02 ± 0.33
saxpy/default/Float64/1024 0.052 ± 0.027 ms 17 ± 24 μs 3.05 ± 4.6
saxpy/default/Float64/1048576 1.28 ± 0.081 ms 1.17 ± 0.11 ms 1.09 ± 0.12
saxpy/default/Float64/16384 0.0562 ± 0.0098 ms 0.0602 ± 0.0096 ms 0.934 ± 0.22
saxpy/default/Float64/2048 15.5 ± 36 μs 16.4 ± 5.8 μs 0.944 ± 2.2
saxpy/default/Float64/256 15.8 ± 36 μs 14.2 ± 1.5 μs 1.11 ± 2.6
saxpy/default/Float64/262144 0.327 ± 0.062 ms 0.299 ± 0.04 ms 1.09 ± 0.26
saxpy/default/Float64/32768 0.0812 ± 0.074 ms 0.0778 ± 0.016 ms 1.04 ± 0.98
saxpy/default/Float64/4096 0.0441 ± 0.039 ms 14.4 ± 1.4 μs 3.07 ± 2.7
saxpy/default/Float64/512 15.5 ± 2.6 μs 16.6 ± 31 μs 0.93 ± 1.7
saxpy/default/Float64/64 17.8 ± 9.7 μs 14.8 ± 1.6 μs 1.21 ± 0.67
saxpy/default/Float64/65536 0.12 ± 0.017 ms 0.116 ± 0.025 ms 1.03 ± 0.27
saxpy/static workgroup=(1024,)/Float16/1024 14.4 ± 2.2 μs 13.4 ± 1.1 μs 1.07 ± 0.19
saxpy/static workgroup=(1024,)/Float16/1048576 0.46 ± 0.047 ms 0.42 ± 0.048 ms 1.1 ± 0.17
saxpy/static workgroup=(1024,)/Float16/16384 0.0538 ± 0.011 ms 0.0567 ± 0.0058 ms 0.95 ± 0.22
saxpy/static workgroup=(1024,)/Float16/2048 14.7 ± 3.8 μs 17.1 ± 2 μs 0.858 ± 0.25
saxpy/static workgroup=(1024,)/Float16/256 16.3 ± 4.3 μs 17.1 ± 1.5 μs 0.954 ± 0.26
saxpy/static workgroup=(1024,)/Float16/262144 0.163 ± 0.19 ms 0.146 ± 0.026 ms 1.12 ± 1.3
saxpy/static workgroup=(1024,)/Float16/32768 0.0645 ± 0.012 ms 0.0588 ± 0.012 ms 1.1 ± 0.31
saxpy/static workgroup=(1024,)/Float16/4096 14.8 ± 1.8 μs 16.6 ± 1.6 μs 0.895 ± 0.14
saxpy/static workgroup=(1024,)/Float16/512 16.1 ± 2.7 μs 17.9 ± 2.7 μs 0.898 ± 0.2
saxpy/static workgroup=(1024,)/Float16/64 16.5 ± 1.8 μs 15.8 ± 2 μs 1.04 ± 0.17
saxpy/static workgroup=(1024,)/Float16/65536 0.0744 ± 0.014 ms 0.0724 ± 0.0087 ms 1.03 ± 0.23
saxpy/static workgroup=(1024,)/Float32/1024 13.6 ± 1.7 μs 15 ± 3.1 μs 0.909 ± 0.22
saxpy/static workgroup=(1024,)/Float32/1048576 0.614 ± 0.076 ms 0.572 ± 0.1 ms 1.07 ± 0.23
saxpy/static workgroup=(1024,)/Float32/16384 0.0515 ± 0.043 ms 0.0528 ± 0.041 ms 0.974 ± 1.1
saxpy/static workgroup=(1024,)/Float32/2048 15.3 ± 3.3 μs 14.5 ± 2 μs 1.06 ± 0.27
saxpy/static workgroup=(1024,)/Float32/256 15.3 ± 37 μs 14.8 ± 2.1 μs 1.03 ± 2.5
saxpy/static workgroup=(1024,)/Float32/262144 0.209 ± 0.029 ms 0.192 ± 0.027 ms 1.09 ± 0.21
saxpy/static workgroup=(1024,)/Float32/32768 0.0625 ± 0.014 ms 0.0632 ± 0.0041 ms 0.989 ± 0.23
saxpy/static workgroup=(1024,)/Float32/4096 0.0467 ± 0.035 ms 15.8 ± 2.6 μs 2.96 ± 2.2
saxpy/static workgroup=(1024,)/Float32/512 14.2 ± 4.2 μs 15.6 ± 37 μs 0.909 ± 2.2
saxpy/static workgroup=(1024,)/Float32/64 16.6 ± 2.2 μs 17.9 ± 1.5 μs 0.926 ± 0.14
saxpy/static workgroup=(1024,)/Float32/65536 0.0767 ± 0.014 ms 0.0803 ± 0.011 ms 0.956 ± 0.21
saxpy/static workgroup=(1024,)/Float64/1024 16 ± 2.1 μs 15.3 ± 4.1 μs 1.05 ± 0.31
saxpy/static workgroup=(1024,)/Float64/1048576 1.29 ± 0.1 ms 1.21 ± 0.09 ms 1.06 ± 0.12
saxpy/static workgroup=(1024,)/Float64/16384 0.0648 ± 0.011 ms 0.0696 ± 0.045 ms 0.932 ± 0.62
saxpy/static workgroup=(1024,)/Float64/2048 14.3 ± 1.7 μs 14.3 ± 2.9 μs 1 ± 0.23
saxpy/static workgroup=(1024,)/Float64/256 16.6 ± 2.1 μs 15.7 ± 2.2 μs 1.05 ± 0.2
saxpy/static workgroup=(1024,)/Float64/262144 0.348 ± 0.042 ms 0.321 ± 0.041 ms 1.08 ± 0.19
saxpy/static workgroup=(1024,)/Float64/32768 0.0821 ± 0.088 ms 0.0783 ± 0.016 ms 1.05 ± 1.1
saxpy/static workgroup=(1024,)/Float64/4096 0.0481 ± 0.04 ms 14.1 ± 0.91 μs 3.41 ± 2.8
saxpy/static workgroup=(1024,)/Float64/512 14.2 ± 2 μs 13.5 ± 2.1 μs 1.05 ± 0.22
saxpy/static workgroup=(1024,)/Float64/64 17.5 ± 1.7 μs 18.1 ± 34 μs 0.969 ± 1.9
saxpy/static workgroup=(1024,)/Float64/65536 0.123 ± 0.014 ms 0.118 ± 0.077 ms 1.04 ± 0.69
time_to_load 0.427 ± 0.033 s 0.385 ± 0.032 s 1.11 ± 0.13
main 99223b5... main / 99223b5...
const/@Const/Float32/262144 9 allocs: 0.203 kB 9 allocs: 0.203 kB 1
const/@Const/Float32/65536 9 allocs: 0.203 kB 9 allocs: 0.203 kB 1
const/@Const/Float64/262144 9 allocs: 0.203 kB 9 allocs: 0.203 kB 1
const/@Const/Float64/65536 9 allocs: 0.203 kB 9 allocs: 0.203 kB 1
const/unmarked/Float32/262144 9 allocs: 0.203 kB 9 allocs: 0.203 kB 1
const/unmarked/Float32/65536 9 allocs: 0.203 kB 9 allocs: 0.203 kB 1
const/unmarked/Float64/262144 9 allocs: 0.203 kB 9 allocs: 0.203 kB 1
const/unmarked/Float64/65536 9 allocs: 0.203 kB 9 allocs: 0.203 kB 1
launch/3D static workgroup, dynamic ndrange 9 allocs: 0.219 kB 9 allocs: 0.219 kB 1
launch/3D static workgroup, static ndrange 9 allocs: 0.219 kB 9 allocs: 0.219 kB 1
launch/dynamic workgroup, dynamic ndrange 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
launch/dynamic workgroup, dynamic ndrange, workgroupsize given 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
launch/static workgroup, dynamic ndrange 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
launch/static workgroup, static ndrange 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
partition/dynamic workgroup, dynamic ndrange 2 allocs: 0.0625 kB 2 allocs: 0.0625 kB 1
partition/static workgroup, dynamic ndrange 2 allocs: 32 B 2 allocs: 32 B 1
partition/static workgroup, static ndrange 0 allocs: 0 B 0 allocs: 0 B
saxpy/default/Float16/1024 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float16/1048576 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/default/Float16/16384 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float16/2048 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float16/256 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
saxpy/default/Float16/262144 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/default/Float16/32768 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float16/4096 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float16/512 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float16/64 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
saxpy/default/Float16/65536 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/default/Float32/1024 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float32/1048576 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/default/Float32/16384 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float32/2048 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float32/256 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
saxpy/default/Float32/262144 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/default/Float32/32768 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float32/4096 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float32/512 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float32/64 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
saxpy/default/Float32/65536 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float64/1024 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float64/1048576 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/default/Float64/16384 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float64/2048 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float64/256 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
saxpy/default/Float64/262144 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/default/Float64/32768 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float64/4096 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float64/512 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float64/64 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
saxpy/default/Float64/65536 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/static workgroup=(1024,)/Float16/1024 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float16/1048576 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/static workgroup=(1024,)/Float16/16384 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float16/2048 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float16/256 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
saxpy/static workgroup=(1024,)/Float16/262144 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/static workgroup=(1024,)/Float16/32768 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float16/4096 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float16/512 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float16/64 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
saxpy/static workgroup=(1024,)/Float16/65536 8 allocs: 0.141 kB 12 allocs: 0.25 kB 0.562
saxpy/static workgroup=(1024,)/Float32/1024 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float32/1048576 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/static workgroup=(1024,)/Float32/16384 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float32/2048 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float32/256 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
saxpy/static workgroup=(1024,)/Float32/262144 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/static workgroup=(1024,)/Float32/32768 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float32/4096 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float32/512 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float32/64 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
saxpy/static workgroup=(1024,)/Float32/65536 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float64/1024 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float64/1048576 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/static workgroup=(1024,)/Float64/16384 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float64/2048 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float64/256 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
saxpy/static workgroup=(1024,)/Float64/262144 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/static workgroup=(1024,)/Float64/32768 12 allocs: 0.25 kB 8 allocs: 0.141 kB 1.78
saxpy/static workgroup=(1024,)/Float64/4096 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float64/512 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float64/64 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
saxpy/static workgroup=(1024,)/Float64/65536 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
time_to_load 0.2 k allocs: 11.8 kB 0.2 k allocs: 11.8 kB 1

Benchmark Plots

A plot of the benchmark results have been uploaded as an artifact to the workflow run for this PR.
Go to "Actions"->"Benchmark a pull request"->[the most recent run]->"Artifacts" (at the bottom).

@maleadt
maleadt marked this pull request as ready for review October 3, 2026 18:38
Comment thread src/foreach_index.jl Outdated
`foreach_index(f, y, 1:n)` asked every array for its backend, and a range has none, so it failed
with an error asking to implement `get_backend` for it, although a range can be indexed on any
backend. Like AcceleratedKernels, let ranges, `CartesianIndices` and `LinearIndices`, and Base's
views, reshapes and permutations of them, run on the backend of the other arrays. Unlike it, do
not fall back to the host if there are no other arrays: ask for the backend instead.
@maleadt
maleadt merged commit 0977a07 into main Oct 5, 2026
60 of 63 checks passed
@maleadt
maleadt deleted the tb/foreach_index branch October 5, 2026 16:59
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants