Skip to content

Let a custom mapping define the linear index of its work-items - #848

Merged
vchuravy merged 1 commit into
mainfrom
vc/mapping-linear-index
Oct 7, 2026
Merged

vchuravy merged 1 commit into
mainfrom
vc/mapping-linear-index

Conversation

@vchuravy

@vchuravy vchuravy commented Oct 6, 2026 •

Copy link
Copy Markdown
Member

For a custom NDRange mapping, @index(Global, Linear) is linear_index(__ndrange(ctx), expand(iterspace, group, item)): the position of the expanded index within the context's ndrange. That doesn't fit a mapping whose linear index isn't a function of the expanded index alone. The motivating case is Oceananigans launching kernels over a list of active cells (CliMA/Oceananigans.jl#5799), where the linear index of a work-item is its position in the list. A cell can't be mapped back to that position without a reverse lookup, so that PR overrides the internal __index_Global_Linear for its context type instead. #781 documents that as something not to do, and it only works as long as @index lowers through that function.

This adds a hook on the iteration space:

NDIteration.linear_index(iterspace::NDRange, ndrange, groupidx::CartesianIndex, idx::CartesianIndex)

It defaults to the previous definition, linear_index(ndrange, expand(iterspace, groupidx, idx)), and global_linear calls it for iteration spaces other than KernelAbstractions' own. A mapping specializes it on its NDRange type, e.g. for Oceananigans:

NDIteration.linear_index(r::MappedNDRange, ::MappedIndices, g::CartesianIndex{1}, i::CartesianIndex{1}) = mapped_position(r, g, i)

KernelAbstractions' built-in iteration spaces (identity and offsets) still take the direct path, so their code doesn't change.

Open points

Testing

  • New test: the launch testsuite gets a launch over a list of indices, with a partial last workgroup. It checks that each listed index is visited once, that nothing else is touched, and that @index(Global, Linear) is the position in the list. Without the hook the kernel doesn't compile (linear_index(::ListedIndices, ::CartesianIndex) has no method); with it the test passes on the CPU backend. The GPU back-ends will run it through the shared testsuite.
  • Existing tests: the full test suite passes locally on Julia 1.12 with the hook (run before the new test was added; the launch testsuite including it passes on its own).
  • Oceananigans: a standalone copy of the IndexMap launch from CliMA/Oceananigans.jl@6852216 passes with the hook in place of the __index_Global_Linear override.

🤖 Generated with Claude Code

@codecov

codecov Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 79.83%. Comparing base (254fab8) to head (1656bf6).
⚠️ Report is 1 commits behind head on main.

Additional details and impacted files
@@           Coverage Diff           @@
##             main     #848   +/-   ##
=======================================
  Coverage   79.82%   79.83%           
=======================================
  Files          26       26           
  Lines        2151     2152    +1     
=======================================
+ Hits         1717     1718    +1     
  Misses        434      434           

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@github-actions

github-actions Bot commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Benchmark Results

Show table
main 1656bf6... main / 1656bf6...
const/@Const/Float32/262144 0.526 ± 0.019 ms 0.526 ± 0.021 ms 0.999 ± 0.054
const/@Const/Float32/65536 0.216 ± 0.0057 ms 0.216 ± 0.006 ms 1 ± 0.038
const/@Const/Float64/262144 0.896 ± 0.013 ms 0.889 ± 0.0098 ms 1.01 ± 0.018
const/@Const/Float64/65536 0.368 ± 0.016 ms 0.368 ± 0.014 ms 1 ± 0.059
const/unmarked/Float32/262144 2.29 ± 0.032 ms 2.23 ± 0.017 ms 1.02 ± 0.016
const/unmarked/Float32/65536 0.597 ± 0.019 ms 0.595 ± 0.017 ms 1 ± 0.043
const/unmarked/Float64/262144 3.12 ± 0.027 ms 3.15 ± 0.022 ms 0.989 ± 0.011
const/unmarked/Float64/65536 0.813 ± 0.016 ms 0.811 ± 0.012 ms 1 ± 0.025
launch/3D static workgroup, dynamic ndrange 10.9 ± 0.45 μs 10.6 ± 0.87 μs 1.03 ± 0.096
launch/3D static workgroup, static ndrange 10.7 ± 0.45 μs 10.4 ± 0.52 μs 1.03 ± 0.068
launch/dynamic workgroup, dynamic ndrange 12 ± 0.69 μs 11.7 ± 0.49 μs 1.02 ± 0.073
launch/dynamic workgroup, dynamic ndrange, workgroupsize given 12 ± 39 μs 12 ± 0.64 μs 1 ± 3.3
launch/static workgroup, dynamic ndrange 10.8 ± 0.35 μs 10.3 ± 0.29 μs 1.04 ± 0.045
launch/static workgroup, static ndrange 10.8 ± 0.62 μs 10.8 ± 2.3 μs 1 ± 0.22
partition/dynamic workgroup, dynamic ndrange 0.0587 ± 0.00097 μs 0.0583 ± 0.0014 μs 1.01 ± 0.03
partition/static workgroup, dynamic ndrange 0.0551 ± 0.012 μs 0.0577 ± 0.012 μs 0.954 ± 0.29
partition/static workgroup, static ndrange 1.55 ± 0.01 ns 1.56 ± 0.92 ns 0.994 ± 0.59
saxpy/default/Float16/1024 0.0567 ± 0.0068 ms 0.056 ± 0.0062 ms 1.01 ± 0.16
saxpy/default/Float16/1048576 1.48 ± 0.031 ms 1.46 ± 0.022 ms 1.01 ± 0.026
saxpy/default/Float16/16384 0.0746 ± 0.0095 ms 0.0733 ± 0.0067 ms 1.02 ± 0.16
saxpy/default/Float16/2048 0.0544 ± 0.0063 ms 0.0566 ± 0.0044 ms 0.96 ± 0.13
saxpy/default/Float16/256 11.5 ± 0.23 μs 11.4 ± 0.28 μs 1.01 ± 0.032
saxpy/default/Float16/262144 0.406 ± 0.018 ms 0.405 ± 0.018 ms 1 ± 0.064
saxpy/default/Float16/32768 0.0957 ± 0.0073 ms 0.0945 ± 0.0071 ms 1.01 ± 0.11
saxpy/default/Float16/4096 0.0593 ± 0.0048 ms 0.0591 ± 0.0034 ms 1 ± 0.099
saxpy/default/Float16/512 11.8 ± 9 μs 11.5 ± 0.62 μs 1.03 ± 0.78
saxpy/default/Float16/64 0.0447 ± 0.041 ms 11.2 ± 0.2 μs 3.98 ± 3.7
saxpy/default/Float16/65536 0.14 ± 0.006 ms 0.136 ± 0.0069 ms 1.03 ± 0.069
saxpy/default/Float32/1024 0.0537 ± 0.034 ms 0.0531 ± 0.032 ms 1.01 ± 0.88
saxpy/default/Float32/1048576 0.779 ± 0.061 ms 0.77 ± 0.097 ms 1.01 ± 0.15
saxpy/default/Float32/16384 0.0593 ± 0.0063 ms 0.0589 ± 0.0068 ms 1.01 ± 0.16
saxpy/default/Float32/2048 0.0541 ± 0.0036 ms 0.0566 ± 0.0071 ms 0.957 ± 0.14
saxpy/default/Float32/256 11.4 ± 0.55 μs 11.1 ± 0.33 μs 1.02 ± 0.058
saxpy/default/Float32/262144 0.213 ± 0.034 ms 0.207 ± 0.018 ms 1.03 ± 0.19
saxpy/default/Float32/32768 0.0728 ± 0.0059 ms 0.0698 ± 0.0085 ms 1.04 ± 0.15
saxpy/default/Float32/4096 0.057 ± 0.0039 ms 0.0569 ± 0.004 ms 1 ± 0.099
saxpy/default/Float32/512 13.9 ± 5.2 μs 11.8 ± 5.1 μs 1.17 ± 0.67
saxpy/default/Float32/64 11.3 ± 0.4 μs 11.3 ± 0.65 μs 1.01 ± 0.068
saxpy/default/Float32/65536 0.0965 ± 0.0085 ms 0.091 ± 0.009 ms 1.06 ± 0.14
saxpy/default/Float64/1024 0.0521 ± 0.034 ms 0.054 ± 0.012 ms 0.965 ± 0.67
saxpy/default/Float64/1048576 1.22 ± 0.14 ms 1.16 ± 0.13 ms 1.05 ± 0.17
saxpy/default/Float64/16384 0.0643 ± 0.0076 ms 0.0655 ± 0.0074 ms 0.981 ± 0.16
saxpy/default/Float64/2048 0.0552 ± 0.0061 ms 0.0556 ± 0.0054 ms 0.994 ± 0.15
saxpy/default/Float64/256 14.3 ± 3.7 μs 13.8 ± 5.4 μs 1.03 ± 0.49
saxpy/default/Float64/262144 0.267 ± 0.04 ms 0.265 ± 0.034 ms 1.01 ± 0.2
saxpy/default/Float64/32768 0.0825 ± 0.0094 ms 0.0792 ± 0.0079 ms 1.04 ± 0.16
saxpy/default/Float64/4096 0.0515 ± 0.0096 ms 0.0554 ± 0.0089 ms 0.931 ± 0.23
saxpy/default/Float64/512 21.9 ± 37 μs 20.4 ± 38 μs 1.07 ± 2.7
saxpy/default/Float64/64 11.2 ± 0.37 μs 11 ± 0.37 μs 1.02 ± 0.048
saxpy/default/Float64/65536 0.114 ± 0.017 ms 0.104 ± 0.013 ms 1.09 ± 0.21
saxpy/static workgroup=(1024,)/Float16/1024 0.0527 ± 0.0056 ms 0.0531 ± 0.0022 ms 0.992 ± 0.11
saxpy/static workgroup=(1024,)/Float16/1048576 1.48 ± 0.037 ms 1.47 ± 0.021 ms 1 ± 0.029
saxpy/static workgroup=(1024,)/Float16/16384 0.074 ± 0.0091 ms 0.0744 ± 0.01 ms 0.995 ± 0.18
saxpy/static workgroup=(1024,)/Float16/2048 0.0565 ± 0.0036 ms 0.0561 ± 0.006 ms 1.01 ± 0.12
saxpy/static workgroup=(1024,)/Float16/256 11.8 ± 1.1 μs 11.6 ± 0.27 μs 1.01 ± 0.095
saxpy/static workgroup=(1024,)/Float16/262144 0.408 ± 0.019 ms 0.41 ± 0.019 ms 0.994 ± 0.065
saxpy/static workgroup=(1024,)/Float16/32768 0.094 ± 0.0083 ms 0.0927 ± 0.0066 ms 1.01 ± 0.11
saxpy/static workgroup=(1024,)/Float16/4096 0.0603 ± 0.0035 ms 0.0596 ± 0.0049 ms 1.01 ± 0.1
saxpy/static workgroup=(1024,)/Float16/512 11.6 ± 0.65 μs 11.7 ± 35 μs 0.996 ± 3
saxpy/static workgroup=(1024,)/Float16/64 11.7 ± 0.22 μs 19.1 ± 42 μs 0.611 ± 1.3
saxpy/static workgroup=(1024,)/Float16/65536 0.139 ± 0.0094 ms 0.138 ± 0.0065 ms 1 ± 0.083
saxpy/static workgroup=(1024,)/Float32/1024 16.5 ± 37 μs 22.4 ± 39 μs 0.736 ± 2.1
saxpy/static workgroup=(1024,)/Float32/1048576 0.77 ± 0.082 ms 0.784 ± 0.052 ms 0.982 ± 0.12
saxpy/static workgroup=(1024,)/Float32/16384 0.0621 ± 0.0071 ms 0.0618 ± 0.0077 ms 1 ± 0.17
saxpy/static workgroup=(1024,)/Float32/2048 0.0552 ± 0.0051 ms 0.0556 ± 0.004 ms 0.993 ± 0.12
saxpy/static workgroup=(1024,)/Float32/256 11.7 ± 0.84 μs 12 ± 39 μs 0.975 ± 3.2
saxpy/static workgroup=(1024,)/Float32/262144 0.227 ± 0.03 ms 0.218 ± 0.02 ms 1.04 ± 0.17
saxpy/static workgroup=(1024,)/Float32/32768 0.0711 ± 0.008 ms 0.0718 ± 0.0081 ms 0.991 ± 0.16
saxpy/static workgroup=(1024,)/Float32/4096 0.0518 ± 0.0078 ms 0.0511 ± 0.0096 ms 1.01 ± 0.24
saxpy/static workgroup=(1024,)/Float32/512 0.0515 ± 0.043 ms 0.0495 ± 0.04 ms 1.04 ± 1.2
saxpy/static workgroup=(1024,)/Float32/64 11.8 ± 40 μs 11.5 ± 18 μs 1.02 ± 3.8
saxpy/static workgroup=(1024,)/Float32/65536 0.0977 ± 0.0084 ms 0.0933 ± 0.009 ms 1.05 ± 0.14
saxpy/static workgroup=(1024,)/Float64/1024 0.0529 ± 0.032 ms 0.0522 ± 0.036 ms 1.02 ± 0.94
saxpy/static workgroup=(1024,)/Float64/1048576 1.2 ± 0.16 ms 1.25 ± 0.14 ms 0.964 ± 0.17
saxpy/static workgroup=(1024,)/Float64/16384 0.0665 ± 0.0089 ms 0.0622 ± 0.0068 ms 1.07 ± 0.18
saxpy/static workgroup=(1024,)/Float64/2048 0.0537 ± 0.0038 ms 0.0561 ± 0.0034 ms 0.957 ± 0.089
saxpy/static workgroup=(1024,)/Float64/256 0.0531 ± 0.042 ms 0.0547 ± 0.043 ms 0.971 ± 1.1
saxpy/static workgroup=(1024,)/Float64/262144 0.271 ± 0.033 ms 0.271 ± 0.033 ms 1 ± 0.17
saxpy/static workgroup=(1024,)/Float64/32768 0.0825 ± 0.0095 ms 0.0778 ± 0.008 ms 1.06 ± 0.16
saxpy/static workgroup=(1024,)/Float64/4096 0.054 ± 0.009 ms 0.0543 ± 0.0088 ms 0.994 ± 0.23
saxpy/static workgroup=(1024,)/Float64/512 0.0544 ± 0.0031 ms 0.054 ± 0.036 ms 1.01 ± 0.68
saxpy/static workgroup=(1024,)/Float64/64 0.0505 ± 0.041 ms 11.9 ± 8.6 μs 4.26 ± 4.7
saxpy/static workgroup=(1024,)/Float64/65536 0.111 ± 0.015 ms 0.103 ± 0.0098 ms 1.09 ± 0.18
time_to_load 0.514 ± 0.0012 s 0.514 ± 0.0034 s 1 ± 0.007
main 1656bf6... main / 1656bf6...
const/@Const/Float32/262144 9 allocs: 0.203 kB 9 allocs: 0.203 kB 1
const/@Const/Float32/65536 9 allocs: 0.203 kB 9 allocs: 0.203 kB 1
const/@Const/Float64/262144 9 allocs: 0.203 kB 9 allocs: 0.203 kB 1
const/@Const/Float64/65536 9 allocs: 0.203 kB 9 allocs: 0.203 kB 1
const/unmarked/Float32/262144 9 allocs: 0.203 kB 9 allocs: 0.203 kB 1
const/unmarked/Float32/65536 9 allocs: 0.203 kB 9 allocs: 0.203 kB 1
const/unmarked/Float64/262144 9 allocs: 0.203 kB 9 allocs: 0.203 kB 1
const/unmarked/Float64/65536 9 allocs: 0.203 kB 9 allocs: 0.203 kB 1
launch/3D static workgroup, dynamic ndrange 9 allocs: 0.219 kB 9 allocs: 0.219 kB 1
launch/3D static workgroup, static ndrange 9 allocs: 0.219 kB 9 allocs: 0.219 kB 1
launch/dynamic workgroup, dynamic ndrange 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
launch/dynamic workgroup, dynamic ndrange, workgroupsize given 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
launch/static workgroup, dynamic ndrange 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
launch/static workgroup, static ndrange 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
partition/dynamic workgroup, dynamic ndrange 2 allocs: 0.0625 kB 2 allocs: 0.0625 kB 1
partition/static workgroup, dynamic ndrange 2 allocs: 32 B 2 allocs: 32 B 1
partition/static workgroup, static ndrange 0 allocs: 0 B 0 allocs: 0 B
saxpy/default/Float16/1024 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float16/1048576 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/default/Float16/16384 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float16/2048 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float16/256 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
saxpy/default/Float16/262144 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/default/Float16/32768 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/default/Float16/4096 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float16/512 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float16/64 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
saxpy/default/Float16/65536 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/default/Float32/1024 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float32/1048576 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/default/Float32/16384 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float32/2048 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float32/256 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
saxpy/default/Float32/262144 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/default/Float32/32768 8 allocs: 0.141 kB 12 allocs: 0.25 kB 0.562
saxpy/default/Float32/4096 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float32/512 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float32/64 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
saxpy/default/Float32/65536 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/default/Float64/1024 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float64/1048576 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/default/Float64/16384 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float64/2048 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float64/256 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
saxpy/default/Float64/262144 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/default/Float64/32768 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/default/Float64/4096 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float64/512 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float64/64 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
saxpy/default/Float64/65536 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/static workgroup=(1024,)/Float16/1024 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float16/1048576 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/static workgroup=(1024,)/Float16/16384 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/static workgroup=(1024,)/Float16/2048 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float16/256 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
saxpy/static workgroup=(1024,)/Float16/262144 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/static workgroup=(1024,)/Float16/32768 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/static workgroup=(1024,)/Float16/4096 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float16/512 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float16/64 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
saxpy/static workgroup=(1024,)/Float16/65536 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/static workgroup=(1024,)/Float32/1024 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float32/1048576 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/static workgroup=(1024,)/Float32/16384 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float32/2048 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float32/256 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
saxpy/static workgroup=(1024,)/Float32/262144 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/static workgroup=(1024,)/Float32/32768 12 allocs: 0.25 kB 8 allocs: 0.141 kB 1.78
saxpy/static workgroup=(1024,)/Float32/4096 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float32/512 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float32/64 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
saxpy/static workgroup=(1024,)/Float32/65536 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/static workgroup=(1024,)/Float64/1024 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float64/1048576 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/static workgroup=(1024,)/Float64/16384 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float64/2048 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float64/256 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
saxpy/static workgroup=(1024,)/Float64/262144 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/static workgroup=(1024,)/Float64/32768 8 allocs: 0.141 kB 12 allocs: 0.25 kB 0.562
saxpy/static workgroup=(1024,)/Float64/4096 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float64/512 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float64/64 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
saxpy/static workgroup=(1024,)/Float64/65536 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
time_to_load 0.2 k allocs: 11.8 kB 0.2 k allocs: 11.8 kB 1

Benchmark Plots

A plot of the benchmark results have been uploaded as an artifact to the workflow run for this PR.
Go to "Actions"->"Benchmark a pull request"->[the most recent run]->"Artifacts" (at the bottom).

@vchuravy
vchuravy marked this pull request as ready for review October 6, 2026 14:55
@vchuravy
vchuravy requested a review from giordano October 6, 2026 14:55
giordano added a commit to CliMA/Oceananigans.jl that referenced this pull request Oct 6, 2026
…ns' hook

`@index(Global, Linear)` in a kernel launched over an active cells map
returned the work item's position in the map through a method of
`KernelAbstractions.__index_Global_Linear` for the context of mapped
launches. That's an internal function, which KernelAbstractions
documents as not to be extended, and the method only worked as long as
`@index` lowered through it.

KernelAbstractions now lets a custom `NDRange` mapping define the
linear index of its work items with
`NDIteration.linear_index(iterspace, ndrange, groupidx, idx)`
(JuliaGPU/KernelAbstractions.jl#848), so implement that for
`MappedNDRange` instead, and drop the then unused
`MappedCompilerMetadata`.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EtLmZUPY4g9hi8aJsM4Wrs

@giordano giordano left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Seems to do the job for Oceananigans!

`@index(Global, Linear)` for a custom `NDRange` mapping was the position of
the expanded index in the context's `ndrange`. A mapping whose linear
index isn't a function of the expanded index alone, such as a list of
indices whose linear index is the position in the list, had to override
the internal `__index_Global_Linear` instead.

Add `NDIteration.linear_index(iterspace, ndrange, groupidx, idx)`, which
defaults to the previous definition and which a mapping can specialize on
its `NDRange` type. Test a launch over a list of indices in the launch
testsuite.

Assisted-by: Claude Code
@vchuravy
vchuravy force-pushed the vc/mapping-linear-index branch from 0235cae to 1656bf6 Compare October 6, 2026 18:35
giordano added a commit to CliMA/Oceananigans.jl that referenced this pull request Oct 6, 2026
…ns' hook

`@index(Global, Linear)` in a kernel launched over an active cells map
returned the work item's position in the map through a method of
`KernelAbstractions.__index_Global_Linear` for the context of mapped
launches. That's an internal function, which KernelAbstractions
documents as not to be extended, and the method only worked as long as
`@index` lowered through it.

KernelAbstractions now lets a custom `NDRange` mapping define the
linear index of its work items with
`NDIteration.linear_index(iterspace, ndrange, groupidx, idx)`
(JuliaGPU/KernelAbstractions.jl#848), so implement that for
`MappedNDRange` instead, and drop the then unused
`MappedCompilerMetadata`.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EtLmZUPY4g9hi8aJsM4Wrs
@vchuravy
vchuravy merged commit b9ead6b into main Oct 7, 2026
69 checks passed
@vchuravy
vchuravy deleted the vc/mapping-linear-index branch October 7, 2026 08:17
giordano added a commit to CliMA/Oceananigans.jl that referenced this pull request Oct 7, 2026
…ns' hook

`@index(Global, Linear)` in a kernel launched over an active cells map
returned the work item's position in the map through a method of
`KernelAbstractions.__index_Global_Linear` for the context of mapped
launches. That's an internal function, which KernelAbstractions
documents as not to be extended, and the method only worked as long as
`@index` lowered through it.

KernelAbstractions now lets a custom `NDRange` mapping define the
linear index of its work items with
`NDIteration.linear_index(iterspace, ndrange, groupidx, idx)`
(JuliaGPU/KernelAbstractions.jl#848), so implement that for
`MappedNDRange` instead, and drop the then unused
`MappedCompilerMetadata`.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EtLmZUPY4g9hi8aJsM4Wrs
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants