Repository navigation
Conversation
vchuravy
commented
Oct 5, 2026
Comment on lines
+141
to
+145
| - KernelAbstractions is built on [KernelInterface](@ref kernelinterface), which defines | ||
| `Backend` and the host-side functions (`allocate`, `synchronize`, …) that | ||
| KernelAbstractions re-exports. User code is unaffected, but backends have to implement | ||
| KernelInterface, so KernelAbstractions 0.10 needs a release of the backend package that | ||
| supports it; see the [notes for backend implementations](@ref implementations_notes). |
Member
Author
There was a problem hiding this comment.
Suggested change
| - KernelAbstractions is built on [KernelInterface](@ref kernelinterface), which defines | |
| `Backend` and the host-side functions (`allocate`, `synchronize`, …) that | |
| KernelAbstractions re-exports. User code is unaffected, but backends have to implement | |
| KernelInterface, so KernelAbstractions 0.10 needs a release of the backend package that | |
| supports it; see the [notes for backend implementations](@ref implementations_notes). | |
| - KernelAbstractions is built on [KernelInterface](@ref kernelinterface), which defines | |
| `Backend` and the host-side functions (`allocate`, `synchronize`, …) that | |
| KernelAbstractions re-exports. User code is unaffected, but backends have to implement | |
| KernelInterface; see the [notes for backend implementations](@ref implementations_notes). |
vchuravy
commented
Oct 5, 2026
Comment on lines
+146
to
+148
| - The Enzyme extension has been removed temporarily, and is planned to return in a later | ||
| 0.10 release. Code that differentiates kernels with Enzyme has to stay on | ||
| KernelAbstractions 0.9 until then. |
Member
Author
There was a problem hiding this comment.
Suggested change
| - The Enzyme extension has been removed temporarily, and is planned to return in a later | |
| 0.10 release. Code that differentiates kernels with Enzyme has to stay on | |
| KernelAbstractions 0.9 until then. | |
| - The Enzyme extension has been removed temporarily, and is planned to return in a later | |
| 0.10 release.. |
Contributor
Benchmark ResultsShow table
Benchmark PlotsA plot of the benchmark results have been uploaded as an artifact to the workflow run for this PR. |
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #842 +/- ##
=======================================
Coverage 79.75% 79.75%
=======================================
Files 25 25
Lines 2139 2139
=======================================
Hits 1706 1706
Misses 433 433 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
The 0.10 notes did not mention the Julia 1.10 requirement, that the CPU backend now runs on PoCL (and that `CPU(; static=true)` is gone), the KernelInterface requirement for backends, the temporary removal of the Enzyme extension, `KernelAbstractions.@spawn`, or the removal of `isgpu`. The NUMA example still called `CPU(; static=true)`, which no longer exists. Drop that call and note on its page which of the remarks changed with the PoCL-based CPU backend. Assisted-by: Claude Code
The example pinned Julia's threads with ThreadPinning.jl and relied on `KernelAbstractions.zeros` for a parallel first touch. On 0.10 kernels run on PoCL's threads and `zeros` fills from the calling thread, so pin with `POCL_AFFINITY=1` and initialize the data with a kernel instead. The page now describes 0.10 only. Its numbers are from a system with a single memory domain, where the configurations don't differ; the page says that the gain on a system with multiple domains is not measured. Assisted-by: Claude Code
….10 notes Launching a kernel on the PoCL-based CPU backend costs several microseconds more than on the thread-based backend of 0.9, and a fixed one-dimensional workgroup size such as 64 pads multidimensional ranges. Recommend letting the workgroup size be chosen, and mention that kernels always run with the bounds mask (#845). Assisted-by: Claude Code
vchuravy
force-pushed
the
vc/release-notes-0.10
branch
from
October 9, 2026 08:28
795d078 to
cc5b112
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fills the gaps in the
### 0.10section ofdocs/src/index.mdahead of the 0.10 release, and fixes the NUMA example, which still used an option that no longer exists.Release notes
New entries for:
CPUbackend beingPOCLBackend, running on PoCL's own threads, and the removal ofCPU(; static=true);KernelAbstractions.@spawn(Add KernelAbstractions.@spawn and record_event/wait_event #750);KernelAbstractions.isgpu.CPUbackend, from re-measuring the WaterLily reports (WaterLily performance regression on CPU multi-threading backend for Julia 1.13.0 #696, WaterLily performance regression on CPU multi-threading backend for Julia 1.12.0 #644):64on theCPUbackend. A 1-D workgroup size pads N-d ranges (for(16, 16, 16), 256 workgroups of 64 work-items with 16 inside the range); a stencil over that range took 24 µs per launch with64against 8 µs with the chosen size;NUMA example
examples/numa_aware.jlcalledCPU(; static=true), which throws aMethodErroron main, and its advice was written for the thread-basedCPUbackend of 0.9. The example and its page are rewritten for the PoCL-based backend:POCL_AFFINITY=1pins PoCL's threads. ThreadPinning.jl pins Julia's threads, which no longer run the kernels, so the example drops it.staticoption anymore; the page points tojulia -t/JULIA_KA_CPU_THREADSfor the thread count.KernelAbstractions.zerosfills the array from the calling thread (allocate+fill!), so the example allocates withallocateand initializes with a kernel.The page's numbers are new, measured on main with 16 threads on a Ryzen 9 5950X. That machine has a single NUMA domain, so all four configurations (affinity on/off, parallel/serial initialization) give the same 41.2 GB/s. The page says so, and says that the gain on a system with multiple memory domains has not been measured with the PoCL-based backend. Numbers from a multi-socket machine would be a welcome follow-up.
Testing
The documentation builds locally (
julia --project=docs docs/make.jl, Julia 1.12) and the new cross-references resolve.🤖 Generated with Claude Code