Skip to content

CPU back-end: limit the threads with a sub-device instead of a PoCL extension - #841

Merged
vchuravy merged 1 commit into
mainfrom
tb/pocl-subdevice
Oct 5, 2026
Merged

vchuravy merged 1 commit into
mainfrom
tb/pocl-subdevice

Conversation

@maleadt

@maleadt maleadt commented Oct 5, 2026

Copy link
Copy Markdown
Member

#821 limits the CPU back-end to as many PoCL threads as Julia has, by calling clSetCPUMaxComputeUnitsPOCL, an extension we carried in our PoCL builds. Upstream preferred making sub-devices usable over adding a global setting (pocl/pocl#2371), so the extension is gone again in pocl_standalone_jll 7.2.2, and this PR uses a sub-device instead.

A sub-device used to be just as slow as the whole device, because every command woke up all of PoCL's worker threads. pocl/pocl#2371 now only wakes the workers a command can use, on the sub-device it targets, and 7.2.2 includes that. So a sub-device sized to Julia's thread count behaves like the smaller thread pool the extension gave us. PoCL still starts a thread per hardware thread, but the ones outside the sub-device stay asleep and don't use CPU time.

Median saxpy launch time with -t4 on a 16-core/32-thread Ryzen 9950X:

1k elements 64k 1M
extension (main, 7.2.1) 5 µs 9 µs 27 µs
sub-device (this PR, 7.2.2) 3 µs 10 µs 23–27 µs
sub-device with 7.2.1 31 µs 38 µs 55 µs

The behavior is the same as in #821, with one exception: JULIA_KA_CPU_THREADS can no longer go above the number of hardware threads, since a sub-device can't be larger than its device. PoCL's own variables still can, as the CPU docstring now says. If the device can't be partitioned (PoCL built with OpenMP, or the cpu-tbb device), the whole device is used.

The precompilation workload sets PoCL's thread-count variables again, as before #821, so that precompiling doesn't start a thread per core in every process.

This requires pocl_standalone_jll 7.2.2 (JuliaPackaging/Yggdrasil#15005), so CI will fail until that's registered.

The full test suite passes with -t4 against a local build of the 7.2.2 patch series.

…xtension

clSetCPUMaxComputeUnitsPOCL didn't make it upstream, and pocl_standalone_jll
7.2.2 drops it again. Instead, PoCL now only wakes up the worker threads of
the sub-device a command targets (pocl/pocl#2371), so run kernels on a
sub-device with as many compute units as Julia has threads. Its other
threads stay asleep.

A sub-device can't be larger than its device, so JULIA_KA_CPU_THREADS can
no longer exceed the number of hardware threads; PoCL's variables still
can. The precompilation workload sets those again, so that precompiling
doesn't start a thread per core in every process.
@maleadt maleadt closed this Oct 5, 2026
@maleadt maleadt reopened this Oct 5, 2026
@codecov

codecov Bot commented Oct 5, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 93.33333% with 1 line in your changes missing coverage. Please review.
✅ Project coverage is 79.33%. Comparing base (952c482) to head (8dda43c).

Files with missing lines Patch % Lines
src/pocl/nanoOpenCL.jl 90.00% 1 Missing ⚠️
Additional details and impacted files
@@            Coverage Diff             @@
##             main     #841      +/-   ##
==========================================
- Coverage   79.34%   79.33%   -0.01%     
==========================================
  Files          24       24              
  Lines        2091     2095       +4     
==========================================
+ Hits         1659     1662       +3     
- Misses        432      433       +1     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@vchuravy
vchuravy merged commit 81b365f into main Oct 5, 2026
76 of 113 checks passed
@vchuravy
vchuravy deleted the tb/pocl-subdevice branch October 5, 2026 13:38
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants