Skip to content

Make synchronization cooperative, using GPUToolbox - #995

Merged
maleadt merged 4 commits into
mainfrom
tb/cooperative-wait
Oct 2, 2026
Merged

maleadt merged 4 commits into
mainfrom
tb/cooperative-wait

Conversation

@maleadt

@maleadt maleadt commented Sep 30, 2026 •

Copy link
Copy Markdown
Member

synchronize() already avoided blocking the calling thread, but it did so by polling the command buffer's status in a yield() loop for as long as the GPU was busy. On a task running on one of Julia's worker threads, that loop keeps a CPU core fully busy for the whole wait. device_synchronize() used a different mechanism again: an extra empty command buffer with a completion handler that notifies an AsyncCondition.

This PR switches both to cooperative_wait from GPUToolbox 3.3.1 (JuliaGPU/GPUToolbox.jl#23), the implementation CUDA.jl, oneAPI.jl, OpenCL.jl and KernelAbstractions' POCL back-end are adopting:

  • Short operations are still detected by polling the command buffer's status, so they don't pay for involving another thread.
  • Longer ones are handed to a small pool of worker threads that call waitUntilCompleted (which ObjectiveC.jl already calls GC-safely), while the calling task yields to others.

The logging drain, which used to call waitUntilCompleted on the calling thread, and the in-flight limit of batched queues use the same wait. device_synchronize no longer needs the sentinel command buffer. Because other tasks can now commit work while it waits, it only cleans up after command buffers that have actually completed, instead of force-releasing the roots of everything in every batched queue.

This needs GPUToolbox 3.3.1, which fixes two ways the worker threads could deadlock (JuliaGPU/GPUToolbox.jl#26). Workers no longer run finalizers, which could block before they wake up the waiting task. And waits that can't be polled no longer need a free worker when all of them are busy. The logging drain is such a wait: it also waits for the command buffer's completion handlers, which the buffer's status doesn't cover.

The second commit fixes a problem that made synchronization much less cooperative than intended. synchronize ran entirely inside @autoreleasepool, which takes a single global lock in ObjectiveC.jl. While one task waited, any other task that wanted to launch a kernel or synchronize blocked on that lock. Worse, a task synchronizing in a loop could starve a waiting task indefinitely, because it kept re-acquiring the lock before the waiter got to run:

other = @async while true
    @metal threads=1 short_kernel(b)
    synchronize()
end
@metal threads=1 long_kernel(a)   # ~20 ms
synchronize()                     # before: doesn't return until `other` stops

The pool is now only held around the parts of synchronize that call into Objective-C, not across the wait.

Measured on an M1 (8 GPU cores, macOS 27, Julia 1.12), comparing the old and new implementations in one process, alternating between them in blocks. Overhead is the median wall time of a launch plus synchronize() minus the kernel's GPU time; CPU is the process CPU time per synchronization, including the worker threads.

kernel main thread, before main thread, after spawned task, before spawned task, after
~0 143 µs, 56 µs CPU 143 µs, 56 µs CPU 240 µs, 285 µs CPU 242 µs, 244 µs CPU
100 µs 193 µs, 74 µs CPU 193 µs, 74 µs CPU 248 µs, 385 µs CPU 250 µs, 244 µs CPU
1 ms 192 µs, 146 µs CPU 192 µs, 146 µs CPU 248 µs, 1186 µs CPU 250 µs, 244 µs CPU
5 ms 207 µs, 528 µs CPU 262 µs, 416 µs CPU 251 µs, 5299 µs CPU 271 µs, 247 µs CPU

For comparison, a plain blocking waitUntilCompleted costs 167–223 µs on the main thread and 209–229 µs on a spawned task, at ~50 µs CPU. Metal itself accounts for most of these numbers: it takes 150–250 µs from committing even an empty kernel to reporting it complete.

With several tasks, each synchronizing its own queue:

scenario before after
20 ms kernel while another task launches and synchronizes in a loop >1 s (until the other task gave up) 20.6 ms
8 tasks, 4 × 30 ms and 4 × 10 ms kernels 165 ms (fully serialized) 86 ms

Caveats:

  • Waits that are handed to a worker wake up 20–50 µs later than status polling would. Part of that is waitUntilCompleted also waiting for the command buffer's completion handlers, and part is waking up the worker and the waiting task, which is more expensive on macOS than on Linux.
  • On the main thread the old loop didn't burn much CPU: every yield() there processes libuv events, which takes ~14 µs on macOS, so the thread mostly slept. For the same reason GPUToolbox's polling phase lasts ~3 ms on the main thread, so short waits there behave exactly as before. The CPU savings are for tasks on other threads.
  • Ctrl-C during synchronize now takes effect once the GPU work has completed, like on the other back-ends, instead of immediately.
  • Because long waits use waitUntilCompleted, completion handlers installed with on_completed must not depend on the synchronizing task, e.g. by entering @autoreleasepool while it holds the pool.
  • The in-flight limit can still wait inside a kernel launch's @autoreleasepool, so a launch that hits it still holds the lock while waiting.

The tests don't depend on timing. synchronize() or device_synchronize() waits for a command buffer that is blocked on an MTLSharedEvent, and another task on the same thread opens that gate once the first one is waiting, so the wait can only return if the other task got to run. The tests check this for both functions, that kernel errors are still reported after such a wait, that device_synchronize keeps the roots of work committed while it waited, and that a task launching kernels and synchronizing in a loop can't starve a waiting task: the gate only opens after the loop has synchronized while the other task waits, and the loop keeps going until that wait returns. If the gate isn't opened in time, for example because the wait blocks the thread, a watchdog on a foreign thread opens it after 60 s and the test fails instead of hanging. While the gate is closed, the watchdog also advances the event once a second, because the GPU aborts a command buffer that has been waiting on an event for about 5 s. The last two tests fail on main. With nonblocking_synchronization = false every gated check fails, which is why they are skipped in that mode.

Wait for command buffers with GPUToolbox's cooperative_wait: poll the
command buffer's status first, then call waitUntilCompleted on a worker
thread while the calling task yields. This replaces the unbounded yield
loop in synchronize, the completion-handler sentinel in
device_synchronize, and the blocking wait for logging command buffers.
@autoreleasepool takes a global lock, so holding it across the wait kept
other tasks from using Metal until the wait finished. Worse, a task that
synchronized in a loop could starve a waiting one indefinitely, as it
kept re-acquiring the lock before the waiter got to run.

@github-actions github-actions Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Metal Benchmarks

Details
Benchmark suite Current: f1cee1f Previous: d209fe6 Ratio
array/accumulate/Float32/1d 395708 ns 388291 ns 1.02
array/accumulate/Float32/dims=1 368250 ns 350375 ns 1.05
array/accumulate/Float32/dims=1L 8870542 ns 8763333 ns 1.01
array/accumulate/Float32/dims=2 434875 ns 428458 ns 1.01
array/accumulate/Float32/dims=2L 2696166 ns 2526958 ns 1.07
array/accumulate/Int64/1d 846541 ns 846375 ns 1.00
array/accumulate/Int64/dims=1 911125 ns 904042 ns 1.01
array/accumulate/Int64/dims=1L 9577458 ns 9422041 ns 1.02
array/accumulate/Int64/dims=2 1211875 ns 1211708 ns 1.00
array/accumulate/Int64/dims=2L 6550625 ns 6424250 ns 1.02
array/broadcast 227000 ns 228459 ns 0.99
array/construct 2416 ns 2291 ns 1.05
array/permutedims/2d 455375 ns 452666 ns 1.01
array/permutedims/3d 1025667 ns 1027541 ns 1.00
array/permutedims/4d 1108334 ns 1098416 ns 1.01
array/private/copy 229958 ns 226750 ns 1.01
array/private/copyto!/cpu_to_gpu 214958 ns 211583 ns 1.02
array/private/copyto!/gpu_to_cpu 218125 ns 208917 ns 1.04
array/private/copyto!/gpu_to_gpu 222042 ns 220375 ns 1.01
array/private/iteration/findall/bool 1058666 ns 1048542 ns 1.01
array/private/iteration/findall/int 1230167 ns 1224750 ns 1.00
array/private/iteration/findfirst/bool 1156583 ns 1153792 ns 1.00
array/private/iteration/findfirst/int 962125 ns 1156375 ns 0.83
array/private/iteration/findmin/1d 1202083 ns 1218292 ns 0.99
array/private/iteration/findmin/2d 1032917 ns 1024250 ns 1.01
array/private/iteration/logical 1668750 ns 1661084 ns 1.00
array/private/iteration/scalar 1411625 ns 1381000 ns 1.02
array/random/rand/Float32 418167 ns 414334 ns 1.01
array/random/rand/Int64 498000 ns 494209 ns 1.01
array/random/rand!/Float32 405833 ns 402500 ns 1.01
array/random/rand!/Int64 431834 ns 431167 ns 1.00
array/random/randn/Float32 390709 ns 388208 ns 1.01
array/random/randn!/Float32 375583 ns 367166 ns 1.02
array/reductions/mapreduce/Float32/1d 454500 ns 450000 ns 1.01
array/reductions/mapreduce/Float32/dims=1 352000 ns 349250 ns 1.01
array/reductions/mapreduce/Float32/dims=1L 622833 ns 611750 ns 1.02
array/reductions/mapreduce/Float32/dims=2 359666 ns 355334 ns 1.01
array/reductions/mapreduce/Float32/dims=2L 1240875 ns 1220750 ns 1.02
array/reductions/mapreduce/Int64/1d 644667 ns 628041 ns 1.03
array/reductions/mapreduce/Int64/dims=1 638042 ns 631541 ns 1.01
array/reductions/mapreduce/Int64/dims=1L 1010625 ns 1006459 ns 1.00
array/reductions/mapreduce/Int64/dims=2 793167 ns 788791 ns 1.01
array/reductions/mapreduce/Int64/dims=2L 2198375 ns 2200959 ns 1.00
array/reductions/reduce/Float32/1d 458625 ns 449833 ns 1.02
array/reductions/reduce/Float32/dims=1 361667 ns 356375 ns 1.01
array/reductions/reduce/Float32/dims=1L 609000 ns 607625 ns 1.00
array/reductions/reduce/Float32/dims=2 252083 ns 245042 ns 1.03
array/reductions/reduce/Float32/dims=2L 490542 ns 482709 ns 1.02
array/reductions/reduce/Int64/1d 642333 ns 628250 ns 1.02
array/reductions/reduce/Int64/dims=1 640583 ns 637625 ns 1.00
array/reductions/reduce/Int64/dims=1L 1011167 ns 1004292 ns 1.01
array/reductions/reduce/Int64/dims=2 261042 ns 269292 ns 0.97
array/reductions/reduce/Int64/dims=2L 676917 ns 665459 ns 1.02
array/shared/copy 131125 ns 128833 ns 1.02
array/shared/copyto!/cpu_to_gpu 37208 ns 37834 ns 0.98
array/shared/copyto!/gpu_to_cpu 37542 ns 37542 ns 1
array/shared/copyto!/gpu_to_gpu 38125 ns 37583 ns 1.01
array/shared/iteration/findall/bool 1061250 ns 1046041 ns 1.01
array/shared/iteration/findall/int 1228334 ns 1209208 ns 1.02
array/shared/iteration/findfirst/bool 983084 ns 980292 ns 1.00
array/shared/iteration/findfirst/int 854458 ns 988542 ns 0.86
array/shared/iteration/findmin/1d 1061042 ns 1055334 ns 1.01
array/shared/iteration/findmin/2d 1035583 ns 1032833 ns 1.00
array/shared/iteration/logical 1520041 ns 1528625 ns 0.99
array/shared/iteration/scalar 4886.857142857143 ns 3671.875 ns 1.33
array/sorting/1d 2105250 ns 2083750 ns 1.01
array/sorting/2d 8457959 ns 8302083 ns 1.02
integration/byval/reference 1126917 ns 1117666 ns 1.01
integration/byval/slices=1 1124834 ns 1119833 ns 1.00
integration/byval/slices=2 2029709 ns 2010584 ns 1.01
integration/byval/slices=3 6895500 ns 6549916 ns 1.05
integration/metaldevrt 397667 ns 393166 ns 1.01
kernel/indexing 219166 ns 211500 ns 1.04
kernel/indexing_checked 397625 ns 394709 ns 1.01
kernel/launch 2050.8888888888887 ns 2050.8888888888887 ns 1
kernel/rand 409792 ns 408041 ns 1.00
latency/import 1824430709 ns 1797481542 ns 1.01
latency/precompile 32456308500 ns 32287978750 ns 1.01
latency/ttfp 2281049708 ns 2252476334 ns 1.01
metal/synchronization/context 523.9947643979058 ns 536.4052631578948 ns 0.98
metal/synchronization/stream 465.13775510204084 ns 343.70506912442397 ns 1.35

This comment was automatically generated by workflow using github-action-benchmark.

@codecov

codecov Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 87.09%. Comparing base (d209fe6) to head (f1cee1f).

Additional details and impacted files
@@            Coverage Diff             @@
##             main     #995      +/-   ##
==========================================
+ Coverage   86.90%   87.09%   +0.18%     
==========================================
  Files          92       92              
  Lines        6773     6747      -26     
==========================================
- Hits         5886     5876      -10     
+ Misses        887      871      -16     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

3.3.1 keeps cooperative_wait's worker threads from deadlocking: they no
longer run finalizers, which could block before the waiting task is
notified, and waits that cannot be polled (like the one for logging
command buffers' completion handlers) no longer wait for a worker to
become available when all of them are busy.
The tests relied on kernels running for tens of milliseconds, and checked
a progress counter and the elapsed time. Instead, make synchronization
wait for a command buffer that is gated on a shared event, which another
task opens, so that the wait can only return after that task has run.
While the gate is closed, a watchdog on a foreign thread advances the
event step by step, as the GPU aborts command buffers that wait on an
event for too long. If the test doesn't finish in time, e.g. because
the wait blocks the thread, the watchdog opens the gate and records the
timeout, so that the test fails instead of hanging.
@maleadt
maleadt merged commit 6eb4996 into main Oct 2, 2026
19 checks passed
@maleadt
maleadt deleted the tb/cooperative-wait branch October 2, 2026 05:44
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant