Make synchronization cooperative, using GPUToolbox - #995
Merged
Merged
Conversation
Wait for command buffers with GPUToolbox's cooperative_wait: poll the command buffer's status first, then call waitUntilCompleted on a worker thread while the calling task yields. This replaces the unbounded yield loop in synchronize, the completion-handler sentinel in device_synchronize, and the blocking wait for logging command buffers.
@autoreleasepool takes a global lock, so holding it across the wait kept other tasks from using Metal until the wait finished. Worse, a task that synchronized in a loop could starve a waiting one indefinitely, as it kept re-acquiring the lock before the waiter got to run.
Contributor
There was a problem hiding this comment.
Metal Benchmarks
Details
| Benchmark suite | Current: f1cee1f | Previous: d209fe6 | Ratio |
|---|---|---|---|
array/accumulate/Float32/1d |
395708 ns |
388291 ns |
1.02 |
array/accumulate/Float32/dims=1 |
368250 ns |
350375 ns |
1.05 |
array/accumulate/Float32/dims=1L |
8870542 ns |
8763333 ns |
1.01 |
array/accumulate/Float32/dims=2 |
434875 ns |
428458 ns |
1.01 |
array/accumulate/Float32/dims=2L |
2696166 ns |
2526958 ns |
1.07 |
array/accumulate/Int64/1d |
846541 ns |
846375 ns |
1.00 |
array/accumulate/Int64/dims=1 |
911125 ns |
904042 ns |
1.01 |
array/accumulate/Int64/dims=1L |
9577458 ns |
9422041 ns |
1.02 |
array/accumulate/Int64/dims=2 |
1211875 ns |
1211708 ns |
1.00 |
array/accumulate/Int64/dims=2L |
6550625 ns |
6424250 ns |
1.02 |
array/broadcast |
227000 ns |
228459 ns |
0.99 |
array/construct |
2416 ns |
2291 ns |
1.05 |
array/permutedims/2d |
455375 ns |
452666 ns |
1.01 |
array/permutedims/3d |
1025667 ns |
1027541 ns |
1.00 |
array/permutedims/4d |
1108334 ns |
1098416 ns |
1.01 |
array/private/copy |
229958 ns |
226750 ns |
1.01 |
array/private/copyto!/cpu_to_gpu |
214958 ns |
211583 ns |
1.02 |
array/private/copyto!/gpu_to_cpu |
218125 ns |
208917 ns |
1.04 |
array/private/copyto!/gpu_to_gpu |
222042 ns |
220375 ns |
1.01 |
array/private/iteration/findall/bool |
1058666 ns |
1048542 ns |
1.01 |
array/private/iteration/findall/int |
1230167 ns |
1224750 ns |
1.00 |
array/private/iteration/findfirst/bool |
1156583 ns |
1153792 ns |
1.00 |
array/private/iteration/findfirst/int |
962125 ns |
1156375 ns |
0.83 |
array/private/iteration/findmin/1d |
1202083 ns |
1218292 ns |
0.99 |
array/private/iteration/findmin/2d |
1032917 ns |
1024250 ns |
1.01 |
array/private/iteration/logical |
1668750 ns |
1661084 ns |
1.00 |
array/private/iteration/scalar |
1411625 ns |
1381000 ns |
1.02 |
array/random/rand/Float32 |
418167 ns |
414334 ns |
1.01 |
array/random/rand/Int64 |
498000 ns |
494209 ns |
1.01 |
array/random/rand!/Float32 |
405833 ns |
402500 ns |
1.01 |
array/random/rand!/Int64 |
431834 ns |
431167 ns |
1.00 |
array/random/randn/Float32 |
390709 ns |
388208 ns |
1.01 |
array/random/randn!/Float32 |
375583 ns |
367166 ns |
1.02 |
array/reductions/mapreduce/Float32/1d |
454500 ns |
450000 ns |
1.01 |
array/reductions/mapreduce/Float32/dims=1 |
352000 ns |
349250 ns |
1.01 |
array/reductions/mapreduce/Float32/dims=1L |
622833 ns |
611750 ns |
1.02 |
array/reductions/mapreduce/Float32/dims=2 |
359666 ns |
355334 ns |
1.01 |
array/reductions/mapreduce/Float32/dims=2L |
1240875 ns |
1220750 ns |
1.02 |
array/reductions/mapreduce/Int64/1d |
644667 ns |
628041 ns |
1.03 |
array/reductions/mapreduce/Int64/dims=1 |
638042 ns |
631541 ns |
1.01 |
array/reductions/mapreduce/Int64/dims=1L |
1010625 ns |
1006459 ns |
1.00 |
array/reductions/mapreduce/Int64/dims=2 |
793167 ns |
788791 ns |
1.01 |
array/reductions/mapreduce/Int64/dims=2L |
2198375 ns |
2200959 ns |
1.00 |
array/reductions/reduce/Float32/1d |
458625 ns |
449833 ns |
1.02 |
array/reductions/reduce/Float32/dims=1 |
361667 ns |
356375 ns |
1.01 |
array/reductions/reduce/Float32/dims=1L |
609000 ns |
607625 ns |
1.00 |
array/reductions/reduce/Float32/dims=2 |
252083 ns |
245042 ns |
1.03 |
array/reductions/reduce/Float32/dims=2L |
490542 ns |
482709 ns |
1.02 |
array/reductions/reduce/Int64/1d |
642333 ns |
628250 ns |
1.02 |
array/reductions/reduce/Int64/dims=1 |
640583 ns |
637625 ns |
1.00 |
array/reductions/reduce/Int64/dims=1L |
1011167 ns |
1004292 ns |
1.01 |
array/reductions/reduce/Int64/dims=2 |
261042 ns |
269292 ns |
0.97 |
array/reductions/reduce/Int64/dims=2L |
676917 ns |
665459 ns |
1.02 |
array/shared/copy |
131125 ns |
128833 ns |
1.02 |
array/shared/copyto!/cpu_to_gpu |
37208 ns |
37834 ns |
0.98 |
array/shared/copyto!/gpu_to_cpu |
37542 ns |
37542 ns |
1 |
array/shared/copyto!/gpu_to_gpu |
38125 ns |
37583 ns |
1.01 |
array/shared/iteration/findall/bool |
1061250 ns |
1046041 ns |
1.01 |
array/shared/iteration/findall/int |
1228334 ns |
1209208 ns |
1.02 |
array/shared/iteration/findfirst/bool |
983084 ns |
980292 ns |
1.00 |
array/shared/iteration/findfirst/int |
854458 ns |
988542 ns |
0.86 |
array/shared/iteration/findmin/1d |
1061042 ns |
1055334 ns |
1.01 |
array/shared/iteration/findmin/2d |
1035583 ns |
1032833 ns |
1.00 |
array/shared/iteration/logical |
1520041 ns |
1528625 ns |
0.99 |
array/shared/iteration/scalar |
4886.857142857143 ns |
3671.875 ns |
1.33 |
array/sorting/1d |
2105250 ns |
2083750 ns |
1.01 |
array/sorting/2d |
8457959 ns |
8302083 ns |
1.02 |
integration/byval/reference |
1126917 ns |
1117666 ns |
1.01 |
integration/byval/slices=1 |
1124834 ns |
1119833 ns |
1.00 |
integration/byval/slices=2 |
2029709 ns |
2010584 ns |
1.01 |
integration/byval/slices=3 |
6895500 ns |
6549916 ns |
1.05 |
integration/metaldevrt |
397667 ns |
393166 ns |
1.01 |
kernel/indexing |
219166 ns |
211500 ns |
1.04 |
kernel/indexing_checked |
397625 ns |
394709 ns |
1.01 |
kernel/launch |
2050.8888888888887 ns |
2050.8888888888887 ns |
1 |
kernel/rand |
409792 ns |
408041 ns |
1.00 |
latency/import |
1824430709 ns |
1797481542 ns |
1.01 |
latency/precompile |
32456308500 ns |
32287978750 ns |
1.01 |
latency/ttfp |
2281049708 ns |
2252476334 ns |
1.01 |
metal/synchronization/context |
523.9947643979058 ns |
536.4052631578948 ns |
0.98 |
metal/synchronization/stream |
465.13775510204084 ns |
343.70506912442397 ns |
1.35 |
This comment was automatically generated by workflow using github-action-benchmark.
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #995 +/- ##
==========================================
+ Coverage 86.90% 87.09% +0.18%
==========================================
Files 92 92
Lines 6773 6747 -26
==========================================
- Hits 5886 5876 -10
+ Misses 887 871 -16 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
3.3.1 keeps cooperative_wait's worker threads from deadlocking: they no longer run finalizers, which could block before the waiting task is notified, and waits that cannot be polled (like the one for logging command buffers' completion handlers) no longer wait for a worker to become available when all of them are busy.
The tests relied on kernels running for tens of milliseconds, and checked a progress counter and the elapsed time. Instead, make synchronization wait for a command buffer that is gated on a shared event, which another task opens, so that the wait can only return after that task has run. While the gate is closed, a watchdog on a foreign thread advances the event step by step, as the GPU aborts command buffers that wait on an event for too long. If the test doesn't finish in time, e.g. because the wait blocks the thread, the watchdog opens the gate and records the timeout, so that the test fails instead of hanging.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
synchronize()already avoided blocking the calling thread, but it did so by polling the command buffer's status in ayield()loop for as long as the GPU was busy. On a task running on one of Julia's worker threads, that loop keeps a CPU core fully busy for the whole wait.device_synchronize()used a different mechanism again: an extra empty command buffer with a completion handler that notifies anAsyncCondition.This PR switches both to
cooperative_waitfrom GPUToolbox 3.3.1 (JuliaGPU/GPUToolbox.jl#23), the implementation CUDA.jl, oneAPI.jl, OpenCL.jl and KernelAbstractions' POCL back-end are adopting:waitUntilCompleted(which ObjectiveC.jl already calls GC-safely), while the calling task yields to others.The logging drain, which used to call
waitUntilCompletedon the calling thread, and the in-flight limit of batched queues use the same wait.device_synchronizeno longer needs the sentinel command buffer. Because other tasks can now commit work while it waits, it only cleans up after command buffers that have actually completed, instead of force-releasing the roots of everything in every batched queue.This needs GPUToolbox 3.3.1, which fixes two ways the worker threads could deadlock (JuliaGPU/GPUToolbox.jl#26). Workers no longer run finalizers, which could block before they wake up the waiting task. And waits that can't be polled no longer need a free worker when all of them are busy. The logging drain is such a wait: it also waits for the command buffer's completion handlers, which the buffer's status doesn't cover.
The second commit fixes a problem that made synchronization much less cooperative than intended.
synchronizeran entirely inside@autoreleasepool, which takes a single global lock in ObjectiveC.jl. While one task waited, any other task that wanted to launch a kernel or synchronize blocked on that lock. Worse, a task synchronizing in a loop could starve a waiting task indefinitely, because it kept re-acquiring the lock before the waiter got to run:The pool is now only held around the parts of
synchronizethat call into Objective-C, not across the wait.Measured on an M1 (8 GPU cores, macOS 27, Julia 1.12), comparing the old and new implementations in one process, alternating between them in blocks. Overhead is the median wall time of a launch plus
synchronize()minus the kernel's GPU time; CPU is the process CPU time per synchronization, including the worker threads.For comparison, a plain blocking
waitUntilCompletedcosts 167–223 µs on the main thread and 209–229 µs on a spawned task, at ~50 µs CPU. Metal itself accounts for most of these numbers: it takes 150–250 µs from committing even an empty kernel to reporting it complete.With several tasks, each synchronizing its own queue:
Caveats:
waitUntilCompletedalso waiting for the command buffer's completion handlers, and part is waking up the worker and the waiting task, which is more expensive on macOS than on Linux.yield()there processes libuv events, which takes ~14 µs on macOS, so the thread mostly slept. For the same reason GPUToolbox's polling phase lasts ~3 ms on the main thread, so short waits there behave exactly as before. The CPU savings are for tasks on other threads.synchronizenow takes effect once the GPU work has completed, like on the other back-ends, instead of immediately.waitUntilCompleted, completion handlers installed withon_completedmust not depend on the synchronizing task, e.g. by entering@autoreleasepoolwhile it holds the pool.@autoreleasepool, so a launch that hits it still holds the lock while waiting.The tests don't depend on timing.
synchronize()ordevice_synchronize()waits for a command buffer that is blocked on anMTLSharedEvent, and another task on the same thread opens that gate once the first one is waiting, so the wait can only return if the other task got to run. The tests check this for both functions, that kernel errors are still reported after such a wait, thatdevice_synchronizekeeps the roots of work committed while it waited, and that a task launching kernels and synchronizing in a loop can't starve a waiting task: the gate only opens after the loop has synchronized while the other task waits, and the loop keeps going until that wait returns. If the gate isn't opened in time, for example because the wait blocks the thread, a watchdog on a foreign thread opens it after 60 s and the test fails instead of hanging. While the gate is closed, the watchdog also advances the event once a second, because the GPU aborts a command buffer that has been waiting on an event for about 5 s. The last two tests fail on main. Withnonblocking_synchronization = falseevery gated check fails, which is why they are skipped in that mode.