Repository navigation
Conversation
`@kernel` tests `__validindex` in every work-item and runs the body under that mask, even in workgroups that lie entirely inside the ndrange. PoCL vectorizes over the work-items of a workgroup and keeps the mask in the vectorized loop, which made cheap kernel bodies about twice as slow as without it (#845). Emit each region of the body twice: without the mask when the whole workgroup lies inside the ndrange, which is the same for all of its work-items, and masked otherwise. `__fullgroup` decides, and only the PoCL backend enables it; elsewhere it is `false`, and the unmasked copy is removed when compiling. This is the single-launch form of #449. Assisted-by: Claude Code
| @inline full_group(ctx, launch::Launch, iterspace, ndrange) = | ||
| if builtin(iterspace) && ndrange isa CartesianIndices | ||
| T = index_type(launch) | ||
| groupsize = narrow(T, size(workitems(iterspace))) | ||
| reduce(&, map((g, w, n) -> g * w <= n, group_position(ctx, launch), groupsize, narrow(T, size(ndrange))); init = true) | ||
| else | ||
| false | ||
| end |
There was a problem hiding this comment.
Write this as function full_group
Benchmark ResultsShow table
Benchmark PlotsA plot of the benchmark results have been uploaded as an artifact to the workflow run for this PR. |
|
This causes problems for kernels with Each region between Example: a neighbour sum through
Results are still correct (n = 4096 and 4099). I haven't measured what the error paths cost on PoCL. They do cost on GPUs. I checked whether enabling
Kernels without Emitting the whole body twice behind one uniform branch ( So for PoCL I'd limit the duplication to kernels without |
|
Testing by @maleadt
|
@kerneltests__validindexin every work-item and runs the body under that mask, also in workgroups that lie entirely inside thendrange. On the CPU backend that costs about 2x for cheap kernel bodies (#845): PoCL vectorizes over the work-items of a workgroup and keeps the mask inside the vectorized loop.This emits each region of the body twice: without the mask when the whole workgroup lies inside the
ndrange, and masked otherwise. The test is the same for all work-items of a workgroup, so PoCL's work-item loop only runs the unmasked copy for full workgroups. It is the single-launch form of #449: it covers padded ranges and tuned workgroup sizes, without a second launch for the remainder (which costs several µs on this backend).__fullgroup(ctx)decides. It isfalseby default, so the unmasked copy is removed when compiling and GPU back-ends are unaffected; the PoCL backend overrides it with__fullgroup_check(ctx). Other backends can opt in the same way if the mask is expensive for them.Alternatives I tried
__validindexreturningfull_group | validindex: depends on LLVM unswitching PoCL's work-item loop. It did for a tuned(128, 32, 1)workgroup (983 → 456 µs) but not for(64, 64, 1), which got 33% slower than main.64case most, but not tuned sizes or padded ranges.Measurements
7-point stencil (
r[I] = x[I±e₁] + x[I±e₂] + x[I±e₃] - 6x[I]) on a 128³ or 130³ interior, Julia 1.13, Ryzen 9 5950X, 1 thread pinned to fixed cores (other jobs were running on the machine), best of 3 interleaved runs of best-of-7, µs per launch:6464(64, 64)(64, 64)Results checked against a host reference for exact and padded ranges.
WaterLily v1.6.1 at 4 threads (Julia 1.13, seconds per 25 steps, measured with the same change before restricting it to PoCL, on top of #844, which doesn't change anything on Julia 1.13):
Compile time (WaterLily's constructor and first step) changed within noise, despite the duplicated body.
Tests
ndrangegives two copies of the body on PoCL, a static one (where every workgroup is full) one. It fails on main.@printcodegen check now uses a staticndrange, since a dynamic one has twoprintfcalls (one per copy). It still checks that a@printlowers to a single call.Closes #845. Related: #449.
🤖 Generated with Claude Code