Repository navigation
Conversation
… CPU `@kernel` runs the body under a per-work-item bounds mask, also in workgroups that lie entirely inside the ndrange. On the CPU backend the mask keeps PoCL from vectorizing the work-item loop, so math functions (`sin`, `exp`, `^`, ...) end up as scalar libm calls instead of PoCL's vector math library, and cheap bodies pay for masked loads and stores (#845, #860). For kernels without `@synchronize`, emit the body twice: unmasked when the whole workgroup lies inside the ndrange (the same for all of its work-items), and masked otherwise. `__fullgroup` decides; only the PoCL backend enables it, elsewhere the unmasked copy is removed when compiling. Each copy is a scope of its own, so that a local captured by a closure isn't assigned twice, which would box it. Kernels with `@synchronize` keep a single masked body: duplicating each region would define variables carried across a barrier on two paths and leave `UndefVarError` checks that LLVM can't remove (#849). Assisted-by: Claude Code
Benchmark ResultsShow table
Benchmark PlotsA plot of the benchmark results have been uploaded as an artifact to the workflow run for this PR. |
|
Running full workgroups unmasked makes some barrier-free kernels much slower on PoCL. Trixi.jl's flux differencing "full sweep" kernel (trixi-framework/Trixi.jl#3329) goes from 9.0 ms to 207–229 ms on 1 thread, and a minimal kernel of the same shape from 2.2 ms to 71 ms. Cause: PoCL's
using KernelAbstractions, BenchmarkTools, Random
# One work-item per node of an N×N element; each work-item loops over its partners in a
# rolled loop (no `@synchronize`), like a flux differencing "full sweep" kernel.
@inline function twopoint(ul::NTuple{4, T}, ur::NTuple{4, T}) where {T}
vl = ul[2] / ul[1]; vr = ur[2] / ur[1]
p = T(0.5) * (ul[4] / ul[1] + ur[4] / ur[1])
v = T(0.5) * (vl + vr)
return (v * (ul[1] + ur[1]), v * (ul[2] + ur[2]) + p, v * (ul[3] + ur[3]), v * (ul[4] + ur[4] + p))
end
# Keep the partner loop rolled, as Julia does for Trixi's larger two-point fluxes
macro rolled(ex)
push!(ex.args[2].args, Expr(:loopinfo, (Symbol("llvm.loop.unroll.disable"),)))
return esc(ex)
end
@kernel inbounds=true function fullsweep!(du, @Const(u), @Const(D), ::Val{N}) where {N}
i, j, e = @index(Global, NTuple)
ui = ntuple(v -> @inbounds(u[v, i, j, e]), Val(4))
acc = ntuple(_ -> zero(eltype(du)), Val(4))
@rolled for ii in 1:N
if ii != i
f = twopoint(ui, ntuple(v -> @inbounds(u[v, ii, j, e]), Val(4)))
acc = acc .+ D[i, ii] .* f
end
end
for v in 1:4
du[v, i, j, e] = acc[v]
end
end
const N = 4
nelem = 128^2
Random.seed!(1)
u = 1 .+ rand(4, N, N, nelem)
D = rand(N, N)
du = similar(u)
backend = CPU()
kernel! = fullsweep!(backend) # default workgroup size, as in Trixi
f!() = (kernel!(du, u, D, Val(N); ndrange = (N, N, nelem)); KernelAbstractions.synchronize(backend))
f!()
t = @belapsed f!() evals = 1 seconds = 5
println("KA ", pkgdir(KernelAbstractions), " POCL_FORCE_PARALLEL_OUTER_LOOP=", get(ENV, "POCL_FORCE_PARALLEL_OUTER_LOOP", "unset"),
": ", round(t * 1e3; digits = 2), " ms, sum(du) = ", sum(du))KA 1c2be35 (
Suggestions:
For the Trixi kernels themselves, this PR gave no speedup on PoCL. They have small work-item loops (4 nodes per direction), which LLVM's loop vectorizer doesn't vectorize below its tiny-trip-count threshold of 16 unless (Investigated with the assistance of Claude Code.) |
@kernelruns the body under a per-work-item bounds mask, also in workgroups that lie entirely inside the ndrange. On the CPU backend the mask keeps PoCL from vectorizing the work-item loop. Math functions then become one scalar libm call per work-item instead of using PoCL's vector math library (SLEEF), and cheap bodies pay for masked loads and stores (#845, #860).This is #849 restricted to kernels without
@synchronize. For those, the body is emitted twice: unmasked when the whole workgroup lies inside the ndrange (a test that is the same for all of its work-items), and masked otherwise.__fullgroupdecides. Only the PoCL backend enables it; elsewhere it'sfalseand the unmasked copy is removed when compiling. On CUDA and AMDGPU it gave nothing measurable for kernels without@synchronize(±3%), see the comment on Run full workgroups without the bounds mask on the CPU backend #849.@synchronize: duplicating each region, as Run full workgroups without the bounds mask on the CPU backend #849 did, defines variables carried across a barrier on two paths, and Julia then emitsUndefVarErrorchecks that LLVM doesn't remove. Those kernels keep the single masked body.let). Otherwise a local captured by a closure is assigned twice and gets boxed, which turns into dynamic calls.Measurements
y[i] = f(x[i])over 2^18Float64, 1 thread, Ryzen 9 5950X, ns per element. These numbers are from #849's branch, which does the same for kernels like these; see the correction on #860:funsafe_indices = truemap!sinexplogx^1.7tanOn this branch, measured with the machine heavily loaded (load average ~20), masked kernels run at 0.8–1.1x the time of the
unsafe_indicesones, where main's are 3.4–8x slower. A 7-point stencil over 128³ with the default workgroup size goes from 1655 to 633 µs.Tests
test/codegen_checks.jl):@synchronizegets two copies of the body with a dynamic ndrange, and one with a static ndrange where every workgroup is full;@synchronizekeeps one body and no exception paths. This check fails if the restriction is removed;letscoping.@printnow uses a static ndrange, since a dynamic one has aprintfper copy;Closes #845. Supersedes #849.
🤖 Generated with Claude Code