You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
CPU backend: the bounds mask makes kernel bodies up to 2.3x slower, even when the workgroups tile the ndrange #845
On the PoCL-based CPU backend every @kernel is compiled with DynamicCheck, so every work-item tests __validindex and the kernel body runs under a mask. This happens even when the workgroups tile the ndrange exactly and there are no padding lanes. For a cheap stencil kernel the mask costs about 2x, which puts main behind KA 0.9.43 for kernel bodies of this kind.
The thread-based CPU backend of 0.9 passed the dynamic result of partition through and ran exact partitions with NoDynamicCheck. The PoCL backend always used DynamicCheck (mkcontext in src/pocl/backend.jl, now generic in src/backend_launch.jl since #801), as the GPU back-ends do. This is the CPU side of #449, which proposes unchecked full workgroups plus a checked last one.
Measurements
7-point stencil over the interior of a 130³ Float32 array (ndrange = (128, 128, 128)), Julia 1.13.1, KA main 0977a07, Ryzen 9 5950X, best of 7, µs per launch. "Unmasked" forces NoDynamicCheck through a local prototype; that is safe here because every partition is exact.
workgroupsize
threads
main (masked)
main (unmasked)
KA 0.9.43
tuned (nothing)
1
986
429
442
64
1
1143
609
811
(64, 64)
1
979
529
448
64
4
322
182
229
With a padded range (130³ with workgroup size 64) the masked kernel takes 1746 µs at 1 thread.
The work-item loop is vectorized at width 8 with and without the mask (POCL_VECTORIZER_REMARKS=1), so the mask doesn't prevent vectorization. I haven't determined where the extra time goes; masked loads and stores in the vectorized loop are my guess.
Not every kernel is slower. WaterLily's conv_diff! is 20–30% faster on main than on 0.9.43 despite the mask. Its effect on WaterLily as a whole is small: about 2% on tgv 2^6 at 4 threads (context: #696, #644).
Reproducer
# WG=64 julia -t1 stencil.jl (WG=nothing to let KA choose)using KernelAbstractions, Printf
using KernelAbstractions:@kernel, @index, @Const, get_backend
@kernelfunctionstencil_k!(r, @Const(x), @Const(I0))
I =@index(Global, Cartesian)
I += I0
@fastmath@inbounds r[I] = x[I -CartesianIndex(1,0,0)] + x[I +CartesianIndex(1,0,0)] +
x[I -CartesianIndex(0,1,0)] + x[I +CartesianIndex(0,1,0)] +
x[I -CartesianIndex(0,0,1)] + x[I +CartesianIndex(0,0,1)] -6f0*x[I]
endconst WG =let s =get(ENV, "WG", "nothing"); s =="nothing"?nothing:eval(Meta.parse(s)); endrun!(r, x, R) = (WG ===nothing?stencil_k!(get_backend(r)) :stencil_k!(get_backend(r), WG))(r, x, R[1]-oneunit(R[1]), ndrange=size(R))
functionbestof(f, args...)
f(args...); t1 =@elapsedf(args...); reps =clamp(round(Int, 0.1/max(t1,1e-7)), 5, 100_000); best =Inffor _ in1:7
t =@elapsedfor _ in1:reps; f(args...); end
best =min(best, t/reps)
end
best
end
n =128
x =rand(Float32, n+2, n+2, n+2); y =zero(x); R =CartesianIndices((2:n+1, 2:n+1, 2:n+1))
@printf("KA %s wg=%s threads=%d: %.1f µs\n", pkgversion(KernelAbstractions), WG, Threads.nthreads(), bestof(run!, y, x, R)*1e6)
Possible fixes
Exact partitions with a given workgroup size. Use NoDynamicCheck when the workgroup size is given and partition reports no padding. I have a prototype of this: about 15 lines in launch_tuple/launch_kernel in src/backend_launch.jl, limited to NDLaunch{Int32}. It gives the "unmasked" 64 rows above. The costs are a second compiled variant per kernel and the case it doesn't cover:
Tuned workgroup size. The kernel is compiled before tuning, and the tuned size must keep the context type (see select_launch), so whether the partition will be exact isn't known at compile time. Options: compile both variants and pick after tuning, or have the tuner prefer sizes that divide the ndrange.
Make the masked body cheaper. Find out why the mask costs this much in PoCL's work-item loop, e.g. compare the final machine code (POCL_LEAVE_KERNEL_COMPILER_TEMP_FILES=1), and see whether a different formulation of the check vectorizes better. Check the bounds of a partial workgroup without a branch per dimension #844 already makes the check branch-free.
This is not a 0.10 release blocker in my view. The 0.10 release notes should mention it next to the higher per-launch cost.
On the PoCL-based
CPUbackend every@kernelis compiled withDynamicCheck, so every work-item tests__validindexand the kernel body runs under a mask. This happens even when the workgroups tile thendrangeexactly and there are no padding lanes. For a cheap stencil kernel the mask costs about 2x, which puts main behind KA 0.9.43 for kernel bodies of this kind.The thread-based
CPUbackend of 0.9 passed thedynamicresult ofpartitionthrough and ran exact partitions withNoDynamicCheck. The PoCL backend always usedDynamicCheck(mkcontextinsrc/pocl/backend.jl, now generic insrc/backend_launch.jlsince #801), as the GPU back-ends do. This is the CPU side of #449, which proposes unchecked full workgroups plus a checked last one.Measurements
7-point stencil over the interior of a 130³
Float32array (ndrange = (128, 128, 128)), Julia 1.13.1, KA main 0977a07, Ryzen 9 5950X, best of 7, µs per launch. "Unmasked" forcesNoDynamicCheckthrough a local prototype; that is safe here because every partition is exact.nothing)64(64, 64)64With a padded range (130³ with workgroup size 64) the masked kernel takes 1746 µs at 1 thread.
The work-item loop is vectorized at width 8 with and without the mask (
POCL_VECTORIZER_REMARKS=1), so the mask doesn't prevent vectorization. I haven't determined where the extra time goes; masked loads and stores in the vectorized loop are my guess.Not every kernel is slower. WaterLily's
conv_diff!is 20–30% faster on main than on 0.9.43 despite the mask. Its effect on WaterLily as a whole is small: about 2% on tgv 2^6 at 4 threads (context: #696, #644).Reproducer
Possible fixes
NoDynamicCheckwhen the workgroup size is given andpartitionreports no padding. I have a prototype of this: about 15 lines inlaunch_tuple/launch_kernelinsrc/backend_launch.jl, limited toNDLaunch{Int32}. It gives the "unmasked"64rows above. The costs are a second compiled variant per kernel and the case it doesn't cover:select_launch), so whether the partition will be exact isn't known at compile time. Options: compile both variants and pick after tuning, or have the tuner prefer sizes that divide thendrange.NoDynamicCheck(), just finish the last partial workgroup withDynamicCheck()#449. An unchecked launch over the full workgroups plus a checked one for the remainder. This also covers padded ranges, at the cost of a second launch, which is not cheap on this backend (several µs per launch).POCL_LEAVE_KERNEL_COMPILER_TEMP_FILES=1), and see whether a different formulation of the check vectorizes better. Check the bounds of a partial workgroup without a branch per dimension #844 already makes the check branch-free.This is not a 0.10 release blocker in my view. The 0.10 release notes should mention it next to the higher per-launch cost.