Skip to content

CPU backend: the bounds mask makes kernel bodies up to 2.3x slower, even when the workgroups tile the ndrange #845

Description

@vchuravy

On the PoCL-based CPU backend every @kernel is compiled with DynamicCheck, so every work-item tests __validindex and the kernel body runs under a mask. This happens even when the workgroups tile the ndrange exactly and there are no padding lanes. For a cheap stencil kernel the mask costs about 2x, which puts main behind KA 0.9.43 for kernel bodies of this kind.

The thread-based CPU backend of 0.9 passed the dynamic result of partition through and ran exact partitions with NoDynamicCheck. The PoCL backend always used DynamicCheck (mkcontext in src/pocl/backend.jl, now generic in src/backend_launch.jl since #801), as the GPU back-ends do. This is the CPU side of #449, which proposes unchecked full workgroups plus a checked last one.

Measurements

7-point stencil over the interior of a 130³ Float32 array (ndrange = (128, 128, 128)), Julia 1.13.1, KA main 0977a07, Ryzen 9 5950X, best of 7, µs per launch. "Unmasked" forces NoDynamicCheck through a local prototype; that is safe here because every partition is exact.

workgroupsize threads main (masked) main (unmasked) KA 0.9.43
tuned (nothing) 1 986 429 442
64 1 1143 609 811
(64, 64) 1 979 529 448
64 4 322 182 229

With a padded range (130³ with workgroup size 64) the masked kernel takes 1746 µs at 1 thread.

The work-item loop is vectorized at width 8 with and without the mask (POCL_VECTORIZER_REMARKS=1), so the mask doesn't prevent vectorization. I haven't determined where the extra time goes; masked loads and stores in the vectorized loop are my guess.

Not every kernel is slower. WaterLily's conv_diff! is 20–30% faster on main than on 0.9.43 despite the mask. Its effect on WaterLily as a whole is small: about 2% on tgv 2^6 at 4 threads (context: #696, #644).

Reproducer

# WG=64 julia -t1 stencil.jl   (WG=nothing to let KA choose)
using KernelAbstractions, Printf
using KernelAbstractions: @kernel, @index, @Const, get_backend
@kernel function stencil_k!(r, @Const(x), @Const(I0))
    I = @index(Global, Cartesian)
    I += I0
    @fastmath @inbounds r[I] = x[I - CartesianIndex(1,0,0)] + x[I + CartesianIndex(1,0,0)] +
        x[I - CartesianIndex(0,1,0)] + x[I + CartesianIndex(0,1,0)] +
        x[I - CartesianIndex(0,0,1)] + x[I + CartesianIndex(0,0,1)] - 6f0*x[I]
end
const WG = let s = get(ENV, "WG", "nothing"); s == "nothing" ? nothing : eval(Meta.parse(s)); end
run!(r, x, R) = (WG === nothing ? stencil_k!(get_backend(r)) : stencil_k!(get_backend(r), WG))(r, x, R[1]-oneunit(R[1]), ndrange=size(R))
function bestof(f, args...)
    f(args...); t1 = @elapsed f(args...); reps = clamp(round(Int, 0.1/max(t1,1e-7)), 5, 100_000); best = Inf
    for _ in 1:7
        t = @elapsed for _ in 1:reps; f(args...); end
        best = min(best, t/reps)
    end
    best
end
n = 128
x = rand(Float32, n+2, n+2, n+2); y = zero(x); R = CartesianIndices((2:n+1, 2:n+1, 2:n+1))
@printf("KA %s wg=%s threads=%d: %.1f µs\n", pkgversion(KernelAbstractions), WG, Threads.nthreads(), bestof(run!, y, x, R)*1e6)

Possible fixes

  1. Exact partitions with a given workgroup size. Use NoDynamicCheck when the workgroup size is given and partition reports no padding. I have a prototype of this: about 15 lines in launch_tuple/launch_kernel in src/backend_launch.jl, limited to NDLaunch{Int32}. It gives the "unmasked" 64 rows above. The costs are a second compiled variant per kernel and the case it doesn't cover:
    • Tuned workgroup size. The kernel is compiled before tuning, and the tuned size must keep the context type (see select_launch), so whether the partition will be exact isn't known at compile time. Options: compile both variants and pick after tuning, or have the tuner prefer sizes that divide the ndrange.
  2. Split the launch, as in On CPU always use NoDynamicCheck(), just finish the last partial workgroup with DynamicCheck() #449. An unchecked launch over the full workgroups plus a checked one for the remainder. This also covers padded ranges, at the cost of a second launch, which is not cheap on this backend (several µs per launch).
  3. Make the masked body cheaper. Find out why the mask costs this much in PoCL's work-item loop, e.g. compare the final machine code (POCL_LEAVE_KERNEL_COMPILER_TEMP_FILES=1), and see whether a different formulation of the check vectorizes better. Check the bounds of a partial workgroup without a branch per dimension #844 already makes the check branch-free.

This is not a 0.10 release blocker in my view. The 0.10 release notes should mention it next to the higher per-launch cost.

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions