Skip to content

Add a low-cost tracing subsystem - #836

Open
vchuravy wants to merge 9 commits into
mainfrom
vc/tracing-v2
Open

vchuravy wants to merge 9 commits into
mainfrom
vc/tracing-v2

Conversation

@vchuravy

@vchuravy vchuravy commented Oct 4, 2026

Copy link
Copy Markdown
Member

Adds a low-cost tracing subsystem that puts named ranges on the timeline of a tracing profiler. This supersedes #703 and #66.

@profiling_range "volume integral" domain = "Trixi" begin
    volume_integral!(du, u, backend)
end
profiling_mark("converged")

Kernel launches are also recorded as ranges, named after the kernel.

Design

NVTX, ITT and roctx annotate host threads, and the profiler attributes the device work launched within a range to that range. So which profiler records a range depends on what the process runs under, not on the backend. For example, running the CPU backend under Nsight Systems gives NVTX ranges, and a GPU backend under VTune gives ITT tasks. For this reason the API takes no backend argument, and there are no KernelInterface hooks. Every range goes to all registered KernelAbstractions.Tracers.

  • Cost when no profiler is attached: a single atomic load (about 0.6 ns), because the tracer list is copy-on-write. The label isn't evaluated, so it can use string interpolation for free. The tracing part of a kernel launch is kept in a separate non-inlined function.
  • Ranges use start/end, not push/pop, so they may end on another thread and needn't nest. A range is ended by the tracers that were registered when it started.
  • @profiling_range uses a scope-free :tryfinally, like @time. The range is ended if the expression throws, and assignments inside it remain visible afterwards.

Tracers

Each tracer registers itself only when its profiler is attached:

Tracer Trigger Active when
NVTXExt NVTX.jl (also loaded by CUDA.jl) NVTX.isactive(), i.e. under nsys
ROCTXExt AMDGPU.jl under rocprofv3 (ROCP_TOOL_LIBRARIES) or legacy rocprof (HSA_TOOLS_LIB)
IntelITTExt IntelITT.jl IntelITT.isactive(), i.e. under VTune
NVTXTTracer built in JULIA_KA_NVTXT=1 or =path-%p.nvtxt

NVTXTTracer revives #66. It writes the NVTXT text format, which Nsight Systems can import with ImportNvtxt, so you can trace without a profiler attached.

Testing

  • The test suite covers the macro, registration, ranges, markers, ranges across many tasks, multiple tracers, NVTXT output (including the environment variable, in a subprocess), and the backend testsuite with a tracer registered.
  • Verified by hand under nsys profile --trace=nvtx: nvtx_sum shows Demo:step 1 and KernelAbstractions:scale! ranges.
  • ROCTX is not covered by CI. No AMDGPU release is compatible with KA 0.10 yet. I exercised the extension's code against /opt/rocm/lib/libroctx64.so (ROCm 7.2) using a stand-in AMDGPU module. I could not test it under rocprofv3, which isn't installed.

🤖 Generated with Claude Code

`@profiling_range` and `profiling_mark` put named ranges and markers on
the timeline of a tracing profiler, and kernel launches are annotated
with the kernel's name. Profilers like NVTX, ITT and roctx annotate host
threads and correlate device work themselves, so ranges go to every
registered `Tracer` regardless of backend. With none registered, an
annotation costs one atomic load.

Tracers: NVTX (NVTXExt), roctx (ROCTXExt, via AMDGPU), Intel ITT
(IntelITTExt), each registering only under its profiler, and a built-in
NVTXT writer enabled with `JULIA_KA_NVTXT`.

Supersedes #703 and #66.

Assisted-by: Claude Code (Opus 5.5)
Comment on lines +3 to +6
KernelAbstractions can put named ranges on the timeline of a tracing profiler, such as
NVIDIA Nsight Systems or Intel VTune, so that you can see which part of your program a
stretch of kernels belongs to. Annotations are cheap when no profiler is listening: a
single atomic load, and the label isn't even built.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
KernelAbstractions can put named ranges on the timeline of a tracing profiler, such as
NVIDIA Nsight Systems or Intel VTune, so that you can see which part of your program a
stretch of kernels belongs to. Annotations are cheap when no profiler is listening: a
single atomic load, and the label isn't even built.
KernelAbstractions can put named ranges on the timeline of a tracing profiler, such as
NVIDIA Nsight Systems or Intel VTune, so that you can see which part of your program a
stretch of kernels belongs to. Annotations are cheap when no profiler is listening.

Comment on lines +33 to +37
Ranges are recorded on the host threads of the process, which is how NVTX, ITT and
roctx work: it is the profiler that attributes the device work launched within a range to
it. So which profiler records the ranges depends on what the process runs under, not on the
backend: running the CPU backend under Nsight Systems gives NVTX ranges, and a GPU backend
under VTune gives ITT tasks.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
Ranges are recorded on the host threads of the process, which is how NVTX, ITT and
roctx work: it is the profiler that attributes the device work launched within a range to
it. So which profiler records the ranges depends on what the process runs under, not on the
backend: running the CPU backend under Nsight Systems gives NVTX ranges, and a GPU backend
under VTune gives ITT tasks.
Ranges are recorded on the host threads of the process, which is how NVTX, ITT and
roctx work: it is the profiler that attributes the device work launched within a range to
it.

This was referenced Oct 4, 2026
@codecov

codecov Bot commented Oct 4, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 87.53709% with 42 lines in your changes missing coverage. Please review.
✅ Project coverage is 76.32%. Comparing base (e70abc3) to head (d6c0967).
⚠️ Report is 3 commits behind head on main.

Files with missing lines Patch % Lines
ext/ROCTXExt.jl 0.00% 29 Missing ⚠️
src/profiling.jl 95.04% 6 Missing ⚠️
src/KernelAbstractions.jl 0.00% 3 Missing ⚠️
src/profiler.jl 98.29% 2 Missing ⚠️
ext/IntelITTExt.jl 92.30% 1 Missing ⚠️
ext/NVTXExt.jl 95.23% 1 Missing ⚠️
Additional details and impacted files
@@            Coverage Diff             @@
##             main     #836      +/-   ##
==========================================
- Coverage   79.06%   76.32%   -2.75%     
==========================================
  Files          24       29       +5     
  Lines        2040     2547     +507     
==========================================
+ Hits         1613     1944     +331     
- Misses        427      603     +176     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

Records the ranges, markers and kernel launches of an expression with a
temporary tracer, and summarizes them per name or, with `trace = true`,
lists them in order. By default kernel launches synchronize their
backend while profiling, so that their ranges measure execution rather
than launch; tracers opt into this with `synchronizes_launches`.

Assisted-by: Claude Code (Opus 5.5)
@github-actions

github-actions Bot commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Benchmark Results

Show table
main d6c0967... main / d6c0967...
const/@Const/Float32/262144 0.525 ± 0.019 ms 0.527 ± 0.02 ms 0.997 ± 0.052
const/@Const/Float32/65536 0.169 ± 0.0078 ms 0.17 ± 0.0089 ms 0.997 ± 0.069
const/@Const/Float64/262144 0.886 ± 0.014 ms 0.899 ± 0.021 ms 0.985 ± 0.027
const/@Const/Float64/65536 0.37 ± 0.015 ms 0.368 ± 0.016 ms 1 ± 0.06
const/unmarked/Float32/262144 2.27 ± 0.015 ms 2.31 ± 0.043 ms 0.984 ± 0.02
const/unmarked/Float32/65536 0.597 ± 0.015 ms 0.6 ± 0.015 ms 0.995 ± 0.035
const/unmarked/Float64/262144 3.14 ± 0.036 ms 3.17 ± 0.03 ms 0.991 ± 0.015
const/unmarked/Float64/65536 0.805 ± 0.011 ms 0.812 ± 0.022 ms 0.992 ± 0.03
launch/3D static workgroup, dynamic ndrange 0.053 ± 0.043 ms 10.6 ± 0.26 μs 5.02 ± 4.1
launch/3D static workgroup, static ndrange 10.5 ± 0.53 μs 10.9 ± 44 μs 0.962 ± 3.9
launch/dynamic workgroup, dynamic ndrange 12 ± 0.25 μs 11.9 ± 0.43 μs 1.01 ± 0.042
launch/dynamic workgroup, dynamic ndrange, workgroupsize given 12.2 ± 24 μs 12.2 ± 0.22 μs 0.993 ± 2
launch/static workgroup, dynamic ndrange 10.7 ± 0.3 μs 10.8 ± 0.28 μs 0.99 ± 0.038
launch/static workgroup, static ndrange 10.7 ± 0.59 μs 10.7 ± 0.52 μs 1.01 ± 0.074
partition/dynamic workgroup, dynamic ndrange 0.0582 ± 0.00087 μs 0.0581 ± 0.00097 μs 1 ± 0.022
partition/static workgroup, dynamic ndrange 0.0558 ± 0.012 μs 0.0593 ± 0.011 μs 0.941 ± 0.26
partition/static workgroup, static ndrange 1.55 ± 0.01 ns 1.55 ± 0.01 ns 1 ± 0.0091
saxpy/default/Float16/1024 0.0564 ± 0.006 ms 0.0567 ± 0.0033 ms 0.993 ± 0.12
saxpy/default/Float16/1048576 1.48 ± 0.029 ms 1.48 ± 0.027 ms 1 ± 0.027
saxpy/default/Float16/16384 0.0737 ± 0.0086 ms 0.0744 ± 0.0085 ms 0.991 ± 0.16
saxpy/default/Float16/2048 0.0549 ± 0.0043 ms 0.0556 ± 0.0029 ms 0.988 ± 0.093
saxpy/default/Float16/256 11.5 ± 0.58 μs 11.6 ± 0.59 μs 0.985 ± 0.071
saxpy/default/Float16/262144 0.406 ± 0.017 ms 0.409 ± 0.017 ms 0.993 ± 0.059
saxpy/default/Float16/32768 0.0956 ± 0.0062 ms 0.0932 ± 0.0075 ms 1.03 ± 0.11
saxpy/default/Float16/4096 0.0598 ± 0.005 ms 0.0608 ± 0.0042 ms 0.983 ± 0.11
saxpy/default/Float16/512 11.4 ± 0.39 μs 11.6 ± 0.5 μs 0.975 ± 0.054
saxpy/default/Float16/64 11.1 ± 0.23 μs 11.4 ± 0.41 μs 0.975 ± 0.041
saxpy/default/Float16/65536 0.137 ± 0.0077 ms 0.142 ± 0.009 ms 0.968 ± 0.082
saxpy/default/Float32/1024 19.4 ± 38 μs 0.0517 ± 0.041 ms 0.376 ± 0.79
saxpy/default/Float32/1048576 0.781 ± 0.095 ms 0.809 ± 0.07 ms 0.964 ± 0.14
saxpy/default/Float32/16384 0.0605 ± 0.0066 ms 0.0627 ± 0.0094 ms 0.965 ± 0.18
saxpy/default/Float32/2048 0.0567 ± 0.0048 ms 0.0554 ± 0.0056 ms 1.02 ± 0.13
saxpy/default/Float32/256 11.1 ± 0.47 μs 11.9 ± 42 μs 0.93 ± 3.3
saxpy/default/Float32/262144 0.221 ± 0.036 ms 0.216 ± 0.024 ms 1.02 ± 0.2
saxpy/default/Float32/32768 0.0723 ± 0.0072 ms 0.0723 ± 0.008 ms 1 ± 0.15
saxpy/default/Float32/4096 0.0554 ± 0.0064 ms 0.0575 ± 0.006 ms 0.962 ± 0.15
saxpy/default/Float32/512 14.3 ± 5.5 μs 0.0505 ± 0.039 ms 0.284 ± 0.25
saxpy/default/Float32/64 11.2 ± 40 μs 11.3 ± 0.37 μs 0.989 ± 3.6
saxpy/default/Float32/65536 0.0962 ± 0.0099 ms 0.0976 ± 0.01 ms 0.986 ± 0.14
saxpy/default/Float64/1024 0.0547 ± 0.038 ms 0.0545 ± 0.0059 ms 1 ± 0.71
saxpy/default/Float64/1048576 1.19 ± 0.19 ms 1.36 ± 0.16 ms 0.876 ± 0.17
saxpy/default/Float64/16384 0.0668 ± 0.0068 ms 0.0676 ± 0.0073 ms 0.988 ± 0.15
saxpy/default/Float64/2048 0.0569 ± 0.005 ms 0.0556 ± 0.0062 ms 1.02 ± 0.15
saxpy/default/Float64/256 14.8 ± 2.3 μs 16.1 ± 39 μs 0.92 ± 2.2
saxpy/default/Float64/262144 0.276 ± 0.041 ms 0.272 ± 0.042 ms 1.01 ± 0.22
saxpy/default/Float64/32768 0.0808 ± 0.01 ms 0.0887 ± 0.01 ms 0.912 ± 0.15
saxpy/default/Float64/4096 0.056 ± 0.01 ms 0.0527 ± 0.0079 ms 1.06 ± 0.25
saxpy/default/Float64/512 28 ± 38 μs 0.0542 ± 0.036 ms 0.515 ± 0.78
saxpy/default/Float64/64 11 ± 0.8 μs 11.1 ± 0.3 μs 0.99 ± 0.077
saxpy/default/Float64/65536 0.108 ± 0.012 ms 0.109 ± 0.017 ms 0.989 ± 0.19
saxpy/static workgroup=(1024,)/Float16/1024 0.0548 ± 0.0056 ms 0.0544 ± 0.0063 ms 1.01 ± 0.15
saxpy/static workgroup=(1024,)/Float16/1048576 1.48 ± 0.03 ms 1.49 ± 0.024 ms 0.999 ± 0.026
saxpy/static workgroup=(1024,)/Float16/16384 0.0739 ± 0.0081 ms 0.0745 ± 0.0094 ms 0.993 ± 0.17
saxpy/static workgroup=(1024,)/Float16/2048 0.057 ± 0.0058 ms 0.0572 ± 0.0039 ms 0.996 ± 0.12
saxpy/static workgroup=(1024,)/Float16/256 11.8 ± 42 μs 20.2 ± 42 μs 0.583 ± 2.4
saxpy/static workgroup=(1024,)/Float16/262144 0.409 ± 0.017 ms 0.412 ± 0.02 ms 0.992 ± 0.064
saxpy/static workgroup=(1024,)/Float16/32768 0.0948 ± 0.0072 ms 0.0974 ± 0.0078 ms 0.973 ± 0.11
saxpy/static workgroup=(1024,)/Float16/4096 0.061 ± 0.0032 ms 0.058 ± 0.0043 ms 1.05 ± 0.096
saxpy/static workgroup=(1024,)/Float16/512 11.6 ± 0.3 μs 21.6 ± 41 μs 0.535 ± 1
saxpy/static workgroup=(1024,)/Float16/64 11.4 ± 0.24 μs 11.6 ± 0.51 μs 0.984 ± 0.048
saxpy/static workgroup=(1024,)/Float16/65536 0.138 ± 0.0076 ms 0.142 ± 0.0094 ms 0.974 ± 0.084
saxpy/static workgroup=(1024,)/Float32/1024 0.0514 ± 0.038 ms 0.0544 ± 0.036 ms 0.945 ± 0.94
saxpy/static workgroup=(1024,)/Float32/1048576 0.792 ± 0.071 ms 0.799 ± 0.08 ms 0.992 ± 0.13
saxpy/static workgroup=(1024,)/Float32/16384 0.0628 ± 0.006 ms 0.0635 ± 0.0091 ms 0.989 ± 0.17
saxpy/static workgroup=(1024,)/Float32/2048 0.0552 ± 0.0063 ms 0.0549 ± 0.0055 ms 1 ± 0.15
saxpy/static workgroup=(1024,)/Float32/256 11.5 ± 0.31 μs 0.052 ± 0.042 ms 0.221 ± 0.18
saxpy/static workgroup=(1024,)/Float32/262144 0.226 ± 0.026 ms 0.226 ± 0.027 ms 0.998 ± 0.17
saxpy/static workgroup=(1024,)/Float32/32768 0.0735 ± 0.0065 ms 0.0718 ± 0.0066 ms 1.02 ± 0.13
saxpy/static workgroup=(1024,)/Float32/4096 0.0502 ± 0.0079 ms 0.0515 ± 0.0094 ms 0.974 ± 0.24
saxpy/static workgroup=(1024,)/Float32/512 14.3 ± 8.9 μs 23.5 ± 41 μs 0.608 ± 1.1
saxpy/static workgroup=(1024,)/Float32/64 12.1 ± 41 μs 19 ± 42 μs 0.639 ± 2.6
saxpy/static workgroup=(1024,)/Float32/65536 0.095 ± 0.0061 ms 0.098 ± 0.01 ms 0.969 ± 0.12
saxpy/static workgroup=(1024,)/Float64/1024 0.0532 ± 0.0061 ms 0.0538 ± 0.0033 ms 0.989 ± 0.13
saxpy/static workgroup=(1024,)/Float64/1048576 1.22 ± 0.14 ms 1.39 ± 0.22 ms 0.879 ± 0.17
saxpy/static workgroup=(1024,)/Float64/16384 0.0627 ± 0.0071 ms 0.0662 ± 0.0073 ms 0.947 ± 0.15
saxpy/static workgroup=(1024,)/Float64/2048 0.0554 ± 0.0073 ms 0.0555 ± 0.0068 ms 0.999 ± 0.18
saxpy/static workgroup=(1024,)/Float64/256 0.0546 ± 0.039 ms 0.0547 ± 0.0052 ms 0.997 ± 0.72
saxpy/static workgroup=(1024,)/Float64/262144 0.277 ± 0.043 ms 0.29 ± 0.038 ms 0.956 ± 0.19
saxpy/static workgroup=(1024,)/Float64/32768 0.0785 ± 0.0083 ms 0.087 ± 0.011 ms 0.903 ± 0.15
saxpy/static workgroup=(1024,)/Float64/4096 0.0547 ± 0.0089 ms 0.056 ± 0.0077 ms 0.977 ± 0.21
saxpy/static workgroup=(1024,)/Float64/512 0.0543 ± 0.0048 ms 0.0537 ± 0.0067 ms 1.01 ± 0.15
saxpy/static workgroup=(1024,)/Float64/64 11.7 ± 0.56 μs 0.0506 ± 0.041 ms 0.23 ± 0.19
saxpy/static workgroup=(1024,)/Float64/65536 0.106 ± 0.011 ms 0.115 ± 0.018 ms 0.921 ± 0.17
time_to_load 0.504 ± 0.0074 s 0.527 ± 0.01 s 0.956 ± 0.023
main d6c0967... main / d6c0967...
const/@Const/Float32/262144 9 allocs: 0.203 kB 9 allocs: 0.203 kB 1
const/@Const/Float32/65536 9 allocs: 0.203 kB 9 allocs: 0.203 kB 1
const/@Const/Float64/262144 9 allocs: 0.203 kB 9 allocs: 0.203 kB 1
const/@Const/Float64/65536 9 allocs: 0.203 kB 9 allocs: 0.203 kB 1
const/unmarked/Float32/262144 9 allocs: 0.203 kB 9 allocs: 0.203 kB 1
const/unmarked/Float32/65536 9 allocs: 0.203 kB 9 allocs: 0.203 kB 1
const/unmarked/Float64/262144 9 allocs: 0.203 kB 9 allocs: 0.203 kB 1
const/unmarked/Float64/65536 9 allocs: 0.203 kB 9 allocs: 0.203 kB 1
launch/3D static workgroup, dynamic ndrange 9 allocs: 0.219 kB 9 allocs: 0.219 kB 1
launch/3D static workgroup, static ndrange 9 allocs: 0.219 kB 9 allocs: 0.219 kB 1
launch/dynamic workgroup, dynamic ndrange 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
launch/dynamic workgroup, dynamic ndrange, workgroupsize given 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
launch/static workgroup, dynamic ndrange 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
launch/static workgroup, static ndrange 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
partition/dynamic workgroup, dynamic ndrange 2 allocs: 0.0625 kB 2 allocs: 0.0625 kB 1
partition/static workgroup, dynamic ndrange 2 allocs: 32 B 2 allocs: 32 B 1
partition/static workgroup, static ndrange 0 allocs: 0 B 0 allocs: 0 B
saxpy/default/Float16/1024 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float16/1048576 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/default/Float16/16384 12 allocs: 0.25 kB 8 allocs: 0.141 kB 1.78
saxpy/default/Float16/2048 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float16/256 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
saxpy/default/Float16/262144 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/default/Float16/32768 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/default/Float16/4096 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float16/512 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float16/64 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
saxpy/default/Float16/65536 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/default/Float32/1024 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float32/1048576 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/default/Float32/16384 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float32/2048 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float32/256 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
saxpy/default/Float32/262144 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/default/Float32/32768 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float32/4096 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float32/512 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float32/64 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
saxpy/default/Float32/65536 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/default/Float64/1024 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float64/1048576 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/default/Float64/16384 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float64/2048 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float64/256 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
saxpy/default/Float64/262144 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/default/Float64/32768 8 allocs: 0.141 kB 12 allocs: 0.25 kB 0.562
saxpy/default/Float64/4096 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float64/512 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float64/64 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
saxpy/default/Float64/65536 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/static workgroup=(1024,)/Float16/1024 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float16/1048576 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/static workgroup=(1024,)/Float16/16384 8 allocs: 0.141 kB 12 allocs: 0.25 kB 0.562
saxpy/static workgroup=(1024,)/Float16/2048 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float16/256 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
saxpy/static workgroup=(1024,)/Float16/262144 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/static workgroup=(1024,)/Float16/32768 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/static workgroup=(1024,)/Float16/4096 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float16/512 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float16/64 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
saxpy/static workgroup=(1024,)/Float16/65536 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/static workgroup=(1024,)/Float32/1024 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float32/1048576 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/static workgroup=(1024,)/Float32/16384 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float32/2048 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float32/256 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
saxpy/static workgroup=(1024,)/Float32/262144 8 allocs: 0.141 kB 12 allocs: 0.25 kB 0.562
saxpy/static workgroup=(1024,)/Float32/32768 12 allocs: 0.25 kB 8 allocs: 0.141 kB 1.78
saxpy/static workgroup=(1024,)/Float32/4096 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float32/512 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float32/64 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
saxpy/static workgroup=(1024,)/Float32/65536 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/static workgroup=(1024,)/Float64/1024 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float64/1048576 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/static workgroup=(1024,)/Float64/16384 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float64/2048 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float64/256 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
saxpy/static workgroup=(1024,)/Float64/262144 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/static workgroup=(1024,)/Float64/32768 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/static workgroup=(1024,)/Float64/4096 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float64/512 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float64/64 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
saxpy/static workgroup=(1024,)/Float64/65536 12 allocs: 0.25 kB 8 allocs: 0.141 kB 1.78
time_to_load 0.2 k allocs: 11.8 kB 0.2 k allocs: 11.8 kB 1

Benchmark Plots

A plot of the benchmark results have been uploaded as an artifact to the workflow run for this PR.
Go to "Actions"->"Benchmark a pull request"->[the most recent run]->"Artifacts" (at the bottom).

@vchuravy
vchuravy added this pull request to stack #838 October 4, 2026 08:38
Tracers are global, so `@profile` recorded every task in the process,
including other profiles. It now records only the task running its
expression and the tasks spawned from it, tracked with a ScopedValue,
numbers those tasks, and nests the trace per task rather than per
thread. It warns about ranges still open when it finishes, i.e. tasks
it wasn't waited for.

Assisted-by: Claude Code (Opus 5.5)
Each task is a range from `@spawn` until its queued work has completed,
named after the call site or the new `name = ...` argument. The range
starts in the spawning task, so `@profile` warns about a task it wasn't
waited for even if the task hasn't run yet. `@profile` attributes a
range to the task it ended on, which puts a spawn's range at the root of
the spawned task.

Assisted-by: Claude Code (Opus 5.5)
- `@profiling_range` only enters the exception handler that ends the
  range when a profiler listens, compiling `expr` twice. This takes the
  cost of an annotated call from +6 ns to +0.3 ns when tracing is off.
  `@goto` and `@label` are therefore not supported in `expr`.
- `profiling_mark` inlines the check for a profiler (+3.6 ns to +0 ns).
- The traced launch path is inferred once with `@nospecializeinfer`,
  rather than along with the launch of every new kernel, which cost
  about 1.8 MiB of compiler allocations on each first launch.

Assisted-by: Claude Code (Opus 5.5)
- Labels written as literals (in `@profiling_range` and `@spawn`), and
  kernel names, are passed to tracers as `Symbol`s. They are fixed in the
  code, so tracers can cache what they derive from them by identity;
  labels computed at run time stay `String`s. Domains are `Symbol`s.
- The NVTX tracer registers `Symbol` labels with NVTX once, and caches
  domains by `Symbol` (a range from ~480 to ~410 ns under nsys).
- The NVTXT tracer formats integers into a reused buffer instead of
  interpolating a string per record, and caches the messages of `Symbol`
  labels (a range from ~720 ns and 1.8 KB to ~400 ns and 160 B).
- Kernel labels are computed once per kernel type, and the traced launch
  calls the same `launch_untraced` as the untraced one, rather than a
  dynamically dispatched copy.

Assisted-by: Claude Code (Opus 5.5)
The label is a constant computed from the kernel's type, and the helper
that starts the range is specialized again, now that the traced launch
reuses the inferred `launch_untraced`. Starting and ending the range goes
from ~310 to ~28 ns, i.e. a launch with a tracer registered costs ~70 ns
more than one without, rather than ~500 ns.

Assisted-by: Claude Code (Opus 5.5)
They take the kernel's label, a constant, rather than the kernel, so only
the call sites are specialized on the kernel.

Assisted-by: Claude Code (Opus 5.5)
The traced launch is inferred and compiled once for all kernels, rather
than as part of every kernel's launch, which added ~0.6 MiB and a few ms
to the compilation of every new kernel. A traced launch pays for a
dynamic dispatch to the untraced one instead (~50 ns).

Assisted-by: Claude Code (Opus 5.5)

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant