Repository navigation
Conversation
`@profiling_range` and `profiling_mark` put named ranges and markers on the timeline of a tracing profiler, and kernel launches are annotated with the kernel's name. Profilers like NVTX, ITT and roctx annotate host threads and correlate device work themselves, so ranges go to every registered `Tracer` regardless of backend. With none registered, an annotation costs one atomic load. Tracers: NVTX (NVTXExt), roctx (ROCTXExt, via AMDGPU), Intel ITT (IntelITTExt), each registering only under its profiler, and a built-in NVTXT writer enabled with `JULIA_KA_NVTXT`. Supersedes #703 and #66. Assisted-by: Claude Code (Opus 5.5)
vchuravy
commented
Oct 4, 2026
Comment on lines
+3
to
+6
| KernelAbstractions can put named ranges on the timeline of a tracing profiler, such as | ||
| NVIDIA Nsight Systems or Intel VTune, so that you can see which part of your program a | ||
| stretch of kernels belongs to. Annotations are cheap when no profiler is listening: a | ||
| single atomic load, and the label isn't even built. |
Member
Author
There was a problem hiding this comment.
Suggested change
| KernelAbstractions can put named ranges on the timeline of a tracing profiler, such as | |
| NVIDIA Nsight Systems or Intel VTune, so that you can see which part of your program a | |
| stretch of kernels belongs to. Annotations are cheap when no profiler is listening: a | |
| single atomic load, and the label isn't even built. | |
| KernelAbstractions can put named ranges on the timeline of a tracing profiler, such as | |
| NVIDIA Nsight Systems or Intel VTune, so that you can see which part of your program a | |
| stretch of kernels belongs to. Annotations are cheap when no profiler is listening. |
Comment on lines
+33
to
+37
| Ranges are recorded on the host threads of the process, which is how NVTX, ITT and | ||
| roctx work: it is the profiler that attributes the device work launched within a range to | ||
| it. So which profiler records the ranges depends on what the process runs under, not on the | ||
| backend: running the CPU backend under Nsight Systems gives NVTX ranges, and a GPU backend | ||
| under VTune gives ITT tasks. |
Member
Author
There was a problem hiding this comment.
Suggested change
| Ranges are recorded on the host threads of the process, which is how NVTX, ITT and | |
| roctx work: it is the profiler that attributes the device work launched within a range to | |
| it. So which profiler records the ranges depends on what the process runs under, not on the | |
| backend: running the CPU backend under Nsight Systems gives NVTX ranges, and a GPU backend | |
| under VTune gives ITT tasks. | |
| Ranges are recorded on the host threads of the process, which is how NVTX, ITT and | |
| roctx work: it is the profiler that attributes the device work launched within a range to | |
| it. |
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## main #836 +/- ##
==========================================
- Coverage 79.06% 76.32% -2.75%
==========================================
Files 24 29 +5
Lines 2040 2547 +507
==========================================
+ Hits 1613 1944 +331
- Misses 427 603 +176 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
Records the ranges, markers and kernel launches of an expression with a temporary tracer, and summarizes them per name or, with `trace = true`, lists them in order. By default kernel launches synchronize their backend while profiling, so that their ranges measure execution rather than launch; tracers opt into this with `synchronizes_launches`. Assisted-by: Claude Code (Opus 5.5)
Contributor
Benchmark ResultsShow table
Benchmark PlotsA plot of the benchmark results have been uploaded as an artifact to the workflow run for this PR. |
vchuravy
added this pull request to stack #838
October 4, 2026 08:38
Tracers are global, so `@profile` recorded every task in the process, including other profiles. It now records only the task running its expression and the tasks spawned from it, tracked with a ScopedValue, numbers those tasks, and nests the trace per task rather than per thread. It warns about ranges still open when it finishes, i.e. tasks it wasn't waited for. Assisted-by: Claude Code (Opus 5.5)
Each task is a range from `@spawn` until its queued work has completed, named after the call site or the new `name = ...` argument. The range starts in the spawning task, so `@profile` warns about a task it wasn't waited for even if the task hasn't run yet. `@profile` attributes a range to the task it ended on, which puts a spawn's range at the root of the spawned task. Assisted-by: Claude Code (Opus 5.5)
- `@profiling_range` only enters the exception handler that ends the range when a profiler listens, compiling `expr` twice. This takes the cost of an annotated call from +6 ns to +0.3 ns when tracing is off. `@goto` and `@label` are therefore not supported in `expr`. - `profiling_mark` inlines the check for a profiler (+3.6 ns to +0 ns). - The traced launch path is inferred once with `@nospecializeinfer`, rather than along with the launch of every new kernel, which cost about 1.8 MiB of compiler allocations on each first launch. Assisted-by: Claude Code (Opus 5.5)
- Labels written as literals (in `@profiling_range` and `@spawn`), and kernel names, are passed to tracers as `Symbol`s. They are fixed in the code, so tracers can cache what they derive from them by identity; labels computed at run time stay `String`s. Domains are `Symbol`s. - The NVTX tracer registers `Symbol` labels with NVTX once, and caches domains by `Symbol` (a range from ~480 to ~410 ns under nsys). - The NVTXT tracer formats integers into a reused buffer instead of interpolating a string per record, and caches the messages of `Symbol` labels (a range from ~720 ns and 1.8 KB to ~400 ns and 160 B). - Kernel labels are computed once per kernel type, and the traced launch calls the same `launch_untraced` as the untraced one, rather than a dynamically dispatched copy. Assisted-by: Claude Code (Opus 5.5)
The label is a constant computed from the kernel's type, and the helper that starts the range is specialized again, now that the traced launch reuses the inferred `launch_untraced`. Starting and ending the range goes from ~310 to ~28 ns, i.e. a launch with a tracer registered costs ~70 ns more than one without, rather than ~500 ns. Assisted-by: Claude Code (Opus 5.5)
They take the kernel's label, a constant, rather than the kernel, so only the call sites are specialized on the kernel. Assisted-by: Claude Code (Opus 5.5)
The traced launch is inferred and compiled once for all kernels, rather than as part of every kernel's launch, which added ~0.6 MiB and a few ms to the compilation of every new kernel. A traced launch pays for a dynamic dispatch to the untraced one instead (~50 ns). Assisted-by: Claude Code (Opus 5.5)
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds a low-cost tracing subsystem that puts named ranges on the timeline of a tracing profiler. This supersedes #703 and #66.
Kernel launches are also recorded as ranges, named after the kernel.
Design
NVTX, ITT and roctx annotate host threads, and the profiler attributes the device work launched within a range to that range. So which profiler records a range depends on what the process runs under, not on the backend. For example, running the CPU backend under Nsight Systems gives NVTX ranges, and a GPU backend under VTune gives ITT tasks. For this reason the API takes no backend argument, and there are no KernelInterface hooks. Every range goes to all registered
KernelAbstractions.Tracers.@profiling_rangeuses a scope-free:tryfinally, like@time. The range is ended if the expression throws, and assignments inside it remain visible afterwards.Tracers
Each tracer registers itself only when its profiler is attached:
NVTXExtNVTX.isactive(), i.e. undernsysROCTXExtrocprofv3(ROCP_TOOL_LIBRARIES) or legacyrocprof(HSA_TOOLS_LIB)IntelITTExtIntelITT.isactive(), i.e. under VTuneNVTXTTracerJULIA_KA_NVTXT=1or=path-%p.nvtxtNVTXTTracerrevives #66. It writes the NVTXT text format, which Nsight Systems can import withImportNvtxt, so you can trace without a profiler attached.Testing
nsys profile --trace=nvtx:nvtx_sumshowsDemo:step 1andKernelAbstractions:scale!ranges./opt/rocm/lib/libroctx64.so(ROCm 7.2) using a stand-inAMDGPUmodule. I could not test it underrocprofv3, which isn't installed.🤖 Generated with Claude Code