Repository navigation
Conversation
The buffer-load check is meant for AMD GPUs, but has_buffer_load started as False and was only updated from an amdgcn ASM stage. On NVIDIA no amdgcn stage exists, so every kernel whose offsets fit in 32 bits was flagged as a potential buffer-load issue and the summary printed the AMD warning. has_buffer_load now starts as None and is set only from a captured amdgcn stage, so the existing `is False` test flags only AMD kernels without buffer_load. When no amdgcn kernel was captured, the summary says the check was skipped. Unit tests cover the NVIDIA-like, AMD without buffer_load and AMD with buffer_load cases.
Performance Benchmark
Iterations: 1 warmup + 20 measured |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
The profiler's buffer-load check is meant for AMD GPUs only, but it also fires on NVIDIA.
Profiler.__init__sethas_buffer_load = False, andpost_warmup_callbackonly updates it when the warmup result has anamdgcnASM stage. NVIDIA kernels have noamdgcnstage, so the value stayedFalse. Then_check_32bit_range(if self.has_buffer_load is False) marked every kernel whose offsets fit in 32 bits as a potential buffer-load issue, and the summary printed the AMD "Potential Buffer Load Issue Detected" warning. That warning was a false positive.Fix
has_buffer_loadnow starts asNone("no AMD kernel seen yet"). Only a capturedamdgcnstage sets it toTrueorFalse. The existingis Falsetest therefore flags only AMD kernels whoseamdgcnstage has nobuffer_load.amdgcnkernel was captured, the Buffer Load section of the summary now printsBuffer Load check skipped: no AMD (amdgcn) kernel was captured.In that case it shows neither the AMD warning nor the "used appropriately" message.tests/unit/test_profiler.pypass a fake warmup result topost_warmup_callback, then 32-bit offsets to_check_32bit_range:ttir/ptxstages, noamdgcn) is not flagged, and the summary shows the skipped message;amdgcnstage withoutbuffer_loadis flagged;amdgcnstage withbuffer_loadis not flagged.Testing
Run on an RTX 4090 with the GPU visible. I first confirmed that the worktree's
tilelensis the one imported.