Skip to content

[FIX] Report the profiler's buffer-load issue only for AMD kernels - #483

Open
mark14wu wants to merge 1 commit into
mainfrom
fix/profiler-buffer-load-amd-only
Open

mark14wu wants to merge 1 commit into
mainfrom
fix/profiler-buffer-load-amd-only

Conversation

@mark14wu

@mark14wu mark14wu commented Oct 3, 2026

Copy link
Copy Markdown
Collaborator

Summary

The profiler's buffer-load check is meant for AMD GPUs only, but it also fires on NVIDIA. Profiler.__init__ set has_buffer_load = False, and post_warmup_callback only updates it when the warmup result has an amdgcn ASM stage. NVIDIA kernels have no amdgcn stage, so the value stayed False. Then _check_32bit_range (if self.has_buffer_load is False) marked every kernel whose offsets fit in 32 bits as a potential buffer-load issue, and the summary printed the AMD "Potential Buffer Load Issue Detected" warning. That warning was a false positive.

Fix

  • has_buffer_load now starts as None ("no AMD kernel seen yet"). Only a captured amdgcn stage sets it to True or False. The existing is False test therefore flags only AMD kernels whose amdgcn stage has no buffer_load.
  • When no amdgcn kernel was captured, the Buffer Load section of the summary now prints Buffer Load check skipped: no AMD (amdgcn) kernel was captured. In that case it shows neither the AMD warning nor the "used appropriately" message.
  • New unit tests in tests/unit/test_profiler.py pass a fake warmup result to post_warmup_callback, then 32-bit offsets to _check_32bit_range:
    • an NVIDIA-like result (ttir/ptx stages, no amdgcn) is not flagged, and the summary shows the skipped message;
    • an amdgcn stage without buffer_load is flagged;
    • an amdgcn stage with buffer_load is not flagged.

Testing

Run on an RTX 4090 with the GPU visible. I first confirmed that the worktree's tilelens is the one imported.

python -m pytest tests/unit/test_profiler.py tests/end_to_end/test_profiler.py -q -p no:cacheprovider
26 passed in 1.81s

python -m pytest tests/unit/test_profiler.py -p no:cacheprovider -k buffer_load -rA
4 passed, 12 deselected

pre-commit run --files tilelens/clients/profiler/profiler.py tests/unit/test_profiler.py
all hooks Passed/Skipped (ruff, ruff-format, mypy, codespell, ...)

The buffer-load check is meant for AMD GPUs, but has_buffer_load started
as False and was only updated from an amdgcn ASM stage. On NVIDIA no
amdgcn stage exists, so every kernel whose offsets fit in 32 bits was
flagged as a potential buffer-load issue and the summary printed the AMD
warning.

has_buffer_load now starts as None and is set only from a captured
amdgcn stage, so the existing `is False` test flags only AMD kernels
without buffer_load. When no amdgcn kernel was captured, the summary
says the check was skipped. Unit tests cover the NVIDIA-like, AMD
without buffer_load and AMD with buffer_load cases.
@github-actions

github-actions Bot commented Oct 3, 2026 •

Copy link
Copy Markdown

Performance Benchmark

Benchmark main (min) PR (min) Change Samples
gemm 0.074s 0.075s +0.4% 20 / 20
gemm_oob 0.084s 0.083s -0.4% 20 / 20
indirect_load 0.015s 0.015s -1.1% 20 / 20
nested_loop 0.156s 0.155s -0.3% 20 / 20
block_pointer_loop_advance 0.088s 0.088s +0.3% 20 / 20
liger_jsd 0.103s 0.103s +0.3% 20 / 20
flaggems_layernorm 0.269s 0.271s +0.7% 20 / 20
swiglu 0.125s 0.125s +0.5% 20 / 20
cross_entropy 0.702s 0.703s +0.2% 20 / 20
fused_linear_jsd 0.156s 0.157s +0.4% 20 / 20
Total 1.771s 1.775s +0.2% N/A

Iterations: 1 warmup + 20 measured
Samples are shown as main / PR; long pytest benchmarks may use fewer samples.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant