Skip to content

Use matrix kernels for global-scale gather_qqmm - #4481

Open
dhiltgen wants to merge 1 commit into
ml-explore:mainfrom
dhiltgen:g-qqmm-dispatch
Open

Use matrix kernels for global-scale gather_qqmm#4481
dhiltgen wants to merge 1 commit into
ml-explore:mainfrom
dhiltgen:g-qqmm-dispatch

Conversation

@dhiltgen

@dhiltgen dhiltgen commented Sep 8, 2026

Copy link
Copy Markdown
Contributor

This wires up the new gather_qmm kernels to optimize gather_qqmm with global scale.

Performance

Using the NVIDIA Model Optimizer on Qwen/Qwen3.6-35B-A3B with nvfp4_mlp_only p2048/g128

main prompt tps this PR prompt tps main generation tps this PR generation tps
M5 Max 1,257.0 3,105.3 81.9 81.7
M3 Ultra 1,627.3 2,673.2 68.5 68.6
  • ☑️ I understand it is strictly prohibited to use AI to write PR description
  • AI usage disclosure: co-developed with coding agent

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant