Skip to content

Add GPT-OSS mixed-width CUDA QMoE recipe - #636

Open
David Fan (jiafatom) wants to merge 4 commits into
mainfrom
jiafa/add-gpt-oss-mixed-width-qmoe
Open

David Fan (jiafatom) wants to merge 4 commits into
mainfrom
jiafa/add-gpt-oss-mixed-width-qmoe

Conversation

@jiafatom

Copy link
Copy Markdown
Contributor

Summary

  • add a GPT-OSS-20B CUDA recipe with INT4 dense weights
  • requantize QMoE gate/up projections to INT2 and down projections to INT4
  • use symmetric blockwise expert quantization with block size 64

Depends on:

Validation

  • shell syntax passes with bash -n
  • required export options are present exactly once
  • git diff --check passes

Copilot AI lite review requested due to automatic review settings September 25, 2026 21:07
Comment thread gpt-oss-20b/int4_cuda_int2_int4_qmoe/README.md Fixed
Comment thread gpt-oss-20b/int4_cuda_int2_int4_qmoe/gpt-oss-20b.sh Fixed

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

The export command has invalid arguments, and required recipe metadata and executable handling are incomplete.

Get a fresh assessment by requesting another Copilot review.

Review effort: Lite
Findings: 2 High severity · 1 Medium severity

Open (3)
What changed in this PR

Adds a GPT-OSS-20B CUDA recipe with INT4 dense weights and mixed INT2/INT4 QMoE expert quantization.

Changes:

  • Documents export prerequisites and inference usage.
  • Adds the Olive export script with block size 64 and projection-specific quantization.
File Description
gpt-oss-20b/​int4_cuda_int2_int4_qmoe/​README.md Documents the recipe, export, and execution.
gpt-oss-20b/​int4_cuda_int2_int4_qmoe/​gpt-oss-20b.sh Defines the mixed-width CUDA export command.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread gpt-oss-20b/int4_cuda_int2_int4_qmoe/gpt-oss-20b.sh
Comment thread gpt-oss-20b/int4_cuda_int2_int4_qmoe/gpt-oss-20b.sh Outdated
Comment thread gpt-oss-20b/int4_cuda_int2_int4_qmoe/README.md
@jiafatom
David Fan (jiafatom) force-pushed the jiafa/add-gpt-oss-mixed-width-qmoe branch from d76be2a to 97be9ad Compare September 28, 2026 03:57
## Export

```bash
bash gpt-oss-20b.sh

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks like this script is just a CLI, can we just move the CLI command to this README directly?

@@ -1,3 +1,5 @@
--extra-index-url https://aiinfra.pkgs.visualstudio.com/PublicPackages/_packaging/ORT-Nightly/pypi/simple
accelerate
kernels>=0.16,<0.17

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

what is this used for?

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants