docs: measure the LFM2.5 text recipes against llama.cpp - #638
Open
Yuri Khrustalev (ykhrustalev) wants to merge 2 commits into
Open
Yuri Khrustalev (ykhrustalev) wants to merge 2 commits into
Yuri Khrustalev (ykhrustalev) wants to merge 2 commits into
Conversation
Every CPU, CUDA and WebGPU README gets KL divergence and top-1 agreement against the FP32 model next to LiquidAI's Q4_K_M and Q8_0 GGUFs, replacing the "equivalent" labels, and the CUDA READMEs note that PyPI's CUDA wheels need driver 580+. The WebGPU fp16_int4 recipe is the least accurate on all four models, not the most.
Olive main imports requests at startup without declaring it (microsoft/Olive#2676), so olive run failed in a fresh environment set up as the READMEs say.
Copilot started reviewing on behalf of
Yuri Khrustalev (ykhrustalev)
September 26, 2026 13:35
View session
Contributor
There was a problem hiding this comment.
Copilot review overview
🔵 Needs a closer look
One or more issues must be addressed before approval.
Review effort: Lite
Findings: None
What changed in this PR
Documents measured quantization quality for LFM2.5 recipes and fixes fresh-environment setup dependencies.
Changes:
- Adds KLD, top-1 agreement, size, and comparison tables to CPU, CUDA, and WebGPU READMEs.
- Corrects recipe accuracy descriptions and adds CUDA driver requirements.
- Adds missing
requestsdependency to all recipe environments.
| File | Description |
|---|---|
| LiquidAI-LFM2.5-350M/webgpu/requirements.txt | Updated as part of this pull request. |
| LiquidAI-LFM2.5-350M/webgpu/README.md | Updated as part of this pull request. |
| LiquidAI-LFM2.5-350M/cuda/requirements.txt | Updated as part of this pull request. |
| LiquidAI-LFM2.5-350M/cuda/README.md | Updated as part of this pull request. |
| LiquidAI-LFM2.5-350M/cpu/requirements.txt | Updated as part of this pull request. |
| LiquidAI-LFM2.5-350M/cpu/README.md | Updated as part of this pull request. |
| LiquidAI-LFM2.5-230M/webgpu/requirements.txt | Updated as part of this pull request. |
| LiquidAI-LFM2.5-230M/webgpu/README.md | Updated as part of this pull request. |
| LiquidAI-LFM2.5-230M/cuda/requirements.txt | Updated as part of this pull request. |
| LiquidAI-LFM2.5-230M/cuda/README.md | Updated as part of this pull request. |
| LiquidAI-LFM2.5-230M/cpu/requirements.txt | Updated as part of this pull request. |
| LiquidAI-LFM2.5-230M/cpu/README.md | Updated as part of this pull request. |
| LiquidAI-LFM2.5-2.6B/webgpu/requirements.txt | Updated as part of this pull request. |
| LiquidAI-LFM2.5-2.6B/webgpu/README.md | Updated as part of this pull request. |
| LiquidAI-LFM2.5-2.6B/cuda/requirements.txt | Updated as part of this pull request. |
| LiquidAI-LFM2.5-2.6B/cuda/README.md | Updated as part of this pull request. |
| LiquidAI-LFM2.5-2.6B/cpu/requirements.txt | Updated as part of this pull request. |
| LiquidAI-LFM2.5-2.6B/cpu/README.md | Updated as part of this pull request. |
| LiquidAI-LFM2.5-1.2B-Instruct/webgpu/requirements.txt | Updated as part of this pull request. |
| LiquidAI-LFM2.5-1.2B-Instruct/webgpu/README.md | Updated as part of this pull request. |
| LiquidAI-LFM2.5-1.2B-Instruct/cuda/requirements.txt | Updated as part of this pull request. |
| LiquidAI-LFM2.5-1.2B-Instruct/cuda/README.md | Updated as part of this pull request. |
| LiquidAI-LFM2.5-1.2B-Instruct/cpu/requirements.txt | Updated as part of this pull request. |
| LiquidAI-LFM2.5-1.2B-Instruct/cpu/README.md | Updated as part of this pull request. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
The LFM2.5 text READMEs call the int4 and int8 recipes "Q4_K_M equivalent" and "Q8_0 equivalent" without numbers, and call the WebGPU fp16_int4 recipe the most accurate. Measured against llama.cpp, int4 trails Q4_K_M on most models and fp16_int4 is the least accurate recipe. The recipes also fail to run in a fresh environment set up as the READMEs say.
Solution
requests, which Olivemainimports without declaring ([Bug]: import olive fails out of the box on 0.13.0 - the telemetry exporter imports requests, which is not a declared dependency Olive#2676)Testing
llama-perplexity --kl-divergence; ONNX CPU and WebGPU on an M3 Ultra, CUDA on an A10microsoft/Olive#2692 and microsoft/onnxruntime#32814 close most of the int4 size and quality gap; the tables need a refresh once they ship.