Skip to content

docs: measure the LFM2.5 text recipes against llama.cpp - #638

Open
Yuri Khrustalev (ykhrustalev) wants to merge 2 commits into
microsoft:mainfrom
ykhrustalev:ykhrustalev/measure-lfm2-5-text-recipes-against-llama-cpp
Open

Yuri Khrustalev (ykhrustalev) wants to merge 2 commits into
microsoft:mainfrom
ykhrustalev:ykhrustalev/measure-lfm2-5-text-recipes-against-llama-cpp

Conversation

@ykhrustalev

Copy link
Copy Markdown
Contributor

Problem
The LFM2.5 text READMEs call the int4 and int8 recipes "Q4_K_M equivalent" and "Q8_0 equivalent" without numbers, and call the WebGPU fp16_int4 recipe the most accurate. Measured against llama.cpp, int4 trails Q4_K_M on most models and fp16_int4 is the least accurate recipe. The recipes also fail to run in a fresh environment set up as the READMEs say.

Solution

Testing

  • Scoring: FP32 reference from PyTorch on 64 × 512-token wikitext-2 chunks, GGUFs through llama-perplexity --kl-divergence; ONNX CPU and WebGPU on an M3 Ultra, CUDA on an A10
  • Built 230M and 1.2B-Instruct on CPU, CUDA and WebGPU in fresh environments set up as the READMEs say; every package answers a chat prompt through onnxruntime-genai (0.17.0 on CPU and WebGPU, a CUDA 12 build of the same code on the A10)

microsoft/Olive#2692 and microsoft/onnxruntime#32814 close most of the int4 size and quality gap; the tables need a refresh once they ship.

Every CPU, CUDA and WebGPU README gets KL divergence and top-1 agreement against the
FP32 model next to LiquidAI's Q4_K_M and Q8_0 GGUFs, replacing the "equivalent"
labels, and the CUDA READMEs note that PyPI's CUDA wheels need driver 580+. The
WebGPU fp16_int4 recipe is the least accurate on all four models, not the most.
Olive main imports requests at startup without declaring it (microsoft/Olive#2676),
so olive run failed in a fresh environment set up as the READMEs say.
Copilot AI lite review requested due to automatic review settings September 26, 2026 13:35

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

One or more issues must be addressed before approval.

Review effort: Lite
Findings: None

What changed in this PR

Documents measured quantization quality for LFM2.5 recipes and fixes fresh-environment setup dependencies.

Changes:

  • Adds KLD, top-1 agreement, size, and comparison tables to CPU, CUDA, and WebGPU READMEs.
  • Corrects recipe accuracy descriptions and adds CUDA driver requirements.
  • Adds missing requests dependency to all recipe environments.
File Description
LiquidAI-LFM2.5-350M/​webgpu/​requirements.txt Updated as part of this pull request.
LiquidAI-LFM2.5-350M/​webgpu/​README.md Updated as part of this pull request.
LiquidAI-LFM2.5-350M/​cuda/​requirements.txt Updated as part of this pull request.
LiquidAI-LFM2.5-350M/​cuda/​README.md Updated as part of this pull request.
LiquidAI-LFM2.5-350M/​cpu/​requirements.txt Updated as part of this pull request.
LiquidAI-LFM2.5-350M/​cpu/​README.md Updated as part of this pull request.
LiquidAI-LFM2.5-230M/​webgpu/​requirements.txt Updated as part of this pull request.
LiquidAI-LFM2.5-230M/​webgpu/​README.md Updated as part of this pull request.
LiquidAI-LFM2.5-230M/​cuda/​requirements.txt Updated as part of this pull request.
LiquidAI-LFM2.5-230M/​cuda/​README.md Updated as part of this pull request.
LiquidAI-LFM2.5-230M/​cpu/​requirements.txt Updated as part of this pull request.
LiquidAI-LFM2.5-230M/​cpu/​README.md Updated as part of this pull request.
LiquidAI-LFM2.5-2.6B/​webgpu/​requirements.txt Updated as part of this pull request.
LiquidAI-LFM2.5-2.6B/​webgpu/​README.md Updated as part of this pull request.
LiquidAI-LFM2.5-2.6B/​cuda/​requirements.txt Updated as part of this pull request.
LiquidAI-LFM2.5-2.6B/​cuda/​README.md Updated as part of this pull request.
LiquidAI-LFM2.5-2.6B/​cpu/​requirements.txt Updated as part of this pull request.
LiquidAI-LFM2.5-2.6B/​cpu/​README.md Updated as part of this pull request.
LiquidAI-LFM2.5-1.2B-Instruct/​webgpu/​requirements.txt Updated as part of this pull request.
LiquidAI-LFM2.5-1.2B-Instruct/​webgpu/​README.md Updated as part of this pull request.
LiquidAI-LFM2.5-1.2B-Instruct/​cuda/​requirements.txt Updated as part of this pull request.
LiquidAI-LFM2.5-1.2B-Instruct/​cuda/​README.md Updated as part of this pull request.
LiquidAI-LFM2.5-1.2B-Instruct/​cpu/​requirements.txt Updated as part of this pull request.
LiquidAI-LFM2.5-1.2B-Instruct/​cpu/​README.md Updated as part of this pull request.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants