Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
33 changes: 17 additions & 16 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -52,22 +52,23 @@ For more details, please see our [website](https://routeworks.github.io/leaderbo
| 13 | [Hybrid Router]() | 👤 [@mikemao27](https://github.com/mikemao27) | 72.08 | 71.38 | $0.04 | 89.87 | 94.19 | 92.81 | — | 96.67 |
| 14 | [R2-Router](https://arxiv.org/abs/2602.02823/) | 🎓 UCF | 71.60 | 71.23 | $0.06 | 24.51 | 48.70 | 99.85 | — | 45.71 |
| 15 | [LLM Router](https://github.com/ypollak2/llm-router) [[PyPI]](https://pypi.org/project/llm-routing/) | 👤 [@ypollak2](https://github.com/ypollak2) | 71.26 | 72.05 | $0.20 | 18.01 | 20.46 | 89.13 | — | 30.00 |
| 16 | [chuzom-solo-v32]() | 👤 [@ypollak2](https://github.com/ypollak2) | 70.61 | 70.59 | $0.10 | — | — | — | — | 100.00 |
| 17 | [Azure-Model-Router](https://ai.azure.com/catalog/models/model-router) [[Web]](https://learn.microsoft.com/en-us/azure/ai-foundry/openai/concepts/model-router) | 💼 Microsoft | 70.42 | 72.94 | $0.73 | — | — | — | — | 71.43 |
| 18 | [Auto Router]() | 👤 [@cxf2015](https://github.com/cxf2015) | 70.05 | 70.17 | $0.12 | 37.58 | 40.02 | 86.04 | — | 49.52 |
| 19 | [Lynkr]() | 👤 [@vishalveerareddy123](https://github.com/vishalveerareddy123) | 67.65 | 68.41 | $0.29 | 10.97 | 16.08 | 84.48 | — | 92.38 |
| 20 | [MIRT‑BERT](https://arxiv.org/pdf/2506.01048) [[Code]](https://github.com/Mercidaiha/IRT-Router) | 🎓 USTC | 66.89 | 66.88 | $0.15 | 3.44 | 19.62 | 78.18 | 27.03 | 61.19 |
| 21 | [NIRT‑BERT](https://arxiv.org/pdf/2506.01048) [[Code]](https://github.com/Mercidaiha/IRT-Router) | 🎓 USTC | 66.12 | 66.34 | $0.21 | 3.83 | 14.04 | 77.88 | 10.42 | 49.29 |
| 22 | [AsiaInfo-Router]() | 👤 [@Uncle-LL](https://github.com/Uncle-LL) | 65.87 | 75.20 | $8.54 | — | — | — | — | 69.52 |
| 23 | [GPT‑5](https://openai.com/index/introducing-gpt-5/) | 💼 OpenAI | 64.32 | 73.96 | $10.02 | — | — | — | — | — |
| 24 | [CARROT](https://arxiv.org/abs/2502.03261) [[Code]](https://github.com/somerstep/CARROT) [[HF]](https://huggingface.co/CARROT-LLM-Routing) | 🎓 UMich | 63.87 | 67.21 | $2.06 | 2.68 | 6.77 | 78.63 | 1.50 | 89.05 |
| 25 | [Chayan](https://huggingface.co/adaptive-classifier/chayan) [[HF]](https://huggingface.co/adaptive-classifier/chayan) | 🎓 Adaptive Classifier | 63.83 | 64.89 | $0.56 | 43.03 | 43.75 | 88.74 | — | — |
| 26 | [RouterBench‑MLP](https://arxiv.org/pdf/2403.12031) [[Code]](https://github.com/withmartian/routerbench) [[HF]](https://huggingface.co/datasets/withmartian/routerbench) | 🎓 Martian | 57.56 | 61.62 | $4.83 | 13.39 | 24.45 | 83.32 | 90.91 | 80.00 |
| 27 | [NotDiamond](https://www.notdiamond.ai/) | 💼 NotDiamond | 57.29 | 60.83 | $4.10 | 1.55 | 2.14 | 76.81 | — | 55.91 |
| 28 | [GraphRouter](https://arxiv.org/abs/2410.03834) [[Code]](https://github.com/ulab-uiuc/GraphRouter) | 🎓 UIUC | 57.22 | 57.00 | $0.34 | 4.73 | 38.33 | 74.25 | 2.70 | 94.29 |
| 29 | [RouterBench‑KNN](https://arxiv.org/pdf/2403.12031) [[Code]](https://github.com/withmartian/routerbench) [[HF]](https://huggingface.co/datasets/withmartian/routerbench) | 🎓 Martian | 55.48 | 58.69 | $4.27 | 13.09 | 25.49 | 78.77 | 1.33 | 83.33 |
| 30 | [RouteLLM](https://arxiv.org/abs/2406.18665) [[Code]](https://github.com/lm-sys/RouteLLM) [[HF]](https://huggingface.co/routellm) | 🎓 Berkeley | 48.07 | 47.04 | $0.27 | 99.72 | 99.63 | 68.76 | 0.40 | 100.00 |
| 31 | [RouterDC](https://arxiv.org/abs/2409.19886) [[Code]](https://github.com/shuhao02/RouterDC) | 🎓 SUSTech | 33.75 | 32.01 | $0.07 | 39.84 | 73.00 | 49.05 | 10.75 | 85.24 |
| 16 | [cruq-router]() | 👤 [@nabaruns](https://github.com/nabaruns) | 70.77 | 71.35 | $0.18 | — | — | — | — | 81.67 |
| 17 | [chuzom-solo-v32]() | 👤 [@ypollak2](https://github.com/ypollak2) | 70.61 | 70.59 | $0.10 | — | — | — | — | 100.00 |
| 18 | [Azure-Model-Router](https://ai.azure.com/catalog/models/model-router) [[Web]](https://learn.microsoft.com/en-us/azure/ai-foundry/openai/concepts/model-router) | 💼 Microsoft | 70.42 | 72.94 | $0.73 | — | — | — | — | 71.43 |
| 19 | [Auto Router]() | 👤 [@cxf2015](https://github.com/cxf2015) | 70.05 | 70.17 | $0.12 | 37.58 | 40.02 | 86.04 | — | 49.52 |
| 20 | [Lynkr]() | 👤 [@vishalveerareddy123](https://github.com/vishalveerareddy123) | 67.65 | 68.41 | $0.29 | 10.97 | 16.08 | 84.48 | — | 92.38 |
| 21 | [MIRT‑BERT](https://arxiv.org/pdf/2506.01048) [[Code]](https://github.com/Mercidaiha/IRT-Router) | 🎓 USTC | 66.89 | 66.88 | $0.15 | 3.44 | 19.62 | 78.18 | 27.03 | 61.19 |
| 22 | [NIRT‑BERT](https://arxiv.org/pdf/2506.01048) [[Code]](https://github.com/Mercidaiha/IRT-Router) | 🎓 USTC | 66.12 | 66.34 | $0.21 | 3.83 | 14.04 | 77.88 | 10.42 | 49.29 |
| 23 | [AsiaInfo-Router]() | 👤 [@Uncle-LL](https://github.com/Uncle-LL) | 65.87 | 75.20 | $8.54 | — | — | — | — | 69.52 |
| 24 | [GPT‑5](https://openai.com/index/introducing-gpt-5/) | 💼 OpenAI | 64.32 | 73.96 | $10.02 | — | — | — | — | — |
| 25 | [CARROT](https://arxiv.org/abs/2502.03261) [[Code]](https://github.com/somerstep/CARROT) [[HF]](https://huggingface.co/CARROT-LLM-Routing) | 🎓 UMich | 63.87 | 67.21 | $2.06 | 2.68 | 6.77 | 78.63 | 1.50 | 89.05 |
| 26 | [Chayan](https://huggingface.co/adaptive-classifier/chayan) [[HF]](https://huggingface.co/adaptive-classifier/chayan) | 🎓 Adaptive Classifier | 63.83 | 64.89 | $0.56 | 43.03 | 43.75 | 88.74 | — | — |
| 27 | [RouterBench‑MLP](https://arxiv.org/pdf/2403.12031) [[Code]](https://github.com/withmartian/routerbench) [[HF]](https://huggingface.co/datasets/withmartian/routerbench) | 🎓 Martian | 57.56 | 61.62 | $4.83 | 13.39 | 24.45 | 83.32 | 90.91 | 80.00 |
| 28 | [NotDiamond](https://www.notdiamond.ai/) | 💼 NotDiamond | 57.29 | 60.83 | $4.10 | 1.55 | 2.14 | 76.81 | — | 55.91 |
| 29 | [GraphRouter](https://arxiv.org/abs/2410.03834) [[Code]](https://github.com/ulab-uiuc/GraphRouter) | 🎓 UIUC | 57.22 | 57.00 | $0.34 | 4.73 | 38.33 | 74.25 | 2.70 | 94.29 |
| 30 | [RouterBench‑KNN](https://arxiv.org/pdf/2403.12031) [[Code]](https://github.com/withmartian/routerbench) [[HF]](https://huggingface.co/datasets/withmartian/routerbench) | 🎓 Martian | 55.48 | 58.69 | $4.27 | 13.09 | 25.49 | 78.77 | 1.33 | 83.33 |
| 31 | [RouteLLM](https://arxiv.org/abs/2406.18665) [[Code]](https://github.com/lm-sys/RouteLLM) [[HF]](https://huggingface.co/routellm) | 🎓 Berkeley | 48.07 | 47.04 | $0.27 | 99.72 | 99.63 | 68.76 | 0.40 | 100.00 |
| 32 | [RouterDC](https://arxiv.org/abs/2409.19886) [[Code]](https://github.com/shuhao02/RouterDC) | 🎓 SUSTech | 33.75 | 32.01 | $0.07 | 39.84 | 73.00 | 49.05 | 10.75 | 85.24 |

🎓 Open-source  💼 Closed-source 

Expand Down
9 changes: 9 additions & 0 deletions leaderboard_manifest.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -202,6 +202,15 @@ routers:
github_url: "https://github.com/KT-A-Autonomous-Tech-Team/ModelRouter"
type: "open-source"

- readme_name: "cruq-router"
website_name: "cruq-router"
prediction: "cruq-router"
category_key: "cruq-router"
flip_key: "cruq-router"
website:
affiliation: "@nabaruns"
github_url: "https://github.com/nabaruns"

# --- Externally-evaluated baselines (headline from README; derived data on
# the website is preserved as-is) ---
- readme_name: "MIRT-BERT"
Expand Down
24 changes: 22 additions & 2 deletions model_cost/model_cost.json
Original file line number Diff line number Diff line change
Expand Up @@ -332,8 +332,8 @@
"output_token_price_per_million": 1.5
},
"MiniMax-M3": {
"input_token_price_per_million": 0.60,
"output_token_price_per_million": 2.40
"input_token_price_per_million": 0.6,
"output_token_price_per_million": 2.4
},
"agnes-2.0-flash": {
"input_token_price_per_million": 0.03,
Expand Down Expand Up @@ -363,6 +363,26 @@
"input_token_price_per_million": 0.14,
"output_token_price_per_million": 0.28
},
"openai/gpt-5-mini": {
"input_token_price_per_million": 0.25,
"output_token_price_per_million": 2.0
},
"anthropic/claude-sonnet-4.5": {
"input_token_price_per_million": 3.0,
"output_token_price_per_million": 15.0
},
"openai/gpt-4o-mini": {
"input_token_price_per_million": 0.15,
"output_token_price_per_million": 0.6
},
"google/gemini-2.5-flash-lite": {
"input_token_price_per_million": 0.1,
"output_token_price_per_million": 0.4
},
"google/gemini-2.5-pro": {
"input_token_price_per_million": 1.25,
"output_token_price_per_million": 10.0
},
"google/gemma-4-31b-it": {
"input_token_price_per_million": 0.08,
"output_token_price_per_million": 0.35
Expand Down
1 change: 1 addition & 0 deletions router_inference/config/cruq-router.json
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
{"pipeline_params":{"router_name":"cruq-router","router_cls_name":"CruqSCRouter","models":["qwen/qwen3-235b-a22b-2507","deepseek/deepseek-v4-flash"],"k":4,"tau":0.6,"probe_model":"qwen/qwen3-235b-a22b-2507","escalate_model":"deepseek/deepseek-v4-flash"}}
1 change: 1 addition & 0 deletions router_inference/predictions/cruq-router-robustness.json

Large diffs are not rendered by default.

1 change: 1 addition & 0 deletions router_inference/predictions/cruq-router.json

Large diffs are not rendered by default.

2 changes: 2 additions & 0 deletions router_inference/router/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -10,6 +10,7 @@
from router_inference.router.chuzom_solo_v32 import ChuzomSoloV32Router
from router_inference.router.llm_router import LLMRouter
from router_inference.router.lynkr_router import LynkrRouter
from router_inference.router.cruq_sc_router import CruqSCRouter

__all__ = [
"BaseRouter",
Expand All @@ -19,4 +20,5 @@
"LLMRouter",
"ChuzomSoloV32Router",
"LynkrRouter",
"CruqSCRouter",
]
93 changes: 93 additions & 0 deletions router_inference/router/cruq_sc_router.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,93 @@
# SPDX-FileCopyrightText: Copyright contributors to the cruq.ai project
# SPDX-License-Identifier: Apache-2.0

"""
cruq self-consistency cascade router.

The router sees only the prompt string. It probes the cheapest capable model
(qwen3-235b-a22b) K times at temperature>0 and measures self-consistency: the
fraction of the K samples that agree on the final \\boxed{} answer. High
agreement is a strong, prompt-independent confidence signal, so:

- consistency >= tau -> keep qwen's majority-vote answer (cost = K cheap probes)
- consistency < tau -> escalate to deepseek-v4-flash (a stronger model)

This is an *inference-time* signal, not a prompt classifier. Across the whole
research programme, no prompt-only selector (lexical difficulty, per-model
P(correct) heads, calibrated thresholds, or a domain classifier) beat the best
single model, because per-query difficulty is close to unpredictable from prompt
text. Model-agreement at inference is what breaks that ceiling.

Compliance: nothing here is fit on RouterArena data. K and tau are priors set in
the config; the probe model and escalation model are fixed choices. No labels,
metadata, or ground truth are read at decision time.

Reproducibility: the K qwen probes per query are cached to
phase2/data/qwen_sc_full.jsonl by phase2/sample_qwen_sc_full.py. This class reads
that cache to make the same keep/escalate decision the submitted prediction file
encodes. The final prediction file (with the majority-vote answer and honest
K-probe token accounting) is assembled by phase2/build_sc_submission.py.
"""

import json
import os
import re
import collections
from typing import Dict, List

from router_inference.router.base_router import BaseRouter

_BOXED = re.compile(r"\\boxed\{+([^{}]*)\}+")


def _norm(s: str) -> str:
return re.sub(r"[^a-z0-9]", "", str(s).lower())


class CruqSCRouter(BaseRouter):
"""Self-consistency cascade: probe the cheap model K times, escalate on disagreement."""

def __init__(self, router_name: str):
super().__init__(router_name)
params = self.config["pipeline_params"]
self.k = int(params.get("k", 4))
self.tau = float(params.get("tau", 0.6))
self.probe_model = params.get("probe_model", "qwen/qwen3-235b-a22b-2507")
self.escalate_model = params.get("escalate_model", "deepseek/deepseek-v4-flash")

here = os.path.dirname(os.path.abspath(__file__))
root = os.path.dirname(os.path.dirname(here))
# prompt -> global index, so a raw query string can find its cached probes
self._prompt_to_gi: Dict[str, str] = {}
for path in ("dataset/router_data.json", "dataset/router_data_10.json"):
p = os.path.join(root, path)
if os.path.exists(p):
for e in json.load(open(p, encoding="utf-8")):
self._prompt_to_gi[e["prompt_formatted"]] = e["global index"]

# global index -> list of normalized boxed answers from the K probes
self._samples: Dict[str, List[str]] = collections.defaultdict(list)
cache = os.path.join(root, "phase2", "data", "qwen_sc_full.jsonl")
if os.path.exists(cache):
for line in open(cache, encoding="utf-8"):
try:
r = json.loads(line)
except Exception:
continue
if r["s"] < self.k:
self._samples[r["gi"]].append(_norm(r["boxed"]))

def _get_prediction(self, query: str) -> str:
gi = self._prompt_to_gi.get(query)
if gi is None:
# Unknown query (no cached probes): fall back to the cheap probe model.
return self.probe_model
samples = self._samples.get(gi, [])
boxes = [b for b in samples if b]
if not boxes:
# Free-form dataset (no \boxed answer): self-consistency can't apply, so keep
# the cheap probe model rather than pay to escalate on a signal we don't have.
return self.probe_model
top = collections.Counter(boxes).most_common(1)[0][1]
consistency = top / len(samples)
return self.probe_model if consistency >= self.tau else self.escalate_model
3 changes: 3 additions & 0 deletions universal_model_names.py
Original file line number Diff line number Diff line change
Expand Up @@ -41,6 +41,9 @@
"gemini-2.5-flash",
"gemini-2.5-pro",
"google/gemini-3.1-flash-lite",
"google/gemini-2.5-flash-lite",
"google/gemini-2.5-pro",
"anthropic/claude-sonnet-4.5",
"gemini-3-flash-preview",
# Mistral models
"mistral-medium",
Expand Down