-
Notifications
You must be signed in to change notification settings - Fork 822
Pull requests: open-compass/opencompass
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
[Fix] Support Transformers 5 tokenization
#2575
opened Aug 3, 2026 by
taking-lying-flat
•
Draft
3 tasks done
[Fix] Fix per-turn statistics in dump-res-length and repeat detection issues for multi-round conversations
#2569
opened Jul 27, 2026 by
Myhs-phz
Collaborator
Loading…
[Dataset] Align ARC-AGI-1/2 evaluation with ARC Prize protocol
#2563
opened Jul 21, 2026 by
Vicentvankor
Loading…
fix: use last-match in generic LLM judge postprocess (verdict-injection defense)
#2534
opened Jul 14, 2026 by
AUTHENSOR
Loading…
3 tasks done
feat(dataset): add CHHallu-Src v1 — Chinese history hallucination & source-attribution benchmark
#2525
opened Jul 11, 2026 by
lizhuojunx86
Loading…
fix(MedCalc_Bench): guard calid-69 ground_truth regex match (AttributeError)
#2523
opened Jul 9, 2026 by
WatchTree-19
Loading…
fix(evaluator): guard IndexError on truncated judge choice (JudgeEvaluator/RMBEvaluator)
#2522
opened Jul 9, 2026 by
WatchTree-19
Loading…
Add robust HumanEval postprocess for chat outputs
#2515
opened Jul 8, 2026 by
Ding-god
Loading…
3 of 6 tasks
[Fix] Extract the final answer in GPQA simple-eval predictions
#2496
opened Jun 27, 2026 by
Hibbert133
Contributor
Loading…
4 of 6 tasks
Add optional juryeval integration for LLM-as-Judge metrics
#2465
opened May 31, 2026 by
py-ai-dev
Loading…
[Fix] Combine split eval results in default summarizer
#2451
opened May 15, 2026 by
yhzhu99
Contributor
Loading…
feat: upgrade MiniMax default model to M3
#2418
opened Mar 20, 2026 by
octo-patch
Contributor
Loading…
3 tasks done
[Fix] CEval ModelScope load and HF generate for causal LMs
#2416
opened Mar 19, 2026 by
DeliWang
Loading…
6 tasks
Add support for Azure OpenAI models and managed identity auth
#2415
opened Mar 18, 2026 by
jgbradley1
Loading…
2 of 6 tasks
Previous Next
ProTip!
Mix and match filters to narrow down what you’re looking for.