‹ PublicAI Index
The LLM benchmark aggregator.
Mimo V2.5
Xiaomi · 311B · open weights
Strongest in Science (#64 of 122), weakest in Professional (#128 of 168). Above par in 7 of 15 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5 and Claude Opus 4.7 and ahead of Gemini 3.1 Flash Lite and GPT-4.1 Mini.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Knowledge51.8−11.7#77/1381/2
Academic knowledge52.1−11.8#71/1231/1
Human preference60.3−7.2#84/3421/1
Human preference60.3−7.2#84/3421/1
Coding47−22.1#98/1651/5
Agentic coding46.7−22#92/1571/4
Agents51.3−16.8#103/2682/5
Knowledge work51.7−21.9#68/1782/2
Reasoning46.4−20.5#113/1781/4
Science53.8−9.7#64/1221/1
Mathematics39.6−26.1#113/1401/2
Professional45.2−16.4#128/1681/1
Finance48.2−14.7#99/1511/1
Legal43.3−23.8#116/1511/1
Medical40.9−25.9#123/1401/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Mimo V2.5, left for the other.
§ 3 · Sources
Where the numbers come from
4 publications, 18 figures. Every one links to the page it was read from.
LMArena Text 1434
GDPval-AA 986
AA-Briefcase 752
Vals · Legal Research Bench 9.13%Vals · LegalBench 78.89%Vals · Harvey Legal Agent Benchmark 1.67%Vals · Finance Agent 36.73%Vals · CorpFin 59.91%Vals · TaxEval 71.83%Vals · MortgageTax 59.26%Vals · MedCode 31.89%Vals · MedScribe 72.15%Vals · SWE-bench Verified 71% (Mini-SWE-agent)Vals · Vibe Code Bench 42.17% (OpenHands)Vals · Code Migration 14.23%Vals · GPQA Diamond 81.57%Vals · MMLU Pro 82.93%Vals · ProofBench 16%
Badge
[](https://publicai.io/model-index/m/mimo-v2-5)