‹ PublicAI Index
The LLM benchmark aggregator.
Mimo V2.5 Pro
Xiaomi · 1T · open weights
Strongest in Human preference (#39 of 342), weakest in Jailbreak resistance (#231 of 272). Above par in 13 of 23 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5 and Claude Fable 5.1 and ahead of DeepSeek V3.2 and GPT-4o.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Human preference63.5−4#39/3421/1
Human preference63.5−4#39/3421/1
Knowledge53.5−10#65/1381/2
Academic knowledge54.2−9.7#60/1231/1
Agents53.5−14.6#82/2682/5
Knowledge work54.5−19.1#57/1782/2
Coding48.1−21#88/1651/5
Agentic coding47.8−20.9#84/1571/4
Core abilities50.1−17.3#92/2041/3
General intelligence50.2−19.9#93/2041/3
Reasoning47.4−19.5#105/1781/4
Science54.5−9#61/1221/1
Mathematics41.3−24.4#108/1401/2
Professional47.7−13.9#107/1681/1
Finance51.1−11.8#82/1511/1
Medical46.7−20.1#99/1401/1
Legal44.6−22.5#106/1511/1
Safety49.5−12#198/3371/3
Toxicity avoidance56.4−2.6#60/2721/1
Fairness54.3−15.7#72/3001/2
Secure code58−8.5#85/2741/1
Harm refusal46.3−15.3#220/3001/2
Jailbreak resistance36.1−28.8#231/2721/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Mimo V2.5 Pro, left for the other.
§ 3 · Sources
Where the numbers come from
6 publications, 24 figures. Every one links to the page it was read from.
LMArena Text 1467
Artificial Analysis Intelligence Index 26
GDPval-AA 1107
AA-Briefcase 881
Vals · Legal Research Bench 15.87%Vals · LegalBench 77.13%Vals · Harvey Legal Agent Benchmark 2.08%Vals · Finance Agent 41.5%Vals · CorpFin 61.42%Vals · TaxEval 73.79%Vals · MedCode 32.48%Vals · MedScribe 83.73%Vals · SWE-bench Verified 74% (Mini-SWE-agent)Vals · Vibe Code Bench 34.11% (OpenHands)Vals · Code Migration 21.56%Vals · GPQA Diamond 82.58%Vals · MMLU Pro 84.59%Vals · ProofBench 22%
Enkrypt · Jailbreak risk 24.5%Enkrypt · Harmful content risk 2.2%Enkrypt · CBRN risk 47%Enkrypt · Toxicity risk 1.8%Enkrypt · Bias risk 71.1%Enkrypt · Insecure code risk 17.3%
Badge
[](https://publicai.io/model-index/m/mimo-v2-5-pro)