‹ PublicAI Index
The LLM benchmark aggregator.
Mistral Medium 3.5
Mistral
Strongest in Human preference (#95 of 342), weakest in Reasoning (#169 of 178). Above par in 2 of 17 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5 and Claude Opus 4.7 and ahead of GPT-4.1 Nano and Mistral Small.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Human preference59.5−8#95/3421/1
Human preference59.5−8#95/3421/1
Knowledge43.9−19.6#112/1381/2
Academic knowledge42.6−21.3#101/1231/1
Coding40.5−28.6#148/1651/5
Agentic coding39.5−29.2#145/1571/4
Agents47.2−20.9#149/2682/5
Knowledge work46.4−27.2#99/1782/2
Professional38.8−22.8#159/1681/1
Medical39.9−26.9#125/1401/1
Finance39.7−23.2#130/1511/1
Legal34.8−32.3#151/1511/1
Core abilities43.7−23.7#164/2042/3
General intelligence42−28.1#159/2042/3
Reasoning35.9−31#169/1781/4
Mathematics37.7−28#116/1401/2
Science26.1−37.4#117/1221/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Mistral Medium 3.5, left for the other.
§ 3 · Sources
Where the numbers come from
6 publications, 20 figures. Every one links to the page it was read from.
LMArena Text 1426
Artificial Analysis Intelligence Index 14
GDPval-AA 747
AA-Briefcase 521
Kagi LLM Benchmark 41.4%
Vals · Legal Research Bench 9.13%Vals · CaseLaw 44.16%Vals · Harvey Legal Agent Benchmark 0.42%Vals · Finance Agent 32.1%Vals · CorpFin 58.78%Vals · TaxEval 67.99%Vals · MortgageTax 28.89%Vals · MedCode 33.75%Vals · MedScribe 67.73%Vals · SWE-bench Verified 66.4% (Mini-SWE-agent)Vals · Vibe Code Bench 2.89% (OpenHands)Vals · Code Migration 5.13%Vals · GPQA Diamond 34.85%Vals · MMLU Pro 75.33%Vals · ProofBench 9%
Badge
[](https://publicai.io/model-index/m/mistral-medium-3-5)