‹ PublicAI Index
The LLM benchmark aggregator.
Mistral Small 3.1
Mistral
Strongest in Academic knowledge (#114 of 123), weakest in Agents (#217 of 268). Above par in 0 of 11 scopes. Among the models it meets almost everywhere, it finishes behind Claude Opus 5 and Claude Fable 5 and ahead of Mistral Small and Claude 3.5 Haiku.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Knowledge34.2−29.3#130/1381/2
Academic knowledge31−32.9#114/1231/1
Professional42.1−19.5#148/1681/1
Legal43.1−24#117/1511/1
Medical42.3−24.5#119/1401/1
Finance40.1−22.8#127/1511/1
Reasoning34.4−32.5#174/1781/4
Science27.8−35.7#115/1221/1
Mathematics33.3−32.4#134/1401/2
Agents42.2−25.9#217/2682/5
Knowledge work39.9−33.7#133/1782/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Mistral Small 3.1, left for the other.
§ 3 · Sources
Where the numbers come from
3 publications, 10 figures. Every one links to the page it was read from.
GDPval-AA 363
AA-Briefcase 313
Vals · LegalBench 69.16%Vals · CorpFin 44.17%Vals · TaxEval 58.3%Vals · MortgageTax 61.21%Vals · MedQA 69.1%Vals · GPQA Diamond 44.19%Vals · MMLU Pro 66.02%Vals · AIME 3.54%
Badge
[](https://publicai.io/model-index/m/mistral-small-3-1)