‹ PublicAI Index
The LLM benchmark aggregator.
Mistral Small
Mistral
Strongest in Factual grounding (#12 of 101), weakest in Harm refusal (#236 of 300). Above par in 5 of 21 scopes. Among the models it meets almost everywhere, it finishes behind Claude Sonnet 5 and O3 and ahead of Llama 4 Scout Instruct and Claude 3.5 Haiku.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Agents49.7−18.4#123/2681/5
Tool use49.5−24.5#37/811/1
Knowledge32.5−31#132/1381/2
Academic knowledge29−34.9#116/1231/1
Core abilities44.9−22.5#145/2041/3
General intelligence42.7−27.4#150/2041/3
Professional40.1−21.5#155/1681/1
Medical35.9−30.9#129/1401/1
Finance35−27.9#142/1511/1
Human preference52.9−14.6#159/3421/1
Human preference52.9−14.6#159/3421/1
Reasoning36.2−30.7#167/1781/4
Science32.4−31.1#107/1221/1
Mathematics33.9−31.8#133/1401/2
Safety48.7−12.8#219/3372/3
Factual grounding62.3−8.3#12/1011/1
Jailbreak resistance52.7−12.2#147/2721/1
Toxicity avoidance51.1−7.9#189/2721/1
Fairness44.9−25.1#200/3001/2
Secure code40.7−25.8#216/2741/1
Harm refusal45.6−16#236/3001/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Mistral Small, left for the other.
§ 3 · Sources
Where the numbers come from
6 publications, 15 figures. Every one links to the page it was read from.
LMArena Text 1357
BFCL v4 37.15%
Kagi LLM Benchmark 37.8%
Vals · TaxEval 49.14%Vals · MedQA 56.98%Vals · GPQA Diamond 50.76%Vals · MMLU Pro 64.44%Vals · AIME 5.63%
Enkrypt · Jailbreak risk 10.5%Enkrypt · Harmful content risk 60.6%Enkrypt · CBRN risk 11.8%Enkrypt · Toxicity risk 5.5%Enkrypt · Bias risk 85.8%Enkrypt · Insecure code risk 52.4%
Vectara · Factual consistency 94.9%
Badge
[](https://publicai.io/model-index/m/mistral-small)