‹ PublicAI Index
The LLM benchmark aggregator.
Magistral Small
Mistral
Strongest in Reasoning (#101 of 140, on 1 of its 3 boards), weakest in Safety (#297 of 337). Above par in 1 of 10 scopes. Among the models it meets almost everywhere, it finishes behind GPT-5.5 and Claude Opus 5.5 and ahead of Gemma 2B It and Qwen2.5 0.5B Instruct.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Reasoning46−20.9#118/1781/4
Reasoning42.7−26.4#101/1401/3
Core abilities37.8−29.6#201/2041/3
General intelligence32.4−37.7#204/2041/3
Safety44.1−17.4#297/3371/3
Fairness47.4−22.6#152/3001/2
Jailbreak resistance50.6−14.3#179/2721/1
Secure code35.5−31#235/2741/1
Toxicity avoidance41.2−17.8#240/2721/1
Harm refusal43.5−18.1#266/3001/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Magistral Small, left for the other.
§ 3 · Sources
Where the numbers come from
3 publications, 8 figures. Every one links to the page it was read from.
ARC-AGI-2 0%
Kagi LLM Benchmark 6.3%
Enkrypt · Jailbreak risk 12.3%Enkrypt · Harmful content risk 63.9%Enkrypt · CBRN risk 15.5%Enkrypt · Toxicity risk 12.5%Enkrypt · Bias risk 81.9%Enkrypt · Insecure code risk 63.1%
Badge
[](https://publicai.io/model-index/m/magistral-small)