‹ PublicAI Index
The LLM benchmark aggregator.
Magistral Medium
Mistral
Strongest in Reasoning (#104 of 140, on 1 of its 3 boards), weakest in Safety (#281 of 337). Above par in 0 of 12 scopes. Among the models it meets almost everywhere, it finishes behind Claude Opus 5.5 and Claude Opus 4.7 and ahead of Grok Build 0.1 and Qwen3 30B A3B Instruct.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Reasoning46−20.9#119/1781/4
Reasoning42.7−26.4#104/1401/3
Core abilities37.8−29.6#198/2041/3
General intelligence32.4−37.7#201/2041/3
Human preference47.9−19.6#218/3421/1
Human preference47.9−19.6#218/3421/1
Safety45.6−15.9#281/3371/3
Secure code48.8−17.7#170/2741/1
Fairness45.9−24.1#177/3001/2
Jailbreak resistance46.8−18.1#202/2721/1
Toxicity avoidance47.3−11.7#219/2721/1
Harm refusal42.5−19.1#273/3001/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Magistral Medium, left for the other.
§ 3 · Sources
Where the numbers come from
4 publications, 9 figures. Every one links to the page it was read from.
LMArena Text 1305
ARC-AGI-2 0%
Kagi LLM Benchmark 16.2%
Enkrypt · Jailbreak risk 15.5%Enkrypt · Harmful content risk 66.7%Enkrypt · CBRN risk 16.8%Enkrypt · Toxicity risk 8.2%Enkrypt · Bias risk 84.2%Enkrypt · Insecure code risk 36%
Badge
[](https://publicai.io/model-index/m/magistral-medium)