‹ PublicAI Index
The LLM benchmark aggregator.
Mistral Large
Mistral
Strongest in Fairness (#8 of 300, on 1 of its 2 boards), weakest in Harm refusal (#223 of 300). Above par in 6 of 19 scopes. Among the models it meets almost everywhere, it finishes behind Claude Sonnet 5 and Grok 4.6 and ahead of Gemini 1.5 Flash and Claude 3.5 Haiku.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Professional47.2−14.4#110/1681/1
Medical46.1−20.7#106/1401/1
Finance45.8−17.1#111/1511/1
Safety52.7−8.8#111/3372/3
Fairness70leads#8/3001/2
Factual grounding63.8−6.8#8/1011/1
Jailbreak resistance51.2−13.7#175/2721/1
Secure code44−22.5#199/2741/1
Toxicity avoidance48−11#215/2721/1
Harm refusal46.3−15.3#223/3001/2
Agents50.2−17.9#114/2681/5
Tool use50.4−23.6#34/811/1
Knowledge38−25.5#122/1381/2
Academic knowledge35.6−28.3#107/1231/1
Reasoning36−30.9#168/1781/4
Science30.3−33.2#111/1221/1
Mathematics34.8−30.9#131/1401/2
Human preference48−19.5#216/3421/1
Human preference48−19.5#216/3421/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Mistral Large, left for the other.
§ 3 · Sources
Where the numbers come from
5 publications, 14 figures. Every one links to the page it was read from.
LMArena Text 1306
BFCL v4 38.37%
Vals · TaxEval 63.78%Vals · MedQA 76.22%Vals · GPQA Diamond 47.73%Vals · MMLU Pro 69.71%Vals · AIME 9.17%
Enkrypt · Jailbreak risk 11.8%Enkrypt · Harmful content risk 61.1%Enkrypt · CBRN risk 9.8%Enkrypt · Toxicity risk 7.7%Enkrypt · Bias risk 39.5%Enkrypt · Insecure code risk 45.8%
Vectara · Factual consistency 95.5%
Badge
[](https://publicai.io/model-index/m/mistral-large)