‹ PublicAI Index
The LLM benchmark aggregator.
Mistral Large 3
Mistral
Strongest in Legal (#61 of 151), weakest in Safety (#267 of 337). Above par in 4 of 20 scopes. Among the models it meets almost everywhere, it finishes behind Claude Opus 4.7 and GPT-5.6 Sol and ahead of O3 Mini and GPT-5 Nano.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Professional50.4−11.2#74/1681/1
Legal52.2−14.9#61/1511/1
Medical49.3−17.5#84/1401/1
Finance49.7−13.2#90/1511/1
Knowledge48.5−15#94/1381/2
Academic knowledge48.2−15.7#86/1231/1
Human preference58.3−9.2#110/3421/1
Human preference58.3−9.2#110/3421/1
Reasoning41.5−25.4#147/1782/4
Science44.7−18.8#92/1221/1
Mathematics43.8−21.9#101/1401/2
Reasoning36.6−32.5#129/1401/3
Coding39.7−29.4#154/1651/5
Agentic coding37.6−31.1#153/1571/4
Core abilities44.5−22.9#155/2042/3
General intelligence43−27.1#149/2042/3
Agents41.7−26.4#221/2682/5
Knowledge work39.3−34.3#136/1782/2
Safety46.4−15.1#267/3371/3
Factual grounding38.6−32#89/1011/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Mistral Large 3, left for the other.
§ 3 · Sources
Where the numbers come from
8 publications, 17 figures. Every one links to the page it was read from.
LMArena Text 1413
Artificial Analysis Intelligence Index 9
GDPval-AA 406
AA-Briefcase 229
Kagi LLM Benchmark 50.9%
SimpleBench 20.4%
Vals · CaseLaw 61.41%Vals · LegalBench 79.14%Vals · CorpFin 61.03%Vals · TaxEval 73.06%Vals · MortgageTax 52.11%Vals · MedQA 82.23%Vals · SWE-bench Verified 41.4% (Mini-SWE-agent)Vals · GPQA Diamond 68.43%Vals · MMLU Pro 79.82%Vals · AIME 42.92%
Vectara · Factual consistency 85.5%
Badge
[](https://publicai.io/model-index/m/mistral-large-3)