‹ PublicAI Index
The LLM benchmark aggregator.
Mistral Medium 3.1
Mistral
Strongest in Medical (#96 of 140), weakest in Agents (#193 of 268). Above par in 0 of 10 scopes. Among the models it meets almost everywhere, it finishes behind Claude Opus 5 and Claude Fable 5 and ahead of Command A and Llama 3.3 Instruct Turbo (70B).
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Knowledge43.8−19.7#113/1381/2
Academic knowledge42.6−21.3#102/1231/1
Reasoning45.8−21.1#123/1781/4
Mathematics43.6−22.1#102/1401/2
Professional41.7−19.9#152/1681/1
Medical47.2−19.6#96/1401/1
Finance39.7−23.2#129/1511/1
Legal37.6−29.5#138/1511/1
Agents44.3−23.8#193/2682/5
Knowledge work42.6−31#119/1782/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Mistral Medium 3.1, left for the other.
§ 3 · Sources
Where the numbers come from
3 publications, 9 figures. Every one links to the page it was read from.
GDPval-AA 371
AA-Briefcase 527
Vals · LegalBench 61.98%Vals · CorpFin 50.74%Vals · TaxEval 70.32%Vals · MortgageTax 36.45%Vals · MedQA 78.23%Vals · MMLU Pro 75.29%Vals · AIME 42.29%
Badge
[](https://publicai.io/model-index/m/mistral-medium-3-1)