‹ PublicAI Index
The LLM benchmark aggregator.
Grok 3 Mini
SpaceXAI
Strongest in Code generation (#25 of 77, on 1 of its 2 boards), weakest in Human preference (#154 of 342). Above par in 15 of 16 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5 and Claude Opus 4.7 and ahead of MiniMax M2.5 and GPT-6 Luna.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Professional53.1−8.5#49/1681/1
Legal53.8−13.3#47/1511/1
Medical53.5−13.3#55/1401/1
Finance53.3−9.6#66/1511/1
Coding52.9−16.2#50/1651/5
Code generation55.6−14.4#25/771/2
Core abilities52.9−14.5#62/2041/3
General intelligence54.2−15.9#65/2041/3
Knowledge50.2−13.3#84/1381/2
Academic knowledge50.2−13.7#77/1231/1
Reasoning50.6−16.3#87/1782/4
Mathematics55−10.7#51/1401/2
Science52.2−11.3#70/1221/1
Reasoning42.8−26.3#99/1401/3
Human preference53.5−14#154/3421/1
Human preference53.5−14#154/3421/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Grok 3 Mini, left for the other.
§ 3 · Sources
Where the numbers come from
5 publications, 11 figures. Every one links to the page it was read from.
LMArena Text 1364
ARC-AGI-2 0.4%
LiveCodeBench 66.7%
Kagi LLM Benchmark 61.3%
Vals · LegalBench 83.14%Vals · CorpFin 61.11%Vals · TaxEval 72.98%Vals · MedQA 90.1%Vals · GPQA Diamond 79.29%Vals · MMLU Pro 81.37%Vals · AIME 85%
Badge
[](https://publicai.io/model-index/m/grok-3-mini)