‹ PublicAI Index
The LLM benchmark aggregator.
Grok 3 Mini
SpaceXAI
Strongest in Safe-prompt compliance (#8 of 82), weakest in Harm refusal (#157 of 299). Above par in 19 of 21 scopes. Among the models it meets almost everywhere, it finishes behind Claude Opus 5 and GPT-5.5 and ahead of Grok 4.1 Fast and Kimi K2 Thinking.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Professional53.1−8.5#49/1681/1
Legal53.8−13.3#47/1511/1
Medical53.5−13.3#55/1401/1
Finance53.3−9.6#66/1511/1
Coding52.9−16.2#50/1642/5
Code generation55.6−14.4#25/771/2
Agentic coding51.2−17.5#63/1571/4
Core abilities52.9−14.5#62/2041/3
General intelligence54.2−15.9#65/2041/3
Knowledge50.2−13.3#84/1381/2
Academic knowledge50.2−13.7#77/1231/1
Reasoning50.6−16.3#87/1782/4
Mathematics55−10.7#51/1401/2
Science52.2−11.3#70/1221/1
Reasoning42.8−26.3#99/1401/3
Safety52.7−8.8#109/3361/3
Safe-prompt compliance59.5−0.9#8/821/1
Fairness56.9−13.1#51/2991/2
Harm refusal49.9−11.7#157/2991/2
Human preference53.5−14#154/3411/1
Human preference53.5−14#154/3411/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Grok 3 Mini, left for the other.
§ 3 · Sources
Where the numbers come from
7 publications, 17 figures. Every one links to the page it was read from.
LMArena Text 1364
ARC-AGI-2 0.4%
LiveCodeBench 66.7%
Aider polyglot 49.3%
Kagi LLM Benchmark 61.3%
Vals · LegalBench 83.14%Vals · CorpFin 61.11%Vals · TaxEval 72.98%Vals · MedQA 90.1%Vals · GPQA Diamond 79.29%Vals · MMLU Pro 81.37%Vals · AIME 85%
HELM Safety · HarmBench 57.2%HELM Safety · SimpleSafetyTests 99.3%HELM Safety · Anthropic Red Team 99.5%HELM Safety · BBQ 96.7%HELM Safety · XSTest 98.4%
Badge
[](https://publicai.io/model-index/m/grok-3-mini)