‹ PublicAI Index
The LLM benchmark aggregator.
Grok 3
SpaceXAI
Strongest in Factual grounding (#22 of 101), weakest in Safety (#308 of 337). Above par in 9 of 21 scopes. Among the models it meets almost everywhere, it finishes behind Gemini 3.8 Flash and Claude Opus 5 and ahead of Kimi K2 and GPT OSS 120B.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Professional52.5−9.1#54/1681/1
Legal53.4−13.7#51/1511/1
Finance54−8.9#60/1511/1
Medical50.2−16.6#80/1401/1
Core abilities52.9−14.5#61/2041/3
General intelligence54.2−15.9#64/2041/3
Knowledge48.7−14.8#92/1381/2
Academic knowledge48.4−15.5#85/1231/1
Human preference58.1−9.4#113/3421/1
Human preference58.1−9.4#113/3421/1
Reasoning45.3−21.6#127/1783/4
Science48.7−14.8#82/1221/1
Mathematics48−17.7#87/1401/2
Reasoning41.3−27.8#108/1402/3
Safety42.5−19#308/3372/3
Factual grounding60.5−10.1#22/1011/1
Fairness48.2−21.8#139/3001/2
Secure code40.9−25.6#212/2741/1
Harm refusal45.8−15.8#229/3001/2
Toxicity avoidance26−33#267/2721/1
Jailbreak resistance26−38.9#272/2721/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Grok 3, left for the other.
§ 3 · Sources
Where the numbers come from
7 publications, 18 figures. Every one links to the page it was read from.
LMArena Text 1411
ARC-AGI-2 0%
Kagi LLM Benchmark 61.3%
SimpleBench 36.1%
Vals · LegalBench 82.59%Vals · CorpFin 59.71%Vals · TaxEval 75.88%Vals · MedQA 83.85%Vals · GPQA Diamond 74.24%Vals · MMLU Pro 79.95%Vals · AIME 58.75%
Enkrypt · Jailbreak risk 68.2%Enkrypt · Harmful content risk 59.4%Enkrypt · CBRN risk 12%Enkrypt · Toxicity risk 38.3%Enkrypt · Bias risk 80.6%Enkrypt · Insecure code risk 52%
Vectara · Factual consistency 94.2%
Badge
[](https://publicai.io/model-index/m/grok-3)