‹ PublicAI Index
The LLM benchmark aggregator.
Grok 3
SpaceXAI
Strongest in Factual grounding (#22 of 101), weakest in Safety (#297 of 336). Above par in 13 of 24 scopes. Among the models it meets almost everywhere, it finishes behind O3 and Claude Opus 5 and ahead of Qwen3.5 Flash and GPT-4 Turbo.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Professional52.5−9.1#54/1681/1
Legal53.4−13.7#51/1511/1
Finance54−8.9#60/1511/1
Medical50.2−16.6#80/1401/1
Coding51.9−17.2#59/1641/5
Agentic coding52.3−16.4#57/1571/4
Core abilities52.9−14.5#61/2041/3
General intelligence54.2−15.9#64/2041/3
Knowledge48.7−14.8#92/1381/2
Academic knowledge48.4−15.5#85/1231/1
Human preference58.1−9.4#113/3411/1
Human preference58.1−9.4#113/3411/1
Reasoning45.3−21.6#127/1783/4
Science48.7−14.8#82/1221/1
Mathematics48−17.7#87/1401/2
Reasoning41.3−27.8#108/1402/3
Safety44.1−17.4#297/3363/3
Factual grounding60.5−10.1#22/1011/1
Safe-prompt compliance55.1−5.3#36/821/1
Fairness51.4−18.6#97/2992/2
Secure code40.9−25.6#212/2741/1
Harm refusal43.5−18.1#265/2992/2
Toxicity avoidance26−33#267/2721/1
Jailbreak resistance26−38.9#272/2721/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Grok 3, left for the other.
§ 3 · Sources
Where the numbers come from
9 publications, 24 figures. Every one links to the page it was read from.
LMArena Text 1411
ARC-AGI-2 0%
Aider polyglot 53.3%
Kagi LLM Benchmark 61.3%
SimpleBench 36.1%
Vals · LegalBench 82.59%Vals · CorpFin 59.71%Vals · TaxEval 75.88%Vals · MedQA 83.85%Vals · GPQA Diamond 74.24%Vals · MMLU Pro 79.95%Vals · AIME 58.75%
HELM Safety · HarmBench 45.3%HELM Safety · SimpleSafetyTests 96.8%HELM Safety · Anthropic Red Team 95.5%HELM Safety · BBQ 93.6%HELM Safety · XSTest 96.4%
Enkrypt · Jailbreak risk 68.2%Enkrypt · Harmful content risk 59.4%Enkrypt · CBRN risk 12%Enkrypt · Toxicity risk 38.3%Enkrypt · Bias risk 80.6%Enkrypt · Insecure code risk 52%
Vectara · Factual consistency 94.2%
Badge
[](https://publicai.io/model-index/m/grok-3)