‹ PublicAI Index
The LLM benchmark aggregator.
Grok 4
SpaceXAI
Strongest in Tool use (#8 of 81), weakest in Safety (#299 of 337). Above par in 20 of 25 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5 and Claude Fable 5.1 and ahead of Claude Sonnet 4.6 and GPT-4 Turbo.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Core abilities57.1−10.3#25/2041/3
General intelligence60.2−9.9#26/2041/3
Agents60.1−8#31/2681/5
Tool use68.2−5.8#8/811/1
Reasoning55.3−11.6#37/1783/4
Science58.4−5.1#35/1221/1
Mathematics56.5−9.2#36/1401/2
Reasoning52.8−16.3#49/1402/3
Knowledge54.2−9.3#62/1381/2
Academic knowledge55.1−8.8#57/1231/1
Coding51.4−17.7#64/1652/5
Agentic coding51.7−17#61/1572/4
Professional50.7−10.9#73/1681/1
Legal57.3−9.8#23/1511/1
Medical50.5−16.3#77/1401/1
Finance46.2−16.7#107/1511/1
Human preference58.1−9.4#112/3421/1
Human preference58.1−9.4#112/3421/1
Safety43.4−18.1#299/3372/3
Safe-prompt compliance55.5−4.9#30/821/1
Fairness48.8−21.2#132/3002/2
Secure code50.8−15.7#153/2741/1
Toxicity avoidance52.4−6.6#166/2721/1
Jailbreak resistance26−38.9#270/2721/1
Harm refusal38.7−22.9#292/3002/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Grok 4, left for the other.
§ 3 · Sources
Where the numbers come from
9 publications, 29 figures. Every one links to the page it was read from.
LMArena Text 1411
ARC-AGI-2 29.4%
Aider polyglot 79.6%
BFCL v4 62.97%
Kagi LLM Benchmark 73.6%
SimpleBench 60.5%
Vals · CaseLaw 65.81%Vals · LegalBench 83.19%Vals · CorpFin 66.05%Vals · TaxEval 65.09%Vals · MortgageTax 44.48%Vals · MedQA 92.49%Vals · MedCode 38.08%Vals · MedScribe 78.15%Vals · SWE-bench Verified 57.8% (Mini-SWE-agent)Vals · GPQA Diamond 88.13%Vals · MMLU Pro 85.3%Vals · AIME 90.56%
HELM Safety · HarmBench 39.7%HELM Safety · SimpleSafetyTests 92.2%HELM Safety · Anthropic Red Team 95.7%HELM Safety · BBQ 93.7%HELM Safety · XSTest 96.6%
Enkrypt · Jailbreak risk 43%Enkrypt · Harmful content risk 66.1%Enkrypt · CBRN risk 23.2%Enkrypt · Toxicity risk 4.6%Enkrypt · Bias risk 87.6%Enkrypt · Insecure code risk 32%
Badge
[](https://publicai.io/model-index/m/grok-4)