‹ PublicAI Index
The LLM benchmark aggregator.
Grok 4.5
SpaceXAI
Strongest in Web research (#4 of 4), weakest in Toxicity avoidance (#191 of 272). Above par in 25 of 30 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5 and Claude Fable 5.1 and ahead of Gemini 3.6 Flash and GPT-5.1.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Knowledge58.3−5.2#16/1381/2
Academic knowledge60−3.9#14/1231/1
Core abilities56.9−10.5#26/2042/3
General intelligence61.2−8.9#16/2042/3
Language56.6−14.9#20/571/1
Instruction following56.3−17.4#20/571/1
Data analysis49.3−16.8#36/571/1
Professional54.6−7#37/1681/1
IT operations40.1−33.7#10/111/1
Legal59.9−7.2#13/1511/1
Medical54.9−11.9#39/1401/1
Finance54.4−8.5#54/1511/1
Human preference63.3−4.2#42/3421/1
Human preference63.3−4.2#42/3421/1
Reasoning54.9−12#45/1784/4
Science61.8−1.7#12/1221/1
Reasoning58.3−10.8#29/1403/3
Mathematics47.6−18.1#89/1402/2
Safety55−6.5#52/3371/3
Secure code66.3−0.2#5/2741/1
Fairness57.4−12.6#48/3001/2
Harm refusal52.8−8.8#114/3001/2
Jailbreak resistance52.8−12.1#145/2721/1
Toxicity avoidance51−8#191/2721/1
Agents54.9−13.2#66/2683/5
Web research36.5−27.7#4/41/1
Knowledge work62.2−11.4#30/1782/2
Coding50.1−19#74/1653/5
Agentic coding54.8−13.9#40/1573/4
Code generation32−38#73/771/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Grok 4.5, left for the other.
§ 3 · Sources
Where the numbers come from
10 publications, 38 figures. Every one links to the page it was read from.
LMArena Text 1465
GDPval-AA 1370
AA-Briefcase 1283
Terminal-Bench 12.4% (Grok Build)
ARC-AGI-2 52.6%
LiveBench 75.8LiveBench · Reasoning 87.2LiveBench · Coding 68.6LiveBench · Agentic Coding 56.5LiveBench · Mathematics 90.8LiveBench · Data Analysis 73LiveBench · Language 82.8LiveBench · Instruction Following 71.5
Kagi LLM Benchmark 83.5%
SimpleBench 70%
Vals · Legal Research Bench 37.98%Vals · LegalBench 85.97%Vals · Harvey Legal Agent Benchmark 12.92%Vals · Finance Agent 48.35%Vals · CorpFin 67.41%Vals · TaxEval 71.67%Vals · MortgageTax 61.8%Vals · MedCode 43.29%Vals · MedScribe 86.88%Vals · SRE Bench 0.76%Vals · Web Search Index 38.75% (Exa)Vals · SWE-bench Verified 86.6% (Mini-SWE-agent)Vals · Vibe Code Bench 69% (OpenHands)Vals · Code Migration 36.6%Vals · GPQA Diamond 92.93%Vals · MMLU Pro 89.22%Vals · ProofBench 31%
Enkrypt · Jailbreak risk 10.4%Enkrypt · Harmful content risk 0%Enkrypt · CBRN risk 25.3%Enkrypt · Toxicity risk 5.6%Enkrypt · Bias risk 66.4%Enkrypt · Insecure code risk 0.4%
Badge
[](https://publicai.io/model-index/m/grok-4-5)