‹ PublicAI Index
The LLM benchmark aggregator.
Grok 4.7
SpaceXAI
Strongest in Fairness (#4 of 300, on 1 of its 2 boards), weakest in Reasoning (#89 of 178). Above par in 24 of 26 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5 and Claude Fable 5.1 and ahead of GPT-5.5 and Claude Sonnet 5.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Agents65.5−2.6#5/2682/5
Knowledge work70.1−3.5#4/1782/2
Safety60−1.5#6/3371/3
Fairness70leads#4/3001/2
Secure code65.2−1.3#12/2741/1
Jailbreak resistance61.7−3.2#25/2721/1
Toxicity avoidance57.1−1.9#46/2721/1
Harm refusal55.4−6.2#64/3001/2
Professional57.6−4#13/1681/1
Legal61.2−5.9#7/1511/1
Biology research50.2−18.3#8/161/1
Medical59.9−6.9#10/1401/1
Finance53.7−9.2#62/1511/1
Core abilities58.5−8.9#15/2042/3
Instruction following63−10.7#10/571/1
General intelligence61−9.1#19/2042/3
Data analysis55.9−10.2#23/571/1
Language51.5−20#26/571/1
Coding56−13.1#28/1653/5
Agentic coding58−10.7#22/1573/4
Code generation49.6−20.4#47/771/2
Human preference60.8−6.7#78/3421/1
Human preference60.8−6.7#78/3421/1
Reasoning50.3−16.6#89/1782/4
Reasoning48.4−20.7#68/1401/3
Mathematics51.4−14.3#76/1402/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Grok 4.7, left for the other.
§ 3 · Sources
Where the numbers come from
8 publications, 29 figures. Every one links to the page it was read from.
LMArena Text 1439
Artificial Analysis Intelligence Index 46
GDPval-AA 1695
AA-Briefcase 1657
Terminal-Bench 37.6% (Grok Build)
LiveBench 77.4LiveBench · Reasoning 82.7LiveBench · Coding 77.2LiveBench · Agentic Coding 54LiveBench · Mathematics 95.7LiveBench · Data Analysis 76.9LiveBench · Language 80.1LiveBench · Instruction Following 75.3
Vals · Legal Research Bench 47.12%Vals · LegalBench 84.39%Vals · Harvey Legal Agent Benchmark 12.5%Vals · Finance Agent 52.25%Vals · MedCode 49.55%Vals · MedScribe 89.38%Vals · BioMysteryBench 69.26%Vals · Vibe Code Bench 86.17% (OpenHands)Vals · Code Migration 44.82%Vals · ProofBench 26%
Enkrypt · Jailbreak risk 2.9%Enkrypt · Harmful content risk 0%Enkrypt · CBRN risk 18.3%Enkrypt · Toxicity risk 1.3%Enkrypt · Bias risk 26.6%Enkrypt · Insecure code risk 2.7%
Badge
[](https://publicai.io/model-index/m/grok-4-7)