‹ PublicAI Index
The LLM benchmark aggregator.
Grok 4.20
SpaceXAI
Strongest in Mathematics (#18 of 140, on 1 of its 2 boards), weakest in Coding (#147 of 165). Above par in 8 of 14 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5 and Claude Fable 5.1 and ahead of Qwen3 Max and Mimo V2.5 Pro.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Core abilities57.6−9.8#19/2041/3
General intelligence60.9−9.2#20/2041/3
Reasoning57.8−9.1#23/1782/4
Mathematics58−7.7#18/1401/2
Reasoning56.4−12.7#34/1401/3
Science58.8−4.7#34/1221/1
Knowledge55.2−8.3#47/1381/2
Academic knowledge56.3−7.6#43/1231/1
Professional43.7−17.9#134/1681/1
Finance44.2−18.7#116/1511/1
Medical43.3−23.5#117/1401/1
Legal42.4−24.7#121/1511/1
Coding41.1−28#147/1651/5
Agentic coding40−28.7#143/1571/4
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Grok 4.20, left for the other.
§ 3 · Sources
Where the numbers come from
3 publications, 19 figures. Every one links to the page it was read from.
ARC-AGI-2 65.1%
Kagi LLM Benchmark 75%
Vals · Legal Research Bench 13.94%Vals · CaseLaw 54.45%Vals · LegalBench 77.74%Vals · Harvey Legal Agent Benchmark 0%Vals · Finance Agent 28.49%Vals · CorpFin 63.67%Vals · TaxEval 74.12%Vals · MortgageTax 45.35%Vals · MedQA 94.55%Vals · MedCode 32.16%Vals · MedScribe 63.41%Vals · SWE-bench Verified 72.2% (Mini-SWE-agent)Vals · Vibe Code Bench 4.06% (OpenHands)Vals · Code Migration 0.31%Vals · GPQA Diamond 88.64%Vals · MMLU Pro 86.25%Vals · AIME 96.46%
Badge
[](https://publicai.io/model-index/m/grok-4-20)