‹ PublicAI Index
The LLM benchmark aggregator.
Grok 4 Fast
SpaceXAI
Strongest in Legal (#32 of 151), weakest in Safety (#310 of 337). Above par in 10 of 18 scopes. Among the models it meets almost everywhere, it finishes behind Claude Opus 4.7 and GPT-5.6 Sol and ahead of Grok 4.3 and Mimo V2.5.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Core abilities54.6−12.8#52/2041/3
General intelligence56.6−13.5#52/2041/3
Professional51.9−9.7#62/1681/1
Legal55.9−11.2#32/1511/1
Medical51.3−15.5#71/1401/1
Finance49.9−13#89/1511/1
Reasoning52.8−14.1#66/1782/4
Mathematics56.6−9.1#34/1401/2
Science56.5−7#49/1221/1
Reasoning43.8−25.3#85/1401/3
Knowledge48.4−15.1#95/1381/2
Academic knowledge48.1−15.8#87/1231/1
Human preference57.5−10#120/3421/1
Human preference57.5−10#120/3421/1
Coding37.2−31.9#163/1651/5
Agentic coding35.3−33.4#156/1571/4
Safety42.5−19#310/3371/3
Factual grounding26−44.6#96/1011/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Grok 4 Fast, left for the other.
§ 3 · Sources
Where the numbers come from
5 publications, 17 figures. Every one links to the page it was read from.
LMArena Text 1405
ARC-AGI-2 5.3%
Kagi LLM Benchmark 66.1%
Vals · CaseLaw 65.7%Vals · LegalBench 80.6%Vals · CorpFin 66.9%Vals · TaxEval 75.7%Vals · MortgageTax 42.09%Vals · MedQA 92.07%Vals · MedCode 37.38%Vals · MedScribe 81.63%Vals · SWE-bench Verified 45.4% (Mini-SWE-agent)Vals · Vibe Code Bench 0% (OpenHands)Vals · GPQA Diamond 85.35%Vals · MMLU Pro 79.7%Vals · AIME 91.25%
Vectara · Factual consistency 79.8%
Badge
[](https://publicai.io/model-index/m/grok-4-fast)