‹ PublicAI Index
The LLM benchmark aggregator.
Grok 4.6
SpaceXAI
Strongest in Legal (#2 of 151), weakest in Toxicity avoidance (#125 of 272). Above par in 29 of 30 scopes. Among the models it meets almost everywhere, it finishes behind Muse Spark 1.3 and Claude Opus 5.5 and ahead of GPT-5.5 and GPT-6 Sol.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Agents64−4.1#10/2682/5
Knowledge work68.2−5.4#9/1782/2
Professional57.6−4#12/1681/1
Legal64.1−3#2/1511/1
Biology research55.6−12.9#5/161/1
Cybersecurity51.3−8.4#6/91/1
Medical55.6−11.2#33/1401/1
Finance55.9−7#35/1511/1
Knowledge58.5−5#13/1381/2
Academic knowledge60.2−3.7#11/1231/1
Reasoning59.2−7.7#16/1784/4
Science63−0.5#3/1221/1
Reasoning63.1−6#11/1403/3
Mathematics52.8−12.9#66/1402/2
Safety58.2−3.3#17/3371/3
Secure code66.3−0.2#4/2741/1
Fairness70leads#7/3001/2
Jailbreak resistance56.5−8.4#93/2721/1
Harm refusal53.1−8.5#106/3001/2
Toxicity avoidance54.3−4.7#125/2721/1
Core abilities57.5−9.9#22/2042/3
Language58.3−13.2#16/571/1
Instruction following57−16.7#18/571/1
General intelligence60.7−9.4#24/2042/3
Data analysis50.8−15.3#33/571/1
Coding56.8−12.3#25/1653/5
Agentic coding58.9−9.8#16/1573/4
Code generation48.8−21.2#50/771/2
Human preference62.1−5.4#59/3421/1
Human preference62.1−5.4#59/3421/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Grok 4.6, left for the other.
§ 3 · Sources
Where the numbers come from
10 publications, 38 figures. Every one links to the page it was read from.
LMArena Text 1453
Artificial Analysis Intelligence Index 44
GDPval-AA 1632
AA-Briefcase 1549
Terminal-Bench 20.3% (Grok Build)
ARC-AGI-2 67.1%
LiveBench 78LiveBench · Reasoning 90.5LiveBench · Coding 76.8LiveBench · Agentic Coding 57LiveBench · Mathematics 92.6LiveBench · Data Analysis 73.9LiveBench · Language 83.7LiveBench · Instruction Following 71.9
SimpleBench 75.9%
Vals · Legal Research Bench 48.08%Vals · LegalBench 86.31%Vals · Harvey Legal Agent Benchmark 15.83%Vals · Finance Agent 53.68%Vals · CorpFin 66.16%Vals · TaxEval 71.1%Vals · MortgageTax 64.19%Vals · MedCode 44.71%Vals · MedScribe 86.53%Vals · BioMysteryBench 72.22%Vals · CyberBench 66.01%Vals · SWE-bench Verified 95.6% (Mini-SWE-agent)Vals · Vibe Code Bench 76.24% (OpenHands)Vals · Code Migration 44.57%Vals · GPQA Diamond 94.7%Vals · MMLU Pro 89.4%Vals · ProofBench 51%
Enkrypt · Jailbreak risk 7.3%Enkrypt · Harmful content risk 0.6%Enkrypt · CBRN risk 24%Enkrypt · Toxicity risk 3.3%Enkrypt · Bias risk 35.1%Enkrypt · Insecure code risk 0.4%
Badge
[](https://publicai.io/model-index/m/grok-4-6)