Request a pilot
‹ PublicAI Index

The LLM benchmark aggregator.

Grok 4.6

SpaceXAI

Strongest in Legal (#2 of 151), weakest in Toxicity avoidance (#125 of 272). Above par in 29 of 30 scopes. Among the models it meets almost everywhere, it finishes behind Muse Spark 1.3 and Claude Opus 5.5 and ahead of GPT-5.5 and GPT-6 Sol.

§ 1 · Profile

What it is good at

Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.

Agents64#10/268
Knowledge work68.2#9/178
Professional57.6#12/168
Legal64.1#2/151
Biology research55.6#5/16
Cybersecurity51.3#6/9
Medical55.6#33/140
Finance55.9#35/151
Knowledge58.5#13/138
Academic knowledge60.2#11/123
Reasoning59.2#16/178
Science63#3/122
Reasoning63.1#11/140
Mathematics52.8#66/140
Safety58.2#17/337
Secure code66.3#4/274
Fairness70#7/300
Jailbreak resistance56.5#93/272
Harm refusal53.1#106/300
Toxicity avoidance54.3#125/272
Core abilities57.5#22/204
Language58.3#16/57
Instruction following57#18/57
General intelligence60.7#24/204
Data analysis50.8#33/57
Coding56.8#25/165
Agentic coding58.9#16/157
Code generation48.8#50/77
Human preference62.1#59/342
Human preference62.1#59/342

§ 2 · Head to head

What it beats, and what beats it

The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Grok 4.6, left for the other.

Grok Build 0.1won 19 · lost 0
Mistral Smallwon 19 · lost 0
Mistral Large 3won 19 · lost 0
Grok 3won 20 · lost 0
Command Awon 22 · lost 0
GPT-5.4 Miniwon 22 · lost 0
DeepSeek V3won 25 · lost 0
Grok 4.3won 28 · lost 0
DeepSeek V4 Flashwon 29 · lost 1
GPT-6 Lunawon 25 · lost 1
GPT-4.1 Nanowon 23 · lost 1
GPT-4.1 Miniwon 23 · lost 1
Gemini 2.5 Flashwon 23 · lost 1
DeepSeek V3.2won 23 · lost 1
GPT-5.4 Nanowon 22 · lost 1
Grok 4won 22 · lost 1
Kimi K2won 22 · lost 1
GLM-4.6won 22 · lost 1
Gemini 2.0 Flashwon 21 · lost 1
Qwen3.6 Pluswon 21 · lost 1
Grok 4.1 Fastwon 20 · lost 1
Gemma 4 31Bwon 20 · lost 1
Kimi K2.7 Codewon 18 · lost 1
GLM-4.5won 18 · lost 1
Ling 3.0 Flashwon 17 · lost 1
Gemini 1.5 Prowon 17 · lost 1
Gemini 3.1 Flash Litewon 17 · lost 1
Gemini 2.5 Prowon 23 · lost 2
GPT-4o Miniwon 22 · lost 2
GPT OSS 120Bwon 22 · lost 2
GPT-4.1won 21 · lost 2
Llama 4 Scout Instructwon 20 · lost 2
Qwen3.8 27Bwon 20 · lost 2
Gemini 3.5 Flash Litewon 20 · lost 2
GPT OSS 20Bwon 19 · lost 2
Kimi K2 Thinkingwon 18 · lost 2
Kimi K2.6won 25 · lost 3
Grok 4.5won 25 · lost 3
Qwen3.7 Maxwon 25 · lost 3
DeepSeek R1won 21 · lost 3
Kimi K2.5won 21 · lost 3
Qwen3 235B A22B Instructwon 20 · lost 3
Mimo V2.5 Prowon 20 · lost 3
GLM-5.1won 19 · lost 3
GPT-4owon 18 · lost 3
Claude 3.7 Sonnetwon 18 · lost 3
MiniMax M3won 24 · lost 4
GLM-5.2won 24 · lost 4
Claude Sonnet 4won 21 · lost 4
GPT-5 Miniwon 20 · lost 4
Claude Haiku 4.5won 20 · lost 4
Claude Opus 4won 19 · lost 4
Inklingwon 23 · lost 5
Claude 3.5 Haikuwon 16 · lost 4
O1won 16 · lost 4
O4 Miniwon 19 · lost 5
O3 Miniwon 18 · lost 5
Claude Sonnet 4.6won 20 · lost 6
GPT-5 Nanowon 16 · lost 5
MiniMax M2.7won 16 · lost 5
GPT-5.6 Lunawon 22 · lost 7
DeepSeek V4 Prowon 22 · lost 7
Gemini 3.6 Flashwon 22 · lost 7
Claude 3.5 Sonnetwon 15 · lost 5
O3won 18 · lost 6
GPT-5won 18 · lost 6
GLM-5.3 Flashwon 21 · lost 7
GPT-5.2won 20 · lost 7
GLM-5.3won 17 · lost 6
GPT-5.6 Terrawon 20 · lost 8
Claude Sonnet 4.5won 17 · lost 7
Claude Opus 4.6won 17 · lost 7
GPT-5.1won 15 · lost 7
Claude Opus 4.5won 15 · lost 7
Qwen3.8 Maxwon 19 · lost 9
Muse Spark 1.1won 19 · lost 9
Gemini 3.8 Flashwon 20 · lost 10
Claude Sonnet 5won 18 · lost 10
GPT-5.4won 18 · lost 10
Gemini 3.5 Flashwon 18 · lost 10
Muse Spark 1.2won 14 · lost 8
DeepSeek V4.1 Flashwon 13 · lost 10
Gemini 3.1 Prowon 13 · lost 10
Gemini 3.7 Flashwon 12 · lost 10
Kimi K3won 13 · lost 11
GPT-6 Solwon 14 · lost 12
GPT-5.5won 15 · lost 13
Gemini 3 Prowon 10 · lost 10
Grok 4.7won 13 · lost 13
Claude Opus 4.7won 14 · lost 14
GPT-5.6 Solwon 14 · lost 16
Claude Opus 4.8won 11 · lost 17
Mimo V2.6 Prowon 7 · lost 13
Claude Opus 5won 10 · lost 20
GPT-6 Astrawon 7 · lost 15
Claude Fable 5won 4 · lost 18
Claude Fable 5.1won 2 · lost 21
Claude Opus 5.5won 2 · lost 24
Muse Spark 1.3won 1 · lost 18

§ 3 · Sources

Where the numbers come from

10 publications, 38 figures. Every one links to the page it was read from.

LMArena Text ↗1 measureread 2026-09-26
LMArena Text 1453
Artificial Analysis ↗1 measureread 2026-09-26
Artificial Analysis Intelligence Index 44
GDPval-AA ↗1 measureread 2026-09-26
GDPval-AA 1632
AA-Briefcase ↗1 measureread 2026-09-26
AA-Briefcase 1549
Terminal-Bench ↗1 measureread 2026-09-26
Terminal-Bench 20.3% (Grok Build)
ARC-AGI-2 ↗1 measureread 2026-09-26
ARC-AGI-2 67.1%
LiveBench ↗8 measuresread 2026-09-26
LiveBench 78LiveBench · Reasoning 90.5LiveBench · Coding 76.8LiveBench · Agentic Coding 57LiveBench · Mathematics 92.6LiveBench · Data Analysis 73.9LiveBench · Language 83.7LiveBench · Instruction Following 71.9
SimpleBench ↗1 measureread 2026-09-26
SimpleBench 75.9%
Vals.ai ↗17 measuresread 2026-09-26
Vals · Legal Research Bench 48.08%Vals · LegalBench 86.31%Vals · Harvey Legal Agent Benchmark 15.83%Vals · Finance Agent 53.68%Vals · CorpFin 66.16%Vals · TaxEval 71.1%Vals · MortgageTax 64.19%Vals · MedCode 44.71%Vals · MedScribe 86.53%Vals · BioMysteryBench 72.22%Vals · CyberBench 66.01%Vals · SWE-bench Verified 95.6% (Mini-SWE-agent)Vals · Vibe Code Bench 76.24% (OpenHands)Vals · Code Migration 44.57%Vals · GPQA Diamond 94.7%Vals · MMLU Pro 89.4%Vals · ProofBench 51%
Enkrypt AI Safety Leaderboard ↗6 measuresread 2026-09-26
Enkrypt · Jailbreak risk 7.3%Enkrypt · Harmful content risk 0.6%Enkrypt · CBRN risk 24%Enkrypt · Toxicity risk 3.3%Enkrypt · Bias risk 35.1%Enkrypt · Insecure code risk 0.4%

Badge

PublicAI Index badge for Grok 4.6[![PublicAI Index](https://publicai.io/model-index/badge?model=grok-4-6)](https://publicai.io/model-index/m/grok-4-6)