Request a pilot
‹ PublicAI Index

The LLM benchmark aggregator.

GPT-5.6 Terra

OpenAI

Strongest in Mathematics (#12 of 140), weakest in Toxicity avoidance (#214 of 272). Above par in 31 of 33 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5 and Claude Fable 5.1 and ahead of GPT-5 Mini and Claude Opus 4.5.

§ 1 · Profile

What it is good at

Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.

Reasoning59#17/178
Mathematics59#12/140
Science60.3#24/122
Reasoning58.4#28/140
Expert reasoning57.2—/0✱
Coding57.3#21/165
Agentic coding58.7#18/157
Code generation51.6#38/77
Agents60.3#29/268
Knowledge work63.5#25/178
Tool use57.8—/81✱
Workflow automation54.8—/0✱
Professional54.9#35/168
Finance58.2#16/151
Legal52.7#57/151
Medical53.1#61/140
Safety55.8#38/337
Secure code64.4#17/274
Fairness61#38/300
Harm refusal54.3#82/300
Jailbreak resistance55.1#116/272
Toxicity avoidance48.1#214/272
Human preference63.3#41/342
Human preference63.3#41/342
Core abilities55.1#42/204
Data analysis59.9#12/57
Language56.8#18/57
Instruction following44.2#38/57
General intelligence56.6#51/204
Long context56—/0✱
Knowledge55.7#42/138
Academic knowledge56.8#39/123
Factuality52.1—/0✱

§ 2 · Head to head

What it beats, and what beats it

The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for GPT-5.6 Terra, left for the other.

GPT-5.4 Miniwon 22 · lost 0
DeepSeek V3won 25 · lost 0
DeepSeek V4 Flashwon 29 · lost 1
GPT-4.1 Nanowon 24 · lost 1
GPT-5.4 Nanowon 22 · lost 1
Gemini 3.5 Flash Litewon 21 · lost 1
GPT OSS 20Bwon 20 · lost 1
Mistral Smallwon 19 · lost 1
Grok 3won 19 · lost 1
Grok 4.3won 26 · lost 2
GPT-4.1 Miniwon 23 · lost 2
Gemma 4 31Bwon 23 · lost 2
GPT-4o Miniwon 22 · lost 2
Kimi K2won 22 · lost 2
Llama 4 Scout Instructwon 21 · lost 2
Command Awon 21 · lost 2
Gemini 2.0 Flashwon 20 · lost 2
Qwen3.6 Pluswon 20 · lost 2
Nemotron 3 Ultrawon 19 · lost 2
Gemini 2.5 Flashwon 22 · lost 3
DeepSeek V3.2won 22 · lost 3
GPT-6 Lunawon 22 · lost 3
GLM-4.6won 21 · lost 3
Kimi K2.5won 21 · lost 3
Mimo V2.5 Prowon 20 · lost 3
Grok 4.1 Fastwon 19 · lost 3
GLM-5.1won 19 · lost 3
GPT-4owon 18 · lost 3
Kimi K2.6won 24 · lost 4
Kimi K2 Thinkingwon 17 · lost 3
Claude Haiku 4.5won 21 · lost 4
GPT OSS 120Bwon 20 · lost 4
MiniMax M3won 27 · lost 6
Grok 4won 19 · lost 5
Inklingwon 26 · lost 7
Qwen3.7 Maxwon 22 · lost 6
O3 Miniwon 18 · lost 5
Qwen3.8 27Bwon 20 · lost 6
Claude 3.7 Sonnetwon 16 · lost 5
GPT-5.6 Lunawon 25 · lost 8
Claude 3.5 Haikuwon 15 · lost 5
O1won 15 · lost 5
DeepSeek R1won 18 · lost 6
DeepSeek V4 Prowon 22 · lost 8
O4 Miniwon 18 · lost 7
Claude Sonnet 4won 18 · lost 7
Gemini 2.5 Prowon 18 · lost 7
MiniMax M2.7won 15 · lost 6
GPT-4.1won 17 · lost 7
Qwen3 235B A22B Instructwon 17 · lost 7
Claude 3.5 Sonnetwon 14 · lost 6
GPT-5 Nanowon 15 · lost 7
Gemini 3.6 Flashwon 19 · lost 9
GLM-5.2won 22 · lost 11
Claude Sonnet 4.6won 17 · lost 9
Grok 4.5won 18 · lost 10
GPT-5.1won 14 · lost 8
Claude Opus 4won 14 · lost 9
GLM-5.3 Flashwon 17 · lost 11
Gemini 3.5 Flashwon 17 · lost 11
Claude Sonnet 4.5won 15 · lost 10
GPT-5.2won 16 · lost 12
Gemini 3.1 Prowon 13 · lost 10
O3won 14 · lost 11
GPT-5won 13 · lost 11
Qwen3.8 Maxwon 15 · lost 13
Claude Opus 4.5won 12 · lost 11
GPT-5 Miniwon 13 · lost 12
Claude Opus 4.6won 12 · lost 12
Claude Opus 4.7won 14 · lost 14
GPT-5.4won 13 · lost 15
Muse Spark 1.1won 13 · lost 15
Gemini 3.8 Flashwon 13 · lost 15
Claude Sonnet 5won 15 · lost 18
Grok 4.7won 11 · lost 14
GPT-6 Solwon 11 · lost 14
Gemini 3 Prowon 9 · lost 12
Muse Spark 1.2won 9 · lost 12
GLM-5.3won 10 · lost 14
DeepSeek V4.1 Flashwon 9 · lost 14
Gemini 3.7 Flashwon 8 · lost 14
Grok 4.6won 8 · lost 20
Claude Opus 4.8won 8 · lost 20
GPT-5.5won 8 · lost 20
GPT-5.6 Solwon 8 · lost 22
Kimi K3won 6 · lost 18
Claude Opus 5won 5 · lost 25
GPT-6 Astrawon 2 · lost 18
Claude Opus 5.5won 1 · lost 24
Claude Fable 5.1won 0 · lost 22
Claude Fable 5won 0 · lost 22

§ 3 · Sources

Where the numbers come from

12 publications, 52 figures. Every one links to the page it was read from.

LMArena Text ↗1 measureread 2026-09-26
LMArena Text 1465
Artificial Analysis ↗1 measureread 2026-09-26
Artificial Analysis Intelligence Index 42
GDPval-AA ↗1 measureread 2026-09-26
GDPval-AA 1432
AA-Briefcase ↗1 measureread 2026-09-26
AA-Briefcase 1334
Terminal-Bench ↗1 measureread 2026-09-26
Terminal-Bench 21.5% (Codex)
ARC-AGI-2 ↗1 measureread 2026-09-26
ARC-AGI-2 83.9%
LiveBench ↗8 measuresread 2026-09-26
LiveBench 77.9LiveBench · Reasoning 90.6LiveBench · Coding 78.2LiveBench · Agentic Coding 54.9LiveBench · Mathematics 94.9LiveBench · Data Analysis 79.3LiveBench · Language 82.9LiveBench · Instruction Following 64.6
Kagi LLM Benchmark ↗1 measureread 2026-09-26
Kagi LLM Benchmark 51.3%
SimpleBench ↗1 measureread 2026-09-26
SimpleBench 48.9%
Vals.ai ↗15 measuresread 2026-09-26
Vals · Legal Research Bench 41.35%Vals · LegalBench 85.11%Vals · Harvey Legal Agent Benchmark 0.83%Vals · Finance Agent 54.44%Vals · CorpFin 65.31%Vals · TaxEval 76.17%Vals · MortgageTax 67.33%Vals · MedCode 43.41%Vals · MedScribe 82.87%Vals · SWE-bench Verified 95.4% (Mini-SWE-agent)Vals · Vibe Code Bench 74.59% (OpenHands)Vals · Code Migration 47.8%Vals · GPQA Diamond 90.91%Vals · MMLU Pro 86.66%Vals · ProofBench 74%
Enkrypt AI Safety Leaderboard ↗6 measuresread 2026-09-26
Enkrypt · Jailbreak risk 8.5%Enkrypt · Harmful content risk 2.2%Enkrypt · CBRN risk 20%Enkrypt · Toxicity risk 7.6%Enkrypt · Bias risk 60.7%Enkrypt · Insecure code risk 4.4%
GDPVal-AA 1503tau3-Banking 28.7%Toolathlon Verified 64.8%Automation Bench Public 28%Apex-Agents (pass@1) 25.4%MCPMark 74%WildClawBench 60%Terminal-Bench 2.1 75.7%SciCode 50.1%Humanity's Last Exam (without tools) 38.5%GPQA Diamond 89.6%CritPt 22.9%AA-LCR 73.3%AA-Omniscience Accuracy 45%AA-Omniscience Non-Hallucination 10%

Badge

PublicAI Index badge for GPT-5.6 Terra[![PublicAI Index](https://publicai.io/model-index/badge?model=gpt-5-6-terra)](https://publicai.io/model-index/m/gpt-5-6-terra)

✱ Placed by a one-off publication, not a board that re-ran the model. ✱✱ marks a figure the model’s own publisher printed, which counts for less again. Scores are 0–100 on the PublicAI Index scale, 50 = the average of the models each source lists. Scores belong to their publishers.

Snapshot 2026-09-26 · the full index · JSON API