‹ PublicAI Index
The LLM benchmark aggregator.
GPT-5.6 Terra
OpenAI
Strongest in Mathematics (#12 of 140), weakest in Toxicity avoidance (#214 of 272). Above par in 31 of 33 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5 and Claude Fable 5.1 and ahead of GPT-5 Mini and Claude Opus 4.5.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Reasoning59−7.9#17/1784/4
Mathematics59−6.7#12/1402/2
Science60.3−3.2#24/1221/1
Reasoning58.4−10.7#28/1403/3
Expert reasoning57.2−3.7—/0✱✱
Coding57.3−11.8#21/1653/5
Agentic coding58.7−10#18/1573/4
Code generation51.6−18.4#38/771/2
Agents60.3−7.8#29/2682/5
Knowledge work63.5−10.1#25/1782/2
Tool use57.8−16.2—/81✱0/1
Workflow automation54.8−7.2—/0✱✱
Professional54.9−6.7#35/1681/1
Finance58.2−4.7#16/1511/1
Legal52.7−14.4#57/1511/1
Medical53.1−13.7#61/1401/1
Safety55.8−5.7#38/3371/3
Secure code64.4−2.1#17/2741/1
Fairness61−9#38/3001/2
Harm refusal54.3−7.3#82/3001/2
Jailbreak resistance55.1−9.8#116/2721/1
Toxicity avoidance48.1−10.9#214/2721/1
Human preference63.3−4.2#41/3421/1
Human preference63.3−4.2#41/3421/1
Core abilities55.1−12.3#42/2043/3
Data analysis59.9−6.2#12/571/1
Language56.8−14.7#18/571/1
Instruction following44.2−29.5#38/571/1
General intelligence56.6−13.5#51/2043/3
Long context56−8.6—/0✱✱
Knowledge55.7−7.8#42/1381/2
Academic knowledge56.8−7.1#39/1231/1
Factuality52.1−6.9—/0✱✱
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for GPT-5.6 Terra, left for the other.
§ 3 · Sources
Where the numbers come from
12 publications, 52 figures. Every one links to the page it was read from.
LMArena Text 1465
Artificial Analysis Intelligence Index 42
GDPval-AA 1432
AA-Briefcase 1334
Terminal-Bench 21.5% (Codex)
ARC-AGI-2 83.9%
LiveBench 77.9LiveBench · Reasoning 90.6LiveBench · Coding 78.2LiveBench · Agentic Coding 54.9LiveBench · Mathematics 94.9LiveBench · Data Analysis 79.3LiveBench · Language 82.9LiveBench · Instruction Following 64.6
Kagi LLM Benchmark 51.3%
SimpleBench 48.9%
Vals · Legal Research Bench 41.35%Vals · LegalBench 85.11%Vals · Harvey Legal Agent Benchmark 0.83%Vals · Finance Agent 54.44%Vals · CorpFin 65.31%Vals · TaxEval 76.17%Vals · MortgageTax 67.33%Vals · MedCode 43.41%Vals · MedScribe 82.87%Vals · SWE-bench Verified 95.4% (Mini-SWE-agent)Vals · Vibe Code Bench 74.59% (OpenHands)Vals · Code Migration 47.8%Vals · GPQA Diamond 90.91%Vals · MMLU Pro 86.66%Vals · ProofBench 74%
Enkrypt · Jailbreak risk 8.5%Enkrypt · Harmful content risk 2.2%Enkrypt · CBRN risk 20%Enkrypt · Toxicity risk 7.6%Enkrypt · Bias risk 60.7%Enkrypt · Insecure code risk 4.4%
GDPVal-AA 1503tau3-Banking 28.7%Toolathlon Verified 64.8%Automation Bench Public 28%Apex-Agents (pass@1) 25.4%MCPMark 74%WildClawBench 60%Terminal-Bench 2.1 75.7%SciCode 50.1%Humanity's Last Exam (without tools) 38.5%GPQA Diamond 89.6%CritPt 22.9%AA-LCR 73.3%AA-Omniscience Accuracy 45%AA-Omniscience Non-Hallucination 10%
Badge
[](https://publicai.io/model-index/m/gpt-5-6-terra)