‹ PublicAI Index
The LLM benchmark aggregator.
GPT-5.6 Sol
OpenAI
Strongest in Science (#2 of 122), weakest in Toxicity avoidance (#227 of 272). Above par in 32 of 37 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5 and Claude Opus 5.5 and ahead of Claude Opus 4.7 and Grok 4.6.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Reasoning62.8−4.1#6/1784/4
Science63.4−0.1#2/1221/1
Mathematics61.8−3.9#8/1402/2
Reasoning63.5−5.6#10/1403/3
Expert reasoning46.3−14.6—/0✱✱
Coding62.1−7#7/1653/5
Code generation63.3−6.7#4/771/2
Agentic coding61.7−7#9/1573/4
Repository Q&A44.5−22.2—/0✱✱
Core abilities61−6.4#7/2042/3
Language65.8−5.7#6/571/1
Data analysis60.8−5.3#7/571/1
Instruction following56.8−16.9#19/571/1
General intelligence60.7−9.4#23/2042/3
Human preference65−2.5#16/3421/1
Human preference65−2.5#16/3421/1
Knowledge58.2−5.3#18/1381/2
Academic knowledge59.8−4.1#16/1231/1
Multimodal understanding53.2−18.9—/19✱0/1
Agents62.1−6#21/2683/5
Web research55.1−9.1#2/41/1
Knowledge work66.9−6.7#14/1782/2
Workflow automation48.6−13.4—/0✱✱
Professional56.3−5.3#26/1681/1
IT operations60.3−13.5#3/111/1
Biology research53.6−14.9#7/161/1
Finance57.3−5.6#25/1511/1
Legal56.1−11#30/1511/1
Medical54.5−12.3#46/1401/1
Cybersecurity60.2leads—/9✱0/1
Safety53.1−8.4#101/3372/3
Secure code63−3.5#32/2741/1
Harm refusal55−6.6#75/3001/2
Fairness53.7−16.3#80/3001/2
Factual grounding43.9−26.7#82/1011/1
Jailbreak resistance54.2−10.7#128/2721/1
Toxicity avoidance46.3−12.7#227/2721/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for GPT-5.6 Sol, left for the other.
§ 3 · Sources
Where the numbers come from
12 publications, 56 figures. Every one links to the page it was read from.
LMArena Text 1483
GDPval-AA 1588
AA-Briefcase 1487
Terminal-Bench 37.3% (Codex)
ARC-AGI-2 92.5%
LiveBench 81LiveBench · Reasoning 91.7LiveBench · Coding 83.9LiveBench · Agentic Coding 56.2LiveBench · Mathematics 96.2LiveBench · Data Analysis 79.8LiveBench · Language 87.7LiveBench · Instruction Following 71.8
Kagi LLM Benchmark 67%
SimpleBench 64.8%
Vals · Legal Research Bench 48.08%Vals · LegalBench 86.97%Vals · Harvey Legal Agent Benchmark 2.5%Vals · Finance Agent 53.76%Vals · CorpFin 64.38%Vals · TaxEval 74.78%Vals · MortgageTax 67.29%Vals · MedCode 43.97%Vals · MedScribe 85.23%Vals · BioMysteryBench 71.11%Vals · SRE Bench 30.53%Vals · Web Search Index 45.24% (Exa)Vals · SWE-bench Verified 96.2% (Mini-SWE-agent)Vals · Vibe Code Bench 80.5% (OpenHands)Vals · Code Migration 52.92%Vals · GPQA Diamond 95.2%Vals · MMLU Pro 89.1%Vals · ProofBench 83%
Enkrypt · Jailbreak risk 9.2%Enkrypt · Harmful content risk 0.6%Enkrypt · CBRN risk 19.2%Enkrypt · Toxicity risk 8.9%Enkrypt · Bias risk 72.1%Enkrypt · Insecure code risk 7.1%
Vectara · Factual consistency 87.6%
GPQA Diamond (Pass@1) 94.1%HLE (Pass@1) 44.5%Terminal-Bench 2.1 (Pass@1) 88.8%Terminal-Bench 3.0 (Pass@1) 34.4%Terminal-Bench 4.0 (Pass@1) 39.9%DeepSWE v1.1 (Resolved) 73%ProgramBench (Almost@1) 23%NL2Repo-Bench (Score) 56.8%CyberGym (Pass@1) 84.5%SEC-Bench Pro (Pass@1) 74.3%ExploitGym (Pass@1) 33.7%AutomationBench (Pass@1) 45.8%Agent's Last Exam (Pass@1) 26.7%Chartography w/ tools (Pass@1) 79.9%BabyVision w/ tools (Pass@1) 88.9%ZeroBench-main w/ tools (Pass@5) 53%
Badge
[](https://publicai.io/model-index/m/gpt-5-6-sol)