‹ PublicAI Index
The LLM benchmark aggregator.
GPT-6 Sol
OpenAI
Strongest in Data analysis (#3 of 57), weakest in Toxicity avoidance (#224 of 272). Above par in 25 of 27 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5 and Claude Fable 5.1 and ahead of Muse Spark 1.1 and Claude Opus 4.8.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Core abilities60.6−6.8#8/2042/3
Data analysis63.1−3#3/571/1
General intelligence63.6−6.5#7/2042/3
Language61.3−10.2#13/571/1
Instruction following51.2−22.5#29/571/1
Reasoning60.8−6.1#8/1783/4
Mathematics62−3.7#7/1402/2
Reasoning60.6−8.5#16/1402/3
Coding59.9−9.2#10/1652/5
Agentic coding60.3−8.4#12/1572/4
Code generation59−11#13/771/2
Agents62.2−5.9#20/2682/5
Knowledge work65.9−7.7#18/1782/2
Human preference62.5−5#52/3421/1
Human preference62.5−5#52/3421/1
Professional52−9.6#60/1681/1
Biology research60−8.5#4/161/1
Medical54.9−11.9#38/1401/1
Finance51.6−11.3#78/1511/1
Legal47−20.1#93/1511/1
Safety52.4−9.1#120/3372/3
Factual grounding58.8−11.8#28/1011/1
Fairness54−16#77/3001/2
Secure code57.6−8.9#90/2741/1
Harm refusal50.3−11.3#151/3001/2
Jailbreak resistance50.7−14.2#178/2721/1
Toxicity avoidance46.6−12.4#224/2721/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for GPT-6 Sol, left for the other.
§ 3 · Sources
Where the numbers come from
9 publications, 29 figures. Every one links to the page it was read from.
LMArena Text 1457
Artificial Analysis Intelligence Index 48
GDPval-AA 1487
AA-Briefcase 1483
LiveBench 79.3LiveBench · Reasoning 88.7LiveBench · Coding 81.8LiveBench · Agentic Coding 52.9LiveBench · Mathematics 96.4LiveBench · Data Analysis 81.2LiveBench · Language 85.3LiveBench · Instruction Following 68.6
SimpleBench 73.1%
Vals · Legal Research Bench 28.85%Vals · Harvey Legal Agent Benchmark 1.67%Vals · Finance Agent 49.05%Vals · MedCode 47.07%Vals · MedScribe 82.03%Vals · BioMysteryBench 74.81%Vals · Vibe Code Bench 87.82% (OpenHands)Vals · Code Migration 57.2%Vals · ProofBench 83%
Enkrypt · Jailbreak risk 12.2%Enkrypt · Harmful content risk 0%Enkrypt · CBRN risk 31.8%Enkrypt · Toxicity risk 8.7%Enkrypt · Bias risk 71.6%Enkrypt · Insecure code risk 18.2%
Vectara · Factual consistency 93.5%
Badge
[](https://publicai.io/model-index/m/gpt-6-sol)