‹ PublicAI Index
The LLM benchmark aggregator.
GPT-6.1 Sol
OpenAI
Strongest in Data analysis (#2 of 58), weakest in Finance (#68 of 157). Above par in 21 of 22 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5 and Claude Opus 5.5 and ahead of Muse Spark 1.2 and Kimi K3.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Core abilities65.9−1.2#3/2122/3
Data analysis65.1−0.5#2/581/1
Language69.9−1.1#2/581/1
General intelligence67−2.9#5/2122/3
Instruction following60.7−12.8#12/581/1
Reasoning65.6−0.6#4/1834/4
Reasoning67.7−0.8#3/1413/3
Mathematics64.4−0.5#5/1452/2
Coding60−8.5#11/1682/5
Agentic coding61.5−6.5#9/1612/4
Code generation55.6−14.4#26/781/2
Agents62.9−4.6#15/2772/5
Knowledge work66.8−5.9#13/1872/2
Human preference64.8−4#19/3451/1
Human preference64.8−4#19/3451/1
Professional54.7−9.4#34/1771/1
Biology research67−2.6#2/231/1
IT operations69.9−3.8#2/161/1
Medical57.7−8.9#22/1461/1
Cybersecurity27.5−35#41/431/1
Legal52.6−14.1#60/1571/1
Finance53.3−9.3#68/1571/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for GPT-6.1 Sol, left for the other.
§ 3 · Sources
Where the numbers come from
8 publications, 25 figures. Every one links to the page it was read from.
LMArena Text 1483
Artificial Analysis Intelligence Index 52
GDPval-AA 1575
AA-Briefcase 1564
ARC-AGI-2 94.2%
LiveBench 81.6LiveBench · Reasoning 92.6LiveBench · Coding 80.4LiveBench · Agentic Coding 54.5LiveBench · Mathematics 96.8LiveBench · Data Analysis 82.7LiveBench · Language 90.1LiveBench · Instruction Following 74.2
SimpleBench 82.9%
Vals · Legal Research Bench 38.46%Vals · Harvey Legal Agent Benchmark 5.42%Vals · Finance Agent 52.03%Vals · MedCode 48.84%Vals · MedScribe 86.45%Vals · BioMysteryBench 79.63%Vals · CyberBench 39.29%Vals · SRE Bench 50.76%Vals · Vibe Code Bench 88.93% (OpenHands)Vals · Code Migration 65.12%Vals · ProofBench 99%
Badge
[](https://publicai.io/model-index/m/gpt-6-1-sol)