‹ PublicAI Index
The LLM benchmark aggregator.
GPT-5.5
OpenAI
Strongest in Data analysis (#2 of 57), weakest in Toxicity avoidance (#217 of 272). Above par in 28 of 30 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5.1 and Claude Fable 5 and ahead of Gemini 3.8 Flash and Muse Spark 1.1.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Core abilities63.7−3.7#4/2042/3
Data analysis63.8−2.3#2/571/1
General intelligence67.2−2.9#5/2042/3
Language65.3−6.2#7/571/1
Instruction following54.9−18.8#23/571/1
Reasoning61.6−5.3#7/1784/4
Mathematics59.8−5.9#9/1401/2
Science61.9−1.6#10/1221/1
Reasoning62.4−6.7#12/1403/3
Human preference64.8−2.7#18/3421/1
Human preference64.8−2.7#18/3421/1
Coding57.6−11.5#19/1652/5
Code generation59.6−10.4#9/771/2
Agentic coding57−11.7#28/1572/4
Knowledge57.2−6.3#24/1381/2
Academic knowledge58.6−5.3#21/1231/1
Professional56.3−5.3#25/1681/1
IT operations42.2−31.6#7/111/1
Finance58.5−4.4#13/1511/1
Medical58.5−8.3#18/1401/1
Legal57.1−10#26/1511/1
Agents57.7−10.4#46/2682/5
Knowledge work60−13.6#35/1782/2
Safety55.2−6.3#51/3372/3
Secure code63.9−2.6#19/2741/1
Fairness60.5−9.5#40/3001/2
Factual grounding51.7−18.9#45/1011/1
Harm refusal54.4−7.2#81/3001/2
Jailbreak resistance54.2−10.7#127/2721/1
Toxicity avoidance47.6−11.4#217/2721/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for GPT-5.5, left for the other.
§ 3 · Sources
Where the numbers come from
10 publications, 37 figures. Every one links to the page it was read from.
LMArena Text 1481
GDPval-AA 1336
AA-Briefcase 1137
ARC-AGI-2 85%
LiveBench 80.2LiveBench · Reasoning 89.7LiveBench · Coding 82.1LiveBench · Agentic Coding 54LiveBench · Mathematics 95.9LiveBench · Data Analysis 81.6LiveBench · Language 87.4LiveBench · Instruction Following 70.7
Kagi LLM Benchmark 88.8%
SimpleBench 69%
Vals · Legal Research Bench 40.38%Vals · CaseLaw 66.24%Vals · LegalBench 86.52%Vals · Harvey Legal Agent Benchmark 3.75%Vals · Finance Agent 51.76%Vals · CorpFin 68.42%Vals · TaxEval 74.98%Vals · MortgageTax 68.76%Vals · MedCode 49.1%Vals · MedScribe 86.87%Vals · SRE Bench 3.82%Vals · SWE-bench Verified 82.6% (Mini-SWE-agent)Vals · Vibe Code Bench 69.85% (OpenHands)Vals · Code Migration 45.16%Vals · GPQA Diamond 93.18%Vals · MMLU Pro 88.14%
Enkrypt · Jailbreak risk 9.2%Enkrypt · Harmful content risk 0.6%Enkrypt · CBRN risk 20.7%Enkrypt · Toxicity risk 8%Enkrypt · Bias risk 61.5%Enkrypt · Insecure code risk 5.3%
Vectara · Factual consistency 90.7%
Badge
[](https://publicai.io/model-index/m/gpt-5-5)