‹ PublicAI Index
The LLM benchmark aggregator.
GPT-5
OpenAI
Strongest in Secure code (#10 of 274), weakest in Toxicity avoidance (#112 of 272). Above par in 24 of 26 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5 and Claude Fable 5.1 and ahead of Gemini 3.8 Flash and GPT-5.2.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Safety58.8−2.7#12/3373/3
Secure code65.4−1.1#10/2741/1
Harm refusal59.3−2.3#15/3002/2
Jailbreak resistance62.6−2.3#17/2721/1
Fairness66.2−3.8#20/3002/2
Safe-prompt compliance56.6−3.8#23/821/1
Factual grounding37.1−33.5#91/1011/1
Toxicity avoidance54.8−4.2#112/2721/1
Professional56.7−4.9#20/1681/1
Legal59.2−7.9#16/1511/1
Medical58.5−8.3#17/1401/1
Finance54.2−8.7#56/1511/1
Core abilities56.8−10.6#28/2041/3
General intelligence59.8−10.3#29/2041/3
Knowledge55.5−8#44/1381/2
Academic knowledge56.6−7.3#40/1231/1
Reasoning53.7−13.2#53/1783/4
Mathematics57.2−8.5#23/1401/2
Science56.6−6.9#46/1221/1
Reasoning49.2−19.9#64/1402/3
Coding52.3−16.8#57/1652/5
Agentic coding52.6−16.1#55/1572/4
Human preference60.4−7.1#83/3421/1
Human preference60.4−7.1#83/3421/1
Agents52.1−16#94/2681/5
Knowledge work53.2−20.4#62/1781/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for GPT-5, left for the other.
§ 3 · Sources
Where the numbers come from
10 publications, 31 figures. Every one links to the page it was read from.
LMArena Text 1435
GDPval-AA 906
ARC-AGI-2 9.9%
Aider polyglot 88%
Kagi LLM Benchmark 72.7%
SimpleBench 56.7%
Vals · CaseLaw 66.45%Vals · LegalBench 86.02%Vals · CorpFin 61.07%Vals · TaxEval 73.39%Vals · MortgageTax 65.45%Vals · MedQA 96.32%Vals · MedCode 49.63%Vals · MedScribe 83.65%Vals · SWE-bench Verified 69% (Mini-SWE-agent)Vals · Vibe Code Bench 20.09% (OpenHands)Vals · GPQA Diamond 85.61%Vals · MMLU Pro 86.54%Vals · AIME 93.37%
HELM Safety · HarmBench 97.6%HELM Safety · SimpleSafetyTests 99.8%HELM Safety · Anthropic Red Team 99.1%HELM Safety · BBQ 96.8%HELM Safety · XSTest 97.1%
Enkrypt · Jailbreak risk 2.1%Enkrypt · Harmful content risk 6.7%Enkrypt · CBRN risk 7.1%Enkrypt · Toxicity risk 2.9%Enkrypt · Bias risk 40.1%Enkrypt · Insecure code risk 2.2%
Vectara · Factual consistency 84.9%
Badge
[](https://publicai.io/model-index/m/gpt-5)