‹ PublicAI Index
The LLM benchmark aggregator.
GPT-5.1
OpenAI
Strongest in Legal (#3 of 151), weakest in Toxicity avoidance (#159 of 272). Above par in 20 of 24 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5 and Claude Fable 5.1 and ahead of Kimi K2.6 and Grok 4.7.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Professional58.9−2.7#8/1681/1
Legal63.8−3.3#3/1511/1
Medical61.5−5.3#4/1401/1
Finance54.6−8.3#50/1511/1
Safety57.3−4.2#23/3373/3
Safe-prompt compliance59.1−1.3#10/821/1
Secure code64.5−2#15/2741/1
Harm refusal58.6−3#20/3002/2
Fairness58.2−11.8#46/3002/2
Factual grounding44.6−26#78/1011/1
Jailbreak resistance56.6−8.3#90/2721/1
Toxicity avoidance52.7−6.3#159/2721/1
Knowledge55.4−8.1#45/1381/2
Academic knowledge56.4−7.5#41/1231/1
Reasoning53.7−13.2#50/1783/4
Mathematics57.2−8.5#25/1401/2
Science57.4−6.1#41/1221/1
Reasoning49−20.1#65/1402/3
Human preference62.4−5.1#55/3421/1
Human preference62.4−5.1#55/3421/1
Coding46.7−22.4#100/1651/5
Agentic coding46.3−22.4#95/1571/4
Agents51.1−17#104/2681/5
Knowledge work51.6−22#69/1781/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for GPT-5.1, left for the other.
§ 3 · Sources
Where the numbers come from
8 publications, 29 figures. Every one links to the page it was read from.
LMArena Text 1456
GDPval-AA 813
ARC-AGI-2 17.6%
SimpleBench 53.2%
Vals · CaseLaw 73.42%Vals · LegalBench 85.68%Vals · CorpFin 63.83%Vals · TaxEval 74.86%Vals · MortgageTax 61.37%Vals · MedQA 96.38%Vals · MedCode 52.73%Vals · MedScribe 88.09%Vals · SWE-bench Verified 69.8% (Mini-SWE-agent)Vals · Vibe Code Bench 24.61% (OpenHands)Vals · GPQA Diamond 86.62%Vals · MMLU Pro 86.38%Vals · AIME 93.33%
HELM Safety · HarmBench 97.6%HELM Safety · SimpleSafetyTests 99.8%HELM Safety · Anthropic Red Team 99.5%HELM Safety · BBQ 88.7%HELM Safety · XSTest 98.2%
Enkrypt · Jailbreak risk 7.2%Enkrypt · Harmful content risk 16.1%Enkrypt · CBRN risk 6.5%Enkrypt · Toxicity risk 4.4%Enkrypt · Bias risk 56.3%Enkrypt · Insecure code risk 4%
Vectara · Factual consistency 87.9%
Badge
[](https://publicai.io/model-index/m/gpt-5-1)