‹ PublicAI Index
The LLM benchmark aggregator.
GPT-5.4
OpenAI
Strongest in Mathematics (#10 of 140), weakest in Fairness (#217 of 300). Above par in 28 of 29 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5 and Claude Fable 5.1 and ahead of GPT-5.6 Terra and Gemini 3.1 Pro.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Reasoning59.6−7.3#14/1783/4
Mathematics59.7−6#10/1402/2
Science60.9−2.6#21/1221/1
Reasoning58.9−10.2#25/1402/3
Human preference64.2−3.3#28/3421/1
Human preference64.2−3.3#28/3421/1
Core abilities56.7−10.7#29/2042/3
Data analysis59.9−6.2#11/571/1
Language56.2−15.3#21/571/1
Instruction following54−19.7#24/571/1
General intelligence56.8−13.3#48/2042/3
Knowledge56.5−7#30/1381/2
Academic knowledge57.8−6.1#27/1231/1
Coding53.4−15.7#44/1652/5
Code generation50.2−19.8#44/771/2
Agentic coding54.3−14.4#45/1572/4
Professional53.8−7.8#45/1681/1
Finance57.1−5.8#28/1511/1
Legal52.1−15#62/1511/1
Medical52.6−14.2#64/1401/1
Safety55.3−6.2#47/3372/3
Secure code63.7−2.8#23/2741/1
Factual grounding57.5−13.1#31/1011/1
Harm refusal56−5.6#48/3001/2
Jailbreak resistance57.7−7.2#72/2721/1
Toxicity avoidance55.1−3.9#103/2721/1
Fairness44−26#217/3001/2
Agents55.9−12.2#57/2681/5
Knowledge work58.8−14.8#38/1781/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for GPT-5.4, left for the other.
§ 3 · Sources
Where the numbers come from
8 publications, 34 figures. Every one links to the page it was read from.
LMArena Text 1475
GDPval-AA 1233
ARC-AGI-2 74%
LiveBench 78LiveBench · Reasoning 88.1LiveBench · Coding 77.5LiveBench · Agentic Coding 53.8LiveBench · Mathematics 94.1LiveBench · Data Analysis 79.3LiveBench · Language 82.6LiveBench · Instruction Following 70.2
Kagi LLM Benchmark 63.8%
Vals · CaseLaw 63.77%Vals · LegalBench 86.04%Vals · Harvey Legal Agent Benchmark 0%Vals · CorpFin 65.27%Vals · TaxEval 73.96%Vals · MortgageTax 68.32%Vals · MedQA 96.09%Vals · MedCode 41.29%Vals · MedScribe 77.55%Vals · SWE-bench Verified 78.2% (Mini-SWE-agent)Vals · Vibe Code Bench 67.42% (OpenHands)Vals · Code Migration 34.98%Vals · GPQA Diamond 91.67%Vals · MMLU Pro 87.48%Vals · AIME 96.67%
Enkrypt · Jailbreak risk 6.3%Enkrypt · Harmful content risk 0%Enkrypt · CBRN risk 16.7%Enkrypt · Toxicity risk 2.7%Enkrypt · Bias risk 87.1%Enkrypt · Insecure code risk 5.8%
Vectara · Factual consistency 93%
Badge
[](https://publicai.io/model-index/m/gpt-5-4)