‹ PublicAI Index
The LLM benchmark aggregator.
GPT-5.4 Nano
OpenAI
Strongest in Factual grounding (#2 of 101), weakest in Core abilities (#192 of 204). Above par in 9 of 24 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5 and Claude Fable 5.1 and ahead of Grok 4.3 and DeepSeek V3.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Safety55.4−6.1#43/3371/3
Factual grounding67.3−3.3#2/1011/1
Reasoning50.4−16.5#88/1783/4
Mathematics55.5−10.2#41/1402/2
Science51−12.5#74/1221/1
Reasoning43.9−25.2#84/1402/3
Knowledge45.8−17.7#106/1381/2
Academic knowledge44.9−19#98/1231/1
Agents50.1−18#117/2682/5
Knowledge work50.2−23.4#83/1782/2
Human preference57.1−10.4#123/3421/1
Human preference57.1−10.4#123/3421/1
Professional45.2−16.4#129/1681/1
Medical48.9−17.9#89/1401/1
Finance47.6−15.3#101/1511/1
Legal39.9−27.2#130/1511/1
Coding42.3−26.8#141/1652/5
Code generation36.5−33.5#66/771/2
Agentic coding44−24.7#114/1572/4
Core abilities39.3−28.1#192/2042/3
Instruction following48.8−24.9#31/571/1
Data analysis40.2−25.9#48/571/1
Language26−45.5#57/571/1
General intelligence40.6−29.5#169/2042/3
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for GPT-5.4 Nano, left for the other.
§ 3 · Sources
Where the numbers come from
8 publications, 30 figures. Every one links to the page it was read from.
LMArena Text 1401
GDPval-AA 937
AA-Briefcase 671
ARC-AGI-2 5.7%
LiveBench 69.6LiveBench · Reasoning 81.1LiveBench · Coding 70.8LiveBench · Agentic Coding 46.8LiveBench · Mathematics 91LiveBench · Data Analysis 67.6LiveBench · Language 62.5LiveBench · Instruction Following 67.2
Kagi LLM Benchmark 39.7%
Vals · Legal Research Bench 6.25%Vals · CaseLaw 51.88%Vals · LegalBench 77.92%Vals · Harvey Legal Agent Benchmark 0%Vals · Finance Agent 38.22%Vals · CorpFin 61.19%Vals · TaxEval 67.42%Vals · MortgageTax 59.1%Vals · MedCode 41.03%Vals · MedScribe 77.09%Vals · SWE-bench Verified 69.8% (Mini-SWE-agent)Vals · Vibe Code Bench 26.1% (OpenHands)Vals · Code Migration 14.47%Vals · GPQA Diamond 77.53%Vals · MMLU Pro 77.17%Vals · AIME 88.75%
Vectara · Factual consistency 96.9%
Badge
[](https://publicai.io/model-index/m/gpt-5-4-nano)