‹ PublicAI Index
The LLM benchmark aggregator.
GPT-5.4 Mini
OpenAI
Strongest in Factual grounding (#17 of 101), weakest in Core abilities (#196 of 204). Above par in 10 of 23 scopes. Among the models it meets almost everywhere, it finishes behind Muse Spark 1.1 and Claude Opus 5 and ahead of GPT-5 Mini and Grok 4 Fast.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Human preference61.5−6#65/3421/1
Human preference61.5−6#65/3421/1
Knowledge53.5−10#66/1381/2
Academic knowledge54.2−9.7#61/1231/1
Safety53.5−8#83/3371/3
Factual grounding61.3−9.3#17/1011/1
Agents51.1−17#105/2682/5
Knowledge work51.4−22.2#71/1782/2
Professional46.6−15#115/1681/1
Finance51.8−11.1#77/1511/1
Legal38.9−28.2#134/1511/1
Coding43.3−25.8#130/1652/5
Code generation38.1−31.9#64/771/2
Agentic coding44.7−24#107/1572/4
Reasoning43.5−23.4#137/1783/4
Science54.9−8.6#59/1221/1
Mathematics45−20.7#98/1402/2
Reasoning35.5−33.6#132/1402/3
Core abilities37.9−29.5#196/2042/3
Data analysis45.6−20.5#41/571/1
Instruction following35.8−37.9#50/571/1
Language34.4−37.1#53/571/1
General intelligence37−33.1#190/2042/3
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for GPT-5.4 Mini, left for the other.
§ 3 · Sources
Where the numbers come from
8 publications, 27 figures. Every one links to the page it was read from.
LMArena Text 1447
GDPval-AA 1000
AA-Briefcase 718
ARC-AGI-2 18.9%
LiveBench 66.4LiveBench · Reasoning 71.3LiveBench · Coding 71.6LiveBench · Agentic Coding 41.7LiveBench · Mathematics 78.5LiveBench · Data Analysis 70.8LiveBench · Language 71LiveBench · Instruction Following 59.8
Kagi LLM Benchmark 37.9%
Vals · Legal Research Bench 12.5%Vals · CaseLaw 51.66%Vals · Harvey Legal Agent Benchmark 0%Vals · Finance Agent 45.36%Vals · CorpFin 60.92%Vals · TaxEval 71.22%Vals · MortgageTax 63.51%Vals · SWE-bench Verified 73% (Mini-SWE-agent)Vals · Vibe Code Bench 47.97% (OpenHands)Vals · Code Migration 12.94%Vals · GPQA Diamond 83.08%Vals · MMLU Pro 84.55%Vals · AIME 95.63%
Vectara · Factual consistency 94.5%
Badge
[](https://publicai.io/model-index/m/gpt-5-4-mini)