‹ PublicAI Index
The LLM benchmark aggregator.
GPT-4.5
OpenAI
Strongest in Harm refusal (#30 of 300, on 1 of its 2 boards), weakest in Reasoning (#136 of 178). Above par in 8 of 10 scopes. Among the models it meets almost everywhere, it finishes behind Claude Opus 5.5 and Muse Spark 1.1 and ahead of Qwen3.7 Max and GLM-5.1.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Safety54.7−6.8#58/3371/3
Harm refusal57.3−4.3#30/3001/2
Safe-prompt compliance52.2−8.2#52/821/1
Fairness52.7−17.3#86/3001/2
Human preference61.3−6.2#68/3421/1
Human preference61.3−6.2#68/3421/1
Coding50.1−19#75/1651/5
Agentic coding50.1−18.6#73/1571/4
Reasoning43.7−23.2#136/1782/4
Reasoning40.9−28.2#110/1402/3
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for GPT-4.5, left for the other.
§ 3 · Sources
Where the numbers come from
5 publications, 9 figures. Every one links to the page it was read from.
LMArena Text 1445
ARC-AGI-2 0.8%
Aider polyglot 44.9%
SimpleBench 34.5%
HELM Safety · HarmBench 95.8%HELM Safety · SimpleSafetyTests 100%HELM Safety · Anthropic Red Team 99.4%HELM Safety · BBQ 92%HELM Safety · XSTest 95.1%
Badge
[](https://publicai.io/model-index/m/gpt-4-5)