‹ PublicAI Index
The LLM benchmark aggregator.
GPT-4o
OpenAI
Strongest in Multimodal understanding (#1 of 19), weakest in Human preference (#181 of 342). Above par in 14 of 24 scopes. Among the models it meets almost everywhere, it finishes behind Grok 4.7 and Claude Fable 5 and ahead of DeepSeek V3 and Grok 4.1 Fast.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Knowledge55.6−7.9#43/1382/2
Multimodal understanding72.1leads#1/191/1
Academic knowledge39.2−24.7#104/1231/1
Safety55.3−6.2#48/3373/3
Safe-prompt compliance57.1−3.3#22/821/1
Toxicity avoidance56.7−2.3#52/2721/1
Factual grounding50.9−19.7#53/1011/1
Secure code61.3−5.2#55/2741/1
Jailbreak resistance59−5.9#57/2721/1
Harm refusal54.6−7#78/3002/2
Fairness52.7−17.3#89/3002/2
Professional49.5−12.1#86/1681/1
Legal52.5−14.6#58/1511/1
Medical52.5−14.3#66/1401/1
Finance45.9−17#110/1511/1
Coding40.5−28.6#149/1652/5
Code generation36.2−33.8#68/771/2
Agentic coding43.2−25.5#119/1571/4
Reasoning35.2−31.7#173/1783/4
Science34.5−29#105/1221/1
Mathematics35.5−30.2#128/1401/2
Reasoning35.3−33.8#134/1402/3
Human preference50.8−16.7#181/3421/1
Human preference50.8−16.7#181/3421/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for GPT-4o, left for the other.
§ 3 · Sources
Where the numbers come from
10 publications, 27 figures. Every one links to the page it was read from.
LMArena Text 1336
ARC-AGI-2 0%
LiveCodeBench 29.5%
Aider polyglot 18.2%
MMMU-Pro 51.9%
SimpleBench 17.8%
Vals · CaseLaw 59.7%Vals · LegalBench 82.21%Vals · CorpFin 45.92%Vals · TaxEval 74.53%Vals · MortgageTax 57.43%Vals · MedQA 88.16%Vals · GPQA Diamond 53.79%Vals · MMLU Pro 72.56%Vals · AIME 11.88%
HELM Safety · HarmBench 82.9%HELM Safety · SimpleSafetyTests 98.5%HELM Safety · Anthropic Red Team 99.1%HELM Safety · BBQ 95.1%HELM Safety · XSTest 97.3%
Enkrypt · Jailbreak risk 5.2%Enkrypt · Harmful content risk 32.2%Enkrypt · CBRN risk 6.2%Enkrypt · Toxicity risk 1.6%Enkrypt · Bias risk 79.3%Enkrypt · Insecure code risk 10.7%
Vectara · Factual consistency 90.4%
Badge
[](https://publicai.io/model-index/m/gpt-4o)