‹ PublicAI Index
The LLM benchmark aggregator.
GPT-4o Mini
OpenAI
Strongest in Multimodal understanding (#6 of 19), weakest in Agents (#227 of 268). Above par in 7 of 26 scopes. Among the models it meets almost everywhere, it finishes behind Claude Sonnet 4 and Claude Opus 4.7 and ahead of Mistral Large and Llama 4 Scout Instruct.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Safety53−8.5#103/3372/3
Safe-prompt compliance54.2−6.2#40/821/1
Jailbreak resistance58.1−6.8#67/2721/1
Toxicity avoidance56.1−2.9#74/2721/1
Secure code57.3−9.2#92/2741/1
Harm refusal53.1−8.5#108/3002/2
Fairness46.4−23.6#161/3002/2
Knowledge42.1−21.4#115/1382/2
Multimodal understanding57.3−14.8#6/191/1
Academic knowledge26.9−37#119/1231/1
Professional42.9−18.7#144/1681/1
Medical44.1−22.7#115/1401/1
Finance39.3−23.6#132/1511/1
Coding37.8−31.3#161/1652/5
Code generation35.1−34.9#70/771/2
Agentic coding39.4−29.3#146/1571/4
Core abilities41.9−25.5#177/2041/3
General intelligence38.3−31.8#184/2041/3
Reasoning32.6−34.3#178/1783/4
Science27.8−35.7#114/1221/1
Mathematics35.4−30.3#129/1401/2
Reasoning33−36.1#137/1402/3
Human preference49.1−18.4#206/3421/1
Human preference49.1−18.4#206/3421/1
Agents41.4−26.7#227/2681/5
Knowledge work37.1−36.5#147/1781/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for GPT-4o Mini, left for the other.
§ 3 · Sources
Where the numbers come from
11 publications, 26 figures. Every one links to the page it was read from.
LMArena Text 1318
GDPval-AA -37
ARC-AGI-2 0%
LiveCodeBench 27.5%
Aider polyglot 3.6%
MMMU-Pro 37.6%
Kagi LLM Benchmark 28.8%
SimpleBench 10.7%
Vals · CorpFin 45.45%Vals · TaxEval 60.55%Vals · MortgageTax 54.49%Vals · MedQA 72.44%Vals · GPQA Diamond 44.19%Vals · MMLU Pro 62.73%Vals · AIME 11.46%
HELM Safety · HarmBench 84.9%HELM Safety · SimpleSafetyTests 97.8%HELM Safety · Anthropic Red Team 98.3%HELM Safety · BBQ 88.2%HELM Safety · XSTest 96%
Enkrypt · Jailbreak risk 5.9%Enkrypt · Harmful content risk 39.4%Enkrypt · CBRN risk 8%Enkrypt · Toxicity risk 2%Enkrypt · Bias risk 86.3%Enkrypt · Insecure code risk 18.7%
Badge
[](https://publicai.io/model-index/m/gpt-4o-mini)