‹ PublicAI Index
The LLM benchmark aggregator.
O4 Mini
OpenAI
Strongest in Code generation (#5 of 77, on 1 of its 2 boards), weakest in Toxicity avoidance (#205 of 272). Above par in 19 of 27 scopes. Among the models it meets almost everywhere, it finishes behind Claude Opus 5.5 and Claude Fable 5 and ahead of DeepSeek V3.2 and DeepSeek R1.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Coding59.3−9.8#12/1652/5
Code generation62.7−7.3#5/771/2
Agentic coding57.1−11.6#26/1571/4
Core abilities55.1−12.3#46/2041/3
General intelligence57.3−12.8#44/2041/3
Agents56.2−11.9#53/2681/5
Tool use61.2−12.8#17/811/1
Professional50.3−11.3#77/1681/1
Finance53.7−9.2#63/1511/1
Legal50.7−16.4#74/1511/1
Medical46.3−20.5#102/1401/1
Safety53.5−8#85/3373/3
Safe-prompt compliance57.3−3.1#18/821/1
Harm refusal56.8−4.8#37/3002/2
Jailbreak resistance58.7−6.2#61/2721/1
Secure code60.4−6.1#64/2741/1
Factual grounding28.2−42.4#93/1011/1
Fairness51.9−18.1#96/3002/2
Toxicity avoidance49.4−9.6#205/2721/1
Knowledge49.3−14.2#88/1381/2
Academic knowledge49.2−14.7#81/1231/1
Reasoning48.2−18.7#101/1783/4
Mathematics54.6−11.1#56/1401/2
Science48.9−14.6#81/1221/1
Reasoning42.9−26.2#97/1402/3
Human preference56.1−11.4#132/3421/1
Human preference56.1−11.4#132/3421/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for O4 Mini, left for the other.
§ 3 · Sources
Where the numbers come from
11 publications, 29 figures. Every one links to the page it was read from.
LMArena Text 1391
ARC-AGI-2 6.1%
LiveCodeBench 80.2%
Aider polyglot 72%
BFCL v4 53.24%
Kagi LLM Benchmark 67.6%
SimpleBench 38.7%
Vals · LegalBench 79.19%Vals · CorpFin 58.97%Vals · TaxEval 74.78%Vals · MortgageTax 64.83%Vals · MedQA 96.02%Vals · MedCode 33.79%Vals · MedScribe 69.14%Vals · GPQA Diamond 74.5%Vals · MMLU Pro 80.56%Vals · AIME 83.67%
HELM Safety · HarmBench 97%HELM Safety · SimpleSafetyTests 100%HELM Safety · Anthropic Red Team 98.2%HELM Safety · BBQ 94%HELM Safety · XSTest 97.4%
Enkrypt · Jailbreak risk 5.4%Enkrypt · Harmful content risk 15%Enkrypt · CBRN risk 12.7%Enkrypt · Toxicity risk 6.7%Enkrypt · Bias risk 79.8%Enkrypt · Insecure code risk 12.4%
Vectara · Factual consistency 81.4%
Badge
[](https://publicai.io/model-index/m/o4-mini)