‹ PublicAI Index
The LLM benchmark aggregator.
O3
OpenAI
Strongest in Tool use (#7 of 81), weakest in Toxicity avoidance (#180 of 272). Above par in 25 of 27 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5 and Claude Fable 5.1 and ahead of GPT-5 Mini and GLM-5.1.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Coding59.9−9.2#11/1652/5
Code generation60.4−9.6#8/771/2
Agentic coding59.5−9.2#13/1571/4
Safety57.7−3.8#20/3372/3
Jailbreak resistance62.8−2.1#13/2721/1
Safe-prompt compliance57.1−3.3#21/821/1
Harm refusal58.3−3.3#22/3002/2
Secure code63−3.5#35/2741/1
Fairness55.1−14.9#68/3002/2
Toxicity avoidance51.9−7.1#180/2721/1
Professional54.2−7.4#40/1681/1
Medical55−11.8#36/1401/1
Legal54.2−12.9#43/1511/1
Finance54.2−8.7#57/1511/1
Core abilities55.1−12.3#45/2041/3
General intelligence57.3−12.8#43/2041/3
Knowledge54.5−9#60/1381/2
Academic knowledge55.5−8.4#55/1231/1
Reasoning52−14.9#75/1783/4
Mathematics55.1−10.6#49/1401/2
Science55.6−7.9#54/1221/1
Reasoning47.7−21.4#70/1402/3
Human preference60.1−7.4#87/3421/1
Human preference60.1−7.4#87/3421/1
Agents52.2−15.9#93/2682/5
Tool use68.3−5.7#7/811/1
Computer use37.5−34.2#33/371/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for O3, left for the other.
§ 3 · Sources
Where the numbers come from
11 publications, 29 figures. Every one links to the page it was read from.
LMArena Text 1432
ARC-AGI-2 6.5%
LiveCodeBench 75.8%
Aider polyglot 81.3%
BFCL v4 63.05%
OSWorld 23%
Kagi LLM Benchmark 67.6%
SimpleBench 53.1%
Vals · LegalBench 83.76%Vals · CorpFin 59.71%Vals · TaxEval 74.57%Vals · MortgageTax 65.7%Vals · MedQA 96.06%Vals · MedCode 47.29%Vals · MedScribe 76.65%Vals · GPQA Diamond 84.09%Vals · MMLU Pro 85.59%Vals · AIME 85.28%
HELM Safety · HarmBench 98.4%HELM Safety · SimpleSafetyTests 99%HELM Safety · Anthropic Red Team 98.3%HELM Safety · BBQ 97.9%HELM Safety · XSTest 97.3%
Enkrypt · Jailbreak risk 2%Enkrypt · Harmful content risk 6.1%Enkrypt · CBRN risk 9.7%Enkrypt · Toxicity risk 5%Enkrypt · Bias risk 77%Enkrypt · Insecure code risk 7.1%
Badge
[](https://publicai.io/model-index/m/o3)