‹ PublicAI Index
The LLM benchmark aggregator.
O1
OpenAI
Strongest in Jailbreak resistance (#4 of 272), weakest in Human preference (#121 of 342). Above par in 18 of 21 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5 and Claude Fable 5.1 and ahead of MiniMax M2.7 and GPT-6 Sol.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Safety58−3.5#18/3372/3
Jailbreak resistance64.1−0.8#4/2721/1
Harm refusal59.5−2.1#14/3002/2
Safe-prompt compliance56.4−4#26/821/1
Secure code61.1−5.4#57/2741/1
Fairness52.8−17.2#85/3002/2
Toxicity avoidance55.3−3.7#98/2721/1
Coding53.7−15.4#41/1651/5
Agentic coding54.5−14.2#42/1571/4
Professional53.6−8#48/1681/1
Medical56.9−9.9#24/1401/1
Finance54−8.9#61/1511/1
Legal51.7−15.4#65/1511/1
Knowledge52.4−11.1#75/1381/2
Academic knowledge52.8−11.1#69/1231/1
Reasoning48.4−18.5#98/1782/4
Mathematics51.4−14.3#75/1401/2
Reasoning45.5−23.6#78/1401/3
Science48−15.5#83/1221/1
Human preference57.2−10.3#121/3421/1
Human preference57.2−10.3#121/3421/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for O1, left for the other.
§ 3 · Sources
Where the numbers come from
6 publications, 20 figures. Every one links to the page it was read from.
LMArena Text 1402
Aider polyglot 61.7%
SimpleBench 40.1%
Vals · LegalBench 80.39%Vals · TaxEval 74.28%Vals · MedQA 96.52%Vals · GPQA Diamond 73.23%Vals · MMLU Pro 83.49%Vals · AIME 71.46%
HELM Safety · HarmBench 96.3%HELM Safety · SimpleSafetyTests 99%HELM Safety · Anthropic Red Team 98.3%HELM Safety · BBQ 97.3%HELM Safety · XSTest 97%
Enkrypt · Jailbreak risk 0.9%Enkrypt · Harmful content risk 3.9%Enkrypt · CBRN risk 3.7%Enkrypt · Toxicity risk 2.6%Enkrypt · Bias risk 82.2%Enkrypt · Insecure code risk 11.1%
Badge
[](https://publicai.io/model-index/m/o1)