‹ PublicAI Index
The LLM benchmark aggregator.
GPT OSS 120B
OpenAI · 117B · open weights
Strongest in Secure code (#7 of 274), weakest in Toxicity avoidance (#236 of 272). Above par in 10 of 26 scopes. Among the models it meets almost everywhere, it finishes behind Claude Opus 5.5 and Claude Fable 5 and ahead of Gemma 4 31B and Claude Haiku 4.5.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Safety54.2−7.3#69/3373/3
Secure code65.9−0.6#7/2741/1
Harm refusal57.7−3.9#25/3002/2
Fairness61.9−8.1#32/3002/2
Safe-prompt compliance46.8−13.6#63/821/1
Factual grounding39.3−31.3#87/1011/1
Jailbreak resistance42.8−22.1#211/2721/1
Toxicity avoidance43.4−15.6#236/2721/1
Reasoning48.6−18.3#96/1782/4
Mathematics57−8.7#28/1401/2
Science51.7−11.8#73/1221/1
Reasoning37.4−31.7#127/1401/3
Knowledge47.9−15.6#99/1381/2
Academic knowledge47.4−16.5#91/1231/1
Professional48.3−13.3#100/1681/1
Medical54.2−12.6#49/1401/1
Finance51−11.9#83/1511/1
Legal41.7−25.4#125/1511/1
Core abilities47.3−20.1#125/2042/3
General intelligence46.5−23.6#126/2042/3
Agents48.6−19.5#133/2681/5
Knowledge work47.9−25.7#91/1781/2
Coding41.9−27.2#142/1652/5
Agentic coding40.7−28#140/1572/4
Human preference52.4−15.1#163/3421/1
Human preference52.4−15.1#163/3421/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for GPT OSS 120B, left for the other.
§ 3 · Sources
Where the numbers come from
10 publications, 27 figures. Every one links to the page it was read from.
LMArena Text 1352
Artificial Analysis Intelligence Index 12
GDPval-AA 596
Aider polyglot 41.8%
Kagi LLM Benchmark 58.6%
SimpleBench 22.1%
Vals · CaseLaw 48.77%Vals · LegalBench 75.94%Vals · CorpFin 58.24%Vals · TaxEval 71.59%Vals · MedQA 91.36%Vals · SWE-bench Verified 33.6% (Mini-SWE-agent)Vals · GPQA Diamond 78.54%Vals · MMLU Pro 79.17%Vals · AIME 92.6%
HELM Safety · HarmBench 100%HELM Safety · SimpleSafetyTests 100%HELM Safety · Anthropic Red Team 99.5%HELM Safety · BBQ 98.5%HELM Safety · XSTest 92.7%
Enkrypt · Jailbreak risk 18.9%Enkrypt · Harmful content risk 14.4%Enkrypt · CBRN risk 13.8%Enkrypt · Toxicity risk 10.9%Enkrypt · Bias risk 60.2%Enkrypt · Insecure code risk 1.3%
Vectara · Factual consistency 85.8%
Badge
[](https://publicai.io/model-index/m/gpt-oss-120b)