‹ PublicAI Index
The LLM benchmark aggregator.
GPT-5 Nano
OpenAI
Strongest in Harm refusal (#4 of 300), weakest in Human preference (#179 of 342). Above par in 14 of 24 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5 and Claude Fable 5.1 and ahead of Grok 3 and Kimi K2.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Safety59.6−1.9#11/3373/3
Harm refusal60.7−0.9#4/3002/2
Secure code65.9−0.6#8/2741/1
Jailbreak resistance62.9−2#11/2721/1
Safe-prompt compliance57.7−2.7#17/821/1
Fairness61.1−8.9#37/3002/2
Factual grounding48.7−21.9#61/1011/1
Toxicity avoidance54.6−4.4#121/2721/1
Agents55.5−12.6#59/2681/5
Tool use59.9−14.1#20/811/1
Core abilities53.2−14.2#60/2041/3
General intelligence54.6−15.5#63/2041/3
Reasoning47.1−19.8#108/1782/4
Mathematics54−11.7#61/1401/2
Reasoning43.3−25.8#93/1401/3
Science41.2−22.3#96/1221/1
Knowledge44.6−18.9#108/1381/2
Academic knowledge43.6−20.3#99/1231/1
Professional43.1−18.5#142/1681/1
Finance46.1−16.8#108/1511/1
Medical45.4−21.4#110/1401/1
Legal35.5−31.6#143/1511/1
Human preference51−16.5#179/3421/1
Human preference51−16.5#179/3421/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for GPT-5 Nano, left for the other.
§ 3 · Sources
Where the numbers come from
8 publications, 26 figures. Every one links to the page it was read from.
LMArena Text 1338
ARC-AGI-2 2.6%
BFCL v4 51.45%
Kagi LLM Benchmark 62.2%
Vals · CaseLaw 52.63%Vals · LegalBench 50.13%Vals · TaxEval 67.38%Vals · MortgageTax 53.62%Vals · MedQA 93.26%Vals · MedCode 30.44%Vals · MedScribe 72.86%Vals · GPQA Diamond 63.38%Vals · MMLU Pro 76.07%Vals · AIME 81.18%
HELM Safety · HarmBench 98.4%HELM Safety · SimpleSafetyTests 100%HELM Safety · Anthropic Red Team 99.6%HELM Safety · BBQ 97.6%HELM Safety · XSTest 97.6%
Enkrypt · Jailbreak risk 1.9%Enkrypt · Harmful content risk 7.2%Enkrypt · CBRN risk 2%Enkrypt · Toxicity risk 3.1%Enkrypt · Bias risk 61%Enkrypt · Insecure code risk 1.3%
Vectara · Factual consistency 89.5%
Badge
[](https://publicai.io/model-index/m/gpt-5-nano)