‹ PublicAI Index
The LLM benchmark aggregator.
GPT-4.1 Nano
OpenAI
Strongest in Safe-prompt compliance (#39 of 82), weakest in Agents (#231 of 268). Above par in 6 of 26 scopes. Among the models it meets almost everywhere, it finishes behind Claude Sonnet 5 and GPT-5 and ahead of Llama 4 Scout Instruct and Mistral Large.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Safety52.9−8.6#106/3372/3
Safe-prompt compliance54.2−6.2#39/821/1
Secure code60.4−6.1#65/2741/1
Harm refusal53.8−7.8#92/3002/2
Jailbreak resistance53.6−11.3#136/2721/1
Toxicity avoidance53.4−5.6#145/2721/1
Fairness45.7−24.3#180/3002/2
Knowledge31.5−32#134/1381/2
Academic knowledge27.8−36.1#118/1231/1
Coding42.4−26.7#140/1651/5
Agentic coding40.8−27.9#139/1571/4
Professional39.3−22.3#157/1681/1
Medical41.8−25#120/1401/1
Finance37.4−25.5#138/1511/1
Legal36.9−30.2#141/1511/1
Reasoning38.3−28.6#160/1782/4
Reasoning42.7−26.4#103/1401/3
Science32.4−31.1#108/1221/1
Mathematics39.4−26.3#114/1401/2
Core abilities43.4−24#166/2041/3
General intelligence40.5−29.6#172/2041/3
Human preference49.5−18#198/3421/1
Human preference49.5−18#198/3421/1
Agents41.2−26.9#231/2682/5
Tool use46.5−27.5#41/811/1
Knowledge work33.8−39.8#165/1781/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for GPT-4.1 Nano, left for the other.
§ 3 · Sources
Where the numbers come from
9 publications, 25 figures. Every one links to the page it was read from.
LMArena Text 1322
GDPval-AA -228
ARC-AGI-2 0%
Aider polyglot 8.9%
BFCL v4 33.05%
Kagi LLM Benchmark 33.3%
Vals · LegalBench 61.06%Vals · CorpFin 42.08%Vals · TaxEval 60.75%Vals · MortgageTax 52.82%Vals · MedQA 68.22%Vals · GPQA Diamond 50.76%Vals · MMLU Pro 63.48%Vals · AIME 26.46%
HELM Safety · HarmBench 86.8%HELM Safety · SimpleSafetyTests 99%HELM Safety · Anthropic Red Team 99.6%HELM Safety · BBQ 87.5%HELM Safety · XSTest 96%
Enkrypt · Jailbreak risk 9.7%Enkrypt · Harmful content risk 40%Enkrypt · CBRN risk 10.5%Enkrypt · Toxicity risk 3.9%Enkrypt · Bias risk 87.1%Enkrypt · Insecure code risk 12.4%
Badge
[](https://publicai.io/model-index/m/gpt-4-1-nano)