‹ PublicAI Index
The LLM benchmark aggregator.
GPT-4.1 Mini
OpenAI
Strongest in Safe-prompt compliance (#19 of 82), weakest in Toxicity avoidance (#177 of 272). Above par in 12 of 26 scopes. Among the models it meets almost everywhere, it finishes behind Claude Sonnet 5 and GPT-5.2 and ahead of MiniMax M3 and GPT-5.4 Nano.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Professional51.2−10.4#69/1681/1
Finance52.4−10.5#73/1511/1
Medical50.6−16.2#76/1401/1
Legal49.9−17.2#77/1511/1
Safety53.3−8.2#92/3372/3
Safe-prompt compliance57.3−3.1#19/821/1
Secure code59.3−7.2#71/2741/1
Harm refusal53.4−8.2#99/3002/2
Fairness49.1−20.9#127/3002/2
Jailbreak resistance52.3−12.6#154/2721/1
Toxicity avoidance52.1−6.9#177/2721/1
Coding47.4−21.7#92/1651/5
Agentic coding46.9−21.8#88/1571/4
Knowledge45.8−17.7#105/1381/2
Academic knowledge45−18.9#97/1231/1
Core abilities48.6−18.8#113/2041/3
General intelligence48−22.1#117/2041/3
Agents49.9−18.2#122/2682/5
Tool use59.1−14.9#22/811/1
Knowledge work42.1−31.5#122/1781/2
Reasoning44.3−22.6#133/1782/4
Science44.3−19.2#93/1221/1
Mathematics45.5−20.2#95/1401/2
Reasoning42.7−26.4#100/1401/3
Human preference55.4−12.1#141/3421/1
Human preference55.4−12.1#141/3421/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for GPT-4.1 Mini, left for the other.
§ 3 · Sources
Where the numbers come from
9 publications, 25 figures. Every one links to the page it was read from.
LMArena Text 1383
GDPval-AA 256
ARC-AGI-2 0%
Aider polyglot 32.4%
BFCL v4 50.45%
Kagi LLM Benchmark 48.6%
Vals · LegalBench 78.04%Vals · CorpFin 57.93%Vals · TaxEval 71.91%Vals · MortgageTax 65.5%Vals · MedQA 84.63%Vals · GPQA Diamond 67.93%Vals · MMLU Pro 77.22%Vals · AIME 49.38%
HELM Safety · HarmBench 85.6%HELM Safety · SimpleSafetyTests 100%HELM Safety · Anthropic Red Team 99.3%HELM Safety · BBQ 92.1%HELM Safety · XSTest 97.4%
Enkrypt · Jailbreak risk 10.8%Enkrypt · Harmful content risk 41.7%Enkrypt · CBRN risk 11.3%Enkrypt · Toxicity risk 4.8%Enkrypt · Bias risk 84.5%Enkrypt · Insecure code risk 14.7%
Badge
[](https://publicai.io/model-index/m/gpt-4-1-mini)