‹ PublicAI Index
The LLM benchmark aggregator.
GPT-5 Mini
OpenAI
Strongest in Harm refusal (#11 of 300), weakest in Toxicity avoidance (#176 of 272). Above par in 22 of 27 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5 and Claude Fable 5.1 and ahead of GLM-5.2 and GPT-5.6 Luna.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Safety58.4−3.1#15/3373/3
Harm refusal59.8−1.8#11/3002/2
Secure code65−1.5#13/2741/1
Safe-prompt compliance57.9−2.5#15/821/1
Jailbreak resistance62.1−2.8#22/2721/1
Fairness61.2−8.8#35/3002/2
Factual grounding42.6−28#83/1011/1
Toxicity avoidance52.1−6.9#176/2721/1
Professional55.4−6.2#32/1681/1
Legal58.4−8.7#20/1511/1
Finance55.1−7.8#45/1511/1
Medical54.5−12.3#47/1401/1
Core abilities56−11.4#35/2041/3
General intelligence58.6−11.5#35/2041/3
Reasoning51.7−15.2#78/1782/4
Mathematics56.7−9#33/1401/2
Science52.9−10.6#67/1221/1
Reasoning43.6−25.5#88/1401/3
Knowledge51−12.5#80/1381/2
Academic knowledge51.3−12.6#74/1231/1
Agents51−17.1#108/2683/5
Tool use62.8−11.2#14/811/1
Knowledge work45.3−28.3#105/1782/2
Coding43.1−26#133/1651/5
Agentic coding42−26.7#128/1571/4
Human preference56−11.5#134/3421/1
Human preference56−11.5#134/3421/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for GPT-5 Mini, left for the other.
§ 3 · Sources
Where the numbers come from
10 publications, 31 figures. Every one links to the page it was read from.
LMArena Text 1390
GDPval-AA 754
AA-Briefcase 428
ARC-AGI-2 4.4%
BFCL v4 55.46%
Kagi LLM Benchmark 70.3%
Vals · CaseLaw 68.49%Vals · LegalBench 81.77%Vals · CorpFin 60.18%Vals · TaxEval 75.22%Vals · MortgageTax 66.89%Vals · MedQA 96.06%Vals · MedCode 43.05%Vals · MedScribe 80.58%Vals · SWE-bench Verified 60.8% (Mini-SWE-agent)Vals · Vibe Code Bench 14.17% (OpenHands)Vals · GPQA Diamond 80.3%Vals · MMLU Pro 82.23%Vals · AIME 91.46%
HELM Safety · HarmBench 97.1%HELM Safety · SimpleSafetyTests 100%HELM Safety · Anthropic Red Team 99.1%HELM Safety · BBQ 96.3%HELM Safety · XSTest 97.7%
Enkrypt · Jailbreak risk 2.6%Enkrypt · Harmful content risk 8.9%Enkrypt · CBRN risk 3.2%Enkrypt · Toxicity risk 4.8%Enkrypt · Bias risk 58.9%Enkrypt · Insecure code risk 3.1%
Vectara · Factual consistency 87.1%
Badge
[](https://publicai.io/model-index/m/gpt-5-mini)