‹ PublicAI Index
The LLM benchmark aggregator.
GPT-4.1
OpenAI
Strongest in Legal (#12 of 151), weakest in Toxicity avoidance (#186 of 272). Above par in 18 of 26 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5 and Claude Fable 5.1 and ahead of GLM-5.1 and GLM-4.5.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Professional56.4−5.2#22/1681/1
Legal60.1−7#12/1511/1
Finance55.9−7#36/1511/1
Medical54.1−12.7#50/1401/1
Safety55.8−5.7#39/3373/3
Safe-prompt compliance58.4−2#12/821/1
Factual grounding61−9.6#20/1011/1
Harm refusal56.2−5.4#44/3002/2
Secure code61.1−5.4#56/2741/1
Jailbreak resistance56.6−8.3#91/2721/1
Fairness50.3−19.7#111/3002/2
Toxicity avoidance51.1−7.9#186/2721/1
Agents56.5−11.6#52/2681/5
Tool use61.7−12.3#16/811/1
Coding51.7−17.4#62/1651/5
Agentic coding52.1−16.6#59/1571/4
Knowledge49.2−14.3#89/1381/2
Academic knowledge49.1−14.8#82/1231/1
Core abilities49.9−17.5#96/2041/3
General intelligence49.8−20.3#97/2041/3
Human preference58.5−9#108/3421/1
Human preference58.5−9#108/3421/1
Reasoning40.9−26#149/1783/4
Science42.6−20.9#94/1221/1
Mathematics42.9−22.8#105/1401/2
Reasoning38.4−30.7#121/1402/3
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for GPT-4.1, left for the other.
§ 3 · Sources
Where the numbers come from
10 publications, 27 figures. Every one links to the page it was read from.
LMArena Text 1415
ARC-AGI-2 0.4%
Aider polyglot 52.4%
BFCL v4 53.96%
Kagi LLM Benchmark 52.3%
SimpleBench 27%
Vals · CaseLaw 69.88%Vals · LegalBench 83.1%Vals · CorpFin 63.05%Vals · TaxEval 75.06%Vals · MortgageTax 65.94%Vals · MedQA 91.18%Vals · GPQA Diamond 65.4%Vals · MMLU Pro 80.5%Vals · AIME 39.58%
HELM Safety · HarmBench 91.7%HELM Safety · SimpleSafetyTests 100%HELM Safety · Anthropic Red Team 99.3%HELM Safety · BBQ 92.6%HELM Safety · XSTest 97.9%
Enkrypt · Jailbreak risk 7.2%Enkrypt · Harmful content risk 25%Enkrypt · CBRN risk 10%Enkrypt · Toxicity risk 5.5%Enkrypt · Bias risk 82.2%Enkrypt · Insecure code risk 11.1%
Vectara · Factual consistency 94.4%
Badge
[](https://publicai.io/model-index/m/gpt-4-1)