‹ PublicAI Index
The LLM benchmark aggregator.
GPT-4 Turbo
OpenAI
Strongest in Safe-prompt compliance (#16 of 82), weakest in Human preference (#195 of 342). Above par in 9 of 16 scopes. Among the models it meets almost everywhere, it finishes behind Claude Sonnet 5 and Claude Opus 4.8 and ahead of Kimi K2.6 and GPT-4.1 Mini.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Safety57.1−4.4#24/3372/3
Safe-prompt compliance57.9−2.5#16/821/1
Harm refusal57−4.6#33/3002/2
Jailbreak resistance60.5−4.4#34/2721/1
Toxicity avoidance56.6−2.4#57/2721/1
Secure code61.1−5.4#58/2741/1
Fairness54.4−15.6#71/3002/2
Professional50.4−11.2#75/1681/1
Legal51.7−15.4#64/1511/1
Medical49.2−17.6#87/1401/1
Coding42.6−26.5#138/1651/5
Code generation35.8−34.2#69/771/2
Reasoning43−23.9#139/1781/4
Reasoning38.7−30.4#117/1401/3
Human preference49.7−17.8#195/3421/1
Human preference49.7−17.8#195/3421/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for GPT-4 Turbo, left for the other.
§ 3 · Sources
Where the numbers come from
6 publications, 16 figures. Every one links to the page it was read from.
LMArena Text 1324
LiveCodeBench 28.7%
SimpleBench 25.1%
Vals · LegalBench 80.46%Vals · MedQA 81.99%
HELM Safety · HarmBench 89.8%HELM Safety · SimpleSafetyTests 99%HELM Safety · Anthropic Red Team 99.7%HELM Safety · BBQ 94.1%HELM Safety · XSTest 97.7%
Enkrypt · Jailbreak risk 3.9%Enkrypt · Harmful content risk 23.9%Enkrypt · CBRN risk 5.3%Enkrypt · Toxicity risk 1.7%Enkrypt · Bias risk 73.6%Enkrypt · Insecure code risk 11.1%
Badge
[](https://publicai.io/model-index/m/gpt-4-turbo)