‹ PublicAI Index
The LLM benchmark aggregator.
GPT-3.5 Turbo
OpenAI
Strongest in Safe-prompt compliance (#60 of 82), weakest in Fairness (#300 of 300). Above par in 1 of 11 scopes. Among the models it meets almost everywhere, it finishes behind GPT-5.1 and GPT-5 and ahead of Mistral Instruct v0.3 (7B) and Zephyr 7B Beta.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Reasoning38.7−28.2#158/1781/4
Reasoning31.8−37.3#140/1401/3
Human preference38.1−29.4#274/3421/1
Human preference38.1−29.4#274/3421/1
Safety44.6−16.9#292/3372/3
Safe-prompt compliance49.3−11.1#60/821/1
Jailbreak resistance52.5−12.4#151/2721/1
Secure code47.3−19.2#184/2741/1
Harm refusal46−15.6#226/3002/2
Toxicity avoidance43.6−15.4#235/2721/1
Fairness33.3−36.7#300/3002/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for GPT-3.5 Turbo, left for the other.
§ 3 · Sources
Where the numbers come from
4 publications, 13 figures. Every one links to the page it was read from.
LMArena Text 1204
SimpleBench 8%
HELM Safety · HarmBench 68.6%HELM Safety · SimpleSafetyTests 92.3%HELM Safety · Anthropic Red Team 98%HELM Safety · BBQ 64.9%HELM Safety · XSTest 93.8%
Enkrypt · Jailbreak risk 10.7%Enkrypt · Harmful content risk 62.8%Enkrypt · CBRN risk 12.8%Enkrypt · Toxicity risk 10.8%Enkrypt · Bias risk 89.9%Enkrypt · Insecure code risk 39.1%
Badge
[](https://publicai.io/model-index/m/gpt-3-5-turbo)