‹ PublicAI Index
The LLM benchmark aggregator.
Granite 3.2 8B Instruct
IBM · 8B
Strongest in Toxicity avoidance (#51 of 272), weakest in Safety (#259 of 336). Above par in 2 of 8 scopes. Among the models it meets almost everywhere, it finishes behind Claude Sonnet 5 and Claude Sonnet 4.5 and ahead of DeepSeek V4 Flash and Gemini 3.6 Flash.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Agents45.6−22.5#172/2671/5
Tool use42−32#60/811/1
Safety46.8−14.7#259/3361/3
Toxicity avoidance56.7−2.3#51/2721/1
Jailbreak resistance55.8−9.1#104/2721/1
Harm refusal48.7−12.9#176/2991/2
Fairness42.9−27.1#244/2991/2
Secure code26.5−40#257/2741/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Granite 3.2 8B Instruct, left for the other.
§ 3 · Sources
Where the numbers come from
2 publications, 7 figures. Every one links to the page it was read from.
BFCL v4 26.87%
Enkrypt · Jailbreak risk 7.9%Enkrypt · Harmful content risk 53.9%Enkrypt · CBRN risk 7.3%Enkrypt · Toxicity risk 1.6%Enkrypt · Bias risk 88.9%Enkrypt · Insecure code risk 81.3%
Badge
[](https://publicai.io/model-index/m/granite-3-2-8b-instruct)