‹ PublicAI Index
The LLM benchmark aggregator.
Granite 3.0 2B Instruct
IBM · 2B
Strongest in Toxicity avoidance (#22 of 272), weakest in Human preference (#299 of 342). Above par in 3 of 8 scopes. Among the models it meets almost everywhere, it finishes behind Claude Sonnet 5 and Claude Sonnet 4 and ahead of Command R and Mimo V2.6 Flash.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Safety49.9−11.6#182/3371/3
Toxicity avoidance58.3−0.7#22/2721/1
Jailbreak resistance54.7−10.2#121/2721/1
Secure code51.7−14.8#142/2741/1
Harm refusal47.9−13.7#190/3001/2
Fairness41.9−28.1#265/3001/2
Human preference33.6−33.9#299/3421/1
Human preference33.6−33.9#299/3421/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Granite 3.0 2B Instruct, left for the other.
§ 3 · Sources
Where the numbers come from
2 publications, 7 figures. Every one links to the page it was read from.
LMArena Text 1157
Enkrypt · Jailbreak risk 8.8%Enkrypt · Harmful content risk 38.3%Enkrypt · CBRN risk 17.8%Enkrypt · Toxicity risk 0.5%Enkrypt · Bias risk 90.4%Enkrypt · Insecure code risk 30.2%
Badge
[](https://publicai.io/model-index/m/granite-3-0-2b-instruct)