‹ PublicAI Index
The LLM benchmark aggregator.
Gemma 3 12B It
Google · 12B · open weights
Strongest in Factual grounding (#7 of 101), weakest in Agents (#251 of 268). Above par in 7 of 12 scopes. Among the models it meets almost everywhere, it finishes behind Claude Sonnet 5 and Inkling and ahead of Gemini 2.5 Pro and Gemma 4 26B A4B.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Safety52.3−9.2#124/3372/3
Factual grounding64.1−6.5#7/1011/1
Toxicity avoidance56.1−2.9#73/2721/1
Jailbreak resistance55.7−9.2#107/2721/1
Secure code51−15.5#150/2741/1
Harm refusal49.3−12.3#164/3001/2
Fairness44.2−25.8#209/3001/2
Human preference51.4−16.1#176/3421/1
Human preference51.4−16.1#176/3421/1
Agents38.9−29.2#251/2682/5
Tool use44.6−29.4#47/811/1
Knowledge work30.5−43.1#176/1781/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Gemma 3 12B It, left for the other.
§ 3 · Sources
Where the numbers come from
5 publications, 10 figures. Every one links to the page it was read from.
LMArena Text 1342
GDPval-AA -424
BFCL v4 30.43%
Enkrypt · Jailbreak risk 8%Enkrypt · Harmful content risk 47.2%Enkrypt · CBRN risk 9.3%Enkrypt · Toxicity risk 2%Enkrypt · Bias risk 86.8%Enkrypt · Insecure code risk 31.6%
Vectara · Factual consistency 95.6%
Badge
[](https://publicai.io/model-index/m/gemma-3-12b-it)