‹ PublicAI Index
The LLM benchmark aggregator.
Gemma 4 12B
Google · 12B
Strongest in Secure code (#52 of 274), weakest in Fairness (#257 of 300). Above par in 6 of 16 scopes. Among the models it meets almost everywhere, it finishes behind Claude Sonnet 5 and Claude Opus 5.5 and ahead of DeepSeek V4 Flash and Llama 4 Maverick Instruct.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Agents46.5−21.6#158/2681/5
Knowledge work44.8−28.8#108/1781/2
Safety49−12.5#206/3371/3
Secure code61.5−5#52/2741/1
Toxicity avoidance54.6−4.4#122/2721/1
Harm refusal48.4−13.2#182/3001/2
Jailbreak resistance40.4−24.5#218/2721/1
Fairness42.4−27.6#257/3001/2
Core abilities47.8−19.6—/204✱0/3
Long context41.7−22.9—/0✱✱
Reasoning52.6−14.3—/178✱0/4
Mathematics53.9−11.8—/140✱0/2
Expert reasoning56.8−4.1—/0✱✱
Coding48.8−20.3—/165✱0/5
Agentic coding47.4−21.3—/157✱0/4
Code generation51−19—/77✱0/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Gemma 4 12B, left for the other.
§ 3 · Sources
Where the numbers come from
3 publications, 13 figures. Every one links to the page it was read from.
GDPval-AA 413
Enkrypt · Jailbreak risk 20.9%Enkrypt · Harmful content risk 4.4%Enkrypt · CBRN risk 34.5%Enkrypt · Toxicity risk 3.1%Enkrypt · Bias risk 89.7%Enkrypt · Insecure code risk 10.2%
Terminal-Bench 2.1 27.3%SciCode 38.2%AA-LCR 61.7%SWE-bench Verified 30.6%HMMT Feb 2026 63.1%HLE 15.7%
Badge
[](https://publicai.io/model-index/m/gemma-4-12b)