‹ PublicAI Index
The LLM benchmark aggregator.
Gemma 3 4B It
Google · 4.3B · open weights
Strongest in Factual grounding (#27 of 101), weakest in Human preference (#220 of 342). Above par in 3 of 13 scopes. Among the models it meets almost everywhere, it finishes behind Claude Sonnet 5 and Inkling and ahead of Grok 4.3 and Grok 4.1 Fast.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Core abilities40.6−26.8#185/2041/3
General intelligence36.5−33.6#192/2041/3
Safety49.6−11.9#191/3372/3
Factual grounding59−11.6#27/1011/1
Jailbreak resistance52.2−12.7#157/2721/1
Fairness46−24#173/3001/2
Toxicity avoidance51.9−7.1#181/2721/1
Secure code44.9−21.6#195/2741/1
Harm refusal47.7−13.9#197/3001/2
Agents42.7−25.4#215/2681/5
Tool use36.8−37.2#74/811/1
Human preference47.7−19.8#220/3421/1
Human preference47.7−19.8#220/3421/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Gemma 3 4B It, left for the other.
§ 3 · Sources
Where the numbers come from
5 publications, 10 figures. Every one links to the page it was read from.
LMArena Text 1303
BFCL v4 19.62%
Kagi LLM Benchmark 25.2%
Enkrypt · Jailbreak risk 10.9%Enkrypt · Harmful content risk 51.1%Enkrypt · CBRN risk 11.5%Enkrypt · Toxicity risk 5%Enkrypt · Bias risk 84%Enkrypt · Insecure code risk 44%
Vectara · Factual consistency 93.6%
Badge
[](https://publicai.io/model-index/m/gemma-3-4b-it)