‹ PublicAI IndexGoogle
The LLM benchmark aggregator.
Gemma 4 E4B
Strongest in Secure code (#29 of 274), weakest in Fairness (#280 of 300). Above par in 2 of 13 scopes. Among the models it meets almost everywhere, it finishes behind Grok 4.7 and Gemini 3 Pro and ahead of MiniMax M3 and Gemma 3 27B It.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Safety48.3−13.2#228/3371/3
Secure code63.5−3#29/2741/1
Toxicity avoidance56.1−2.9#75/2721/1
Harm refusal46.8−14.8#212/3001/2
Jailbreak resistance38.4−26.5#226/2721/1
Fairness39.7−30.3#280/3001/2
Agents41.2−26.9#230/2681/5
Knowledge work36.9−36.7#149/1781/2
Tool use37.8−36.2—/81✱0/1
Reasoning47.5−19.4—/178✱0/4
Science42.2−21.3—/122✱0/1
Coding45.6−23.5—/165✱0/5
Agentic coding44.4−24.3—/157✱0/4
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Gemma 4 E4B, left for the other.
§ 3 · Sources
Where the numbers come from
3 publications, 10 figures. Every one links to the page it was read from.
GDPval-AA -51
Enkrypt · Jailbreak risk 22.6%Enkrypt · Harmful content risk 14.4%Enkrypt · CBRN risk 33.2%Enkrypt · Toxicity risk 2%Enkrypt · Bias risk 93.8%Enkrypt · Insecure code risk 6.2%
tau3-Banking 5%Terminal-Bench 2.1 2%GPQA Diamond 52%
Badge
[](https://publicai.io/model-index/m/gemma-4-e4b)