‹ PublicAI Index
The LLM benchmark aggregator.
Gemma 3 27B It
Google · 27B · open weights
Strongest in Factual grounding (#36 of 101), weakest in Agents (#255 of 268). Above par in 6 of 16 scopes. Among the models it meets almost everywhere, it finishes behind Claude Sonnet 5 and Inkling and ahead of Mistral Large and Nova Micro.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Safety51.6−9.9#142/3372/3
Factual grounding56.5−14.1#36/1011/1
Jailbreak resistance56.8−8.1#85/2721/1
Toxicity avoidance55.3−3.7#99/2721/1
Fairness46.7−23.3#157/3001/2
Harm refusal49.8−11.8#160/3001/2
Secure code48.8−17.7#169/2741/1
Coding41.5−27.6#146/1651/5
Agentic coding39.8−28.9#144/1571/4
Human preference53.6−13.9#151/3421/1
Human preference53.6−13.9#151/3421/1
Core abilities44−23.4#160/2041/3
General intelligence41.4−28.7#163/2041/3
Agents38.7−29.4#255/2682/5
Tool use43.9−30.1#49/811/1
Knowledge work30.5−43.1#175/1781/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Gemma 3 27B It, left for the other.
§ 3 · Sources
Where the numbers come from
7 publications, 12 figures. Every one links to the page it was read from.
LMArena Text 1365
GDPval-AA -424
Aider polyglot 4.9%
BFCL v4 29.47%
Kagi LLM Benchmark 35.1%
Enkrypt · Jailbreak risk 7%Enkrypt · Harmful content risk 43.9%Enkrypt · CBRN risk 9.7%Enkrypt · Toxicity risk 2.6%Enkrypt · Bias risk 83%Enkrypt · Insecure code risk 36%
Vectara · Factual consistency 92.6%
Badge
[](https://publicai.io/model-index/m/gemma-3-27b-it)