‹ PublicAI Index
The LLM benchmark aggregator.
Gemma 4 31B
Google · 31B · open weights
Strongest in Factual grounding (#34 of 101), weakest in Fairness (#270 of 300). Above par in 12 of 28 scopes. Among the models it meets almost everywhere, it finishes behind Claude Opus 5 and Claude Opus 5.5 and ahead of MiniMax M3 and GPT-4.1 Mini.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Core abilities53.7−13.7#56/2041/3
General intelligence55.3−14.8#58/2041/3
Long context49.8−14.8—/0✱✱
Human preference62.1−5.4#60/3421/1
Human preference62.1−5.4#60/3421/1
Professional47.8−13.8#106/1681/1
Finance50.6−12.3#85/1511/1
Legal43.4−23.7#115/1511/1
Safety49.7−11.8#188/3372/3
Factual grounding56.5−14.1#34/1011/1
Secure code60.6−5.9#62/2741/1
Toxicity avoidance52.7−6.3#160/2721/1
Harm refusal48.5−13.1#179/3001/2
Jailbreak resistance42.2−22.7#212/2721/1
Fairness41.2−28.8#270/3001/2
Agents44.6−23.5#188/2682/5
Knowledge work43−30.6#116/1782/2
Tool use45.8−28.2—/81✱0/1
Reasoning49.4−17.5—/178✱0/4
Science49.7−13.8—/122✱0/1
Expert reasoning46−14.9—/0✱✱
Coding50.8−18.3—/165✱0/5
Agentic coding50−18.7—/157✱0/4
Code generation54.2−15.8—/77✱0/2
Knowledge44.1−19.4—/138✱0/2
Factuality41.2−17.8—/0✱✱
Decisions60.2−13.1—/85✱0/1
Routing & classification63.7−10.3—/85✱0/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Gemma 4 31B, left for the other.
§ 3 · Sources
Where the numbers come from
9 publications, 23 figures. Every one links to the page it was read from.
LMArena Text 1453
GDPval-AA 606
AA-Briefcase 367
Kagi LLM Benchmark 63.5%
Vals · CaseLaw 52.63%Vals · MortgageTax 61.37%
Enkrypt · Jailbreak risk 19.4%Enkrypt · Harmful content risk 6.1%Enkrypt · CBRN risk 33.3%Enkrypt · Toxicity risk 4.4%Enkrypt · Bias risk 91.5%Enkrypt · Insecure code risk 12%
Vectara · Factual consistency 92.6%
tau3-Banking 14.8%Terminal-Bench 2.1 43.4%SciCode 43.4%Humanity's Last Exam (without tools) 23.6%GPQA Diamond 85.7%CritPt 1.4%AA-LCR 68.3%AA-Omniscience Accuracy 20%AA-Omniscience Non-Hallucination 15%
Decision correct 77%
Badge
[](https://publicai.io/model-index/m/gemma-4-31b)