‹ PublicAI Index
The LLM benchmark aggregator.
Granite 4.2 3B
IBM · 3B
Strongest in Knowledge work (#158 of 178), weakest in Agents (#252 of 268). Above par in 1 of 14 scopes. Among the models it meets almost everywhere, it finishes behind GPT-5.6 Terra and Claude Sonnet 5 and ahead of Gemma 3 27B It and Gemini 2.0 Flash.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Core abilities42.9−24.5#170/2041/3
General intelligence39.8−30.3#176/2041/3
Human preference46.6−20.9#226/3421/1
Human preference46.6−20.9#226/3421/1
Agents38.8−29.3#252/2682/5
Knowledge work35.4−38.2#158/1782/2
Tool use48.2−25.8—/81✱0/1
Reasoning48.9−18—/178✱0/4
Mathematics52.6−13.1—/140✱0/2
Science43.8−19.7—/122✱0/1
Expert reasoning44−16.9—/0✱✱
Coding47.3−21.8—/165✱0/5
Agentic coding46.7−22—/157✱0/4
Code generation42.6−27.4—/77✱0/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Granite 4.2 3B, left for the other.
§ 3 · Sources
Where the numbers come from
5 publications, 12 figures. Every one links to the page it was read from.
LMArena Text 1292
Artificial Analysis Intelligence Index 9
GDPval-AA 188
AA-Briefcase 99
tau3-Banking 5.6%Terminal-Bench 2.1 13.9%SciCode 24.9%GPQA Diamond 55.9%SWE-bench Verified 32.2%HMMT Feb 2026 57.2%HLE 6.6%BFCL v4 50.8%
Badge
[](https://publicai.io/model-index/m/granite-4-2-3b)