‹ PublicAI Index
The LLM benchmark aggregator.
Granite 4.2 30B
IBM · 30B
Strongest in Knowledge work (#93 of 178, on 1 of its 2 boards), weakest in Human preference (#177 of 342). Above par in 2 of 15 scopes. Among the models it meets almost everywhere, it finishes behind GLM-5.2 and GPT-5.6 Terra and ahead of GPT-4.1 Mini and Granite 4.2 8B.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Agents48.3−19.8#137/2681/5
Knowledge work47.4−26.2#93/1781/2
Tool use45.4−28.6—/81✱0/1
Human preference51.3−16.2#177/3421/1
Human preference51.3−16.2#177/3421/1
Core abilities43.5−23.9—/204✱0/3
Long context26−38.6—/0✱✱
Reasoning46.9−20—/178✱0/4
Science44.7−18.8—/122✱0/1
Expert reasoning36.7−24.2—/0✱✱
Coding49.1−20—/165✱0/5
Agentic coding47.7−21—/157✱0/4
Code generation50−20—/77✱0/2
Knowledge48.2−15.3—/138✱0/2
Factuality47.3−11.7—/0✱✱
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Granite 4.2 30B, left for the other.
§ 3 · Sources
Where the numbers come from
3 publications, 11 figures. Every one links to the page it was read from.
LMArena Text 1341
GDPval-AA 564
tau3-Banking 14.4%Terminal-Bench 2.1 26.6%SciCode 36.6%Humanity's Last Exam (without tools) 11.2%GPQA Diamond 64.4%CritPt 0.3%AA-LCR 46.7%AA-Omniscience Accuracy 10.1%AA-Omniscience Non-Hallucination 74.4%
Badge
[](https://publicai.io/model-index/m/granite-4-2-30b)