§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Agents50.7−17.4—/267✱0/5
Tool use51.7−22.3—/81✱0/1
Core abilities47.9−19.5—/204✱0/3
Long context42.1−22.5—/0✱✱
Reasoning48.3−18.6—/178✱0/4
Science48−15.5—/122✱0/1
Expert reasoning41.4−19.5—/0✱✱
Coding49.1−20—/164✱0/5
Agentic coding48.6−20.1—/157✱0/4
Code generation48.3−21.7—/77✱0/2
Knowledge51.2−12.3—/138✱0/2
Factuality51.8−7.2—/0✱✱
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for G9v3 39A5B, left for the other.
§ 3 · Sources
Where the numbers come from
1 publication, 9 figures. Every one links to the page it was read from.
tau3-Banking 22.1%Terminal-Bench 2.1 32.6%SciCode 34%Humanity's Last Exam (without tools) 17.5%GPQA Diamond 80.5%CritPt 0.3%AA-LCR 62%AA-Omniscience Accuracy 14.9%AA-Omniscience Non-Hallucination 87%
Badge
[](https://publicai.io/model-index/m/g9v3-39a5b)