§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Agents50.8−17.3—/267✱0/5
Tool use52−22—/81✱0/1
Core abilities49−18.4—/204✱0/3
Long context46.1−18.5—/0✱✱
Reasoning49.1−17.8—/178✱0/4
Science49−14.5—/122✱0/1
Expert reasoning45.4−15.5—/0✱✱
Coding48.9−20.2—/164✱0/5
Agentic coding49.1−19.6—/157✱0/4
Code generation46−24—/77✱0/2
Knowledge48.3−15.2—/138✱0/2
Factuality47.5−11.5—/0✱✱
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for K2 Horizon 32B, left for the other.
§ 3 · Sources
Where the numbers come from
1 publication, 9 figures. Every one links to the page it was read from.
tau3-Banking 22.5%Terminal-Bench 2.1 36.6%SciCode 30.2%Humanity's Last Exam (without tools) 22.8%GPQA Diamond 82.3%CritPt 1.4%AA-LCR 65.3%AA-Omniscience Accuracy 16.8%AA-Omniscience Non-Hallucination 58.3%
Badge
[](https://publicai.io/model-index/m/k2-horizon-32b)