§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Agents44.7−23.4—/267✱0/5
Tool use36.5−37.5—/81✱0/1
Reasoning48.6−18.3—/178✱0/4
Mathematics50.2−15.5—/140✱0/2
Science31.6−31.9—/122✱0/1
Coding50.9−18.2—/164✱0/5
Code generation52.2−17.8—/77✱0/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for OpenBMB 1B, left for the other.
§ 3 · Sources
Where the numbers come from
1 publication, 8 figures. Every one links to the page it was read from.
GPQA Diamond 26.3%HMMT Feb 2026 23.3%BFCL v4 25.2%HumanEval+ 65.2%MBPP+ 60.6%AIME 2025 40.4%AIME 2026 40.4%LiveCodeBench v6 33.5%
Badge
[](https://publicai.io/model-index/m/openbmb-1b)