§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Agents51.1−17—/267✱0/5
Tool use52.8−21.2—/81✱0/1
Core abilities53.9−13.5—/204✱0/3
Long context64.3−0.3—/0✱✱
Reasoning49.2−17.7—/178✱0/4
Science50−13.5—/122✱0/1
Expert reasoning44.8−16.1—/0✱✱
Coding51.3−17.8—/164✱0/5
Agentic coding51.2−17.5—/157✱0/4
Code generation54.4−15.6—/77✱0/2
Knowledge46.7−16.8—/138✱0/2
Factuality45−14—/0✱✱
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Muse Glimmer-30B, left for the other.
§ 3 · Sources
Where the numbers come from
1 publication, 9 figures. Every one links to the page it was read from.
tau3-Banking 23.5%Terminal-Bench 2.1 51.7%SciCode 43.6%Humanity's Last Exam (without tools) 22%GPQA Diamond 83.5%CritPt 2.6%AA-LCR 80%AA-Omniscience Accuracy 27%AA-Omniscience Non-Hallucination 18.1%
Badge
[](https://publicai.io/model-index/m/muse-glimmer-30b)