‹ PublicAI Index
The LLM benchmark aggregator.
Muse Spark
Meta
Strongest in Human preference (#10 of 342), weakest in Agents (#113 of 268). Above par in 13 of 15 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5 and Claude Opus 4.7 and ahead of Gemini 3.6 Flash and GLM-5.3 Flash.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Human preference65.6−1.9#10/3421/1
Human preference65.6−1.9#10/3421/1
Professional57.1−4.5#17/1681/1
Medical59.4−7.4#12/1401/1
Finance57.9−5#20/1511/1
Legal56−11.1#31/1511/1
Reasoning57.1−9.8#25/1781/4
Mathematics58.2−7.5#16/1401/2
Science59.5−4#29/1221/1
Knowledge56.3−7.2#33/1381/2
Academic knowledge57.6−6.3#30/1231/1
Coding47.3−21.8#93/1651/5
Agentic coding46.9−21.8#87/1571/4
Agents50.3−17.8#113/2682/5
Knowledge work50.3−23.3#79/1782/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Muse Spark, left for the other.
§ 3 · Sources
Where the numbers come from
4 publications, 14 figures. Every one links to the page it was read from.
LMArena Text 1489
GDPval-AA 986
AA-Briefcase 644
Vals · CaseLaw 63.13%Vals · LegalBench 84.22%Vals · CorpFin 65.11%Vals · TaxEval 77.68%Vals · MedCode 51.31%Vals · MedScribe 85.9%Vals · SWE-bench Verified 74.4% (Mini-SWE-agent)Vals · Vibe Code Bench 19.67% (OpenHands)Vals · GPQA Diamond 89.65%Vals · MMLU Pro 87.32%Vals · AIME 96.88%
Badge
[](https://publicai.io/model-index/m/muse-spark)