‹ PublicAI Index
The LLM benchmark aggregator.
Muse Spark 1.2
Meta
Strongest in Finance (#2 of 151), weakest in Mathematics (#79 of 140). Above par in 20 of 22 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5 and Claude Fable 5.1 and ahead of Qwen3.8 Max and Claude Opus 4.6.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Professional59.7−1.9#6/1681/1
Finance62.6−0.3#2/1511/1
Legal62.5−4.6#5/1511/1
Medical60.1−6.7#9/1401/1
Biology research42.1−26.4#11/161/1
Human preference66.2−1.3#6/3421/1
Human preference66.2−1.3#6/3421/1
Knowledge57.3−6.2#23/1381/2
Academic knowledge58.8−5.1#20/1231/1
Agents60.7−7.4#24/2682/5
Knowledge work64−9.6#21/1782/2
Coding55.8−13.3#29/1652/5
Agentic coding57.4−11.3#23/1572/4
Code generation50.2−19.8#45/771/2
Reasoning55.8−11.1#33/1783/4
Reasoning62.1−7#14/1402/3
Mathematics50.1−15.6#79/1402/2
Core abilities55.1−12.3#44/2041/3
Instruction following61.2−12.5#11/571/1
Data analysis55.2−10.9#25/571/1
Language48.7−22.8#33/571/1
General intelligence55.1−15#60/2041/3
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Muse Spark 1.2, left for the other.
§ 3 · Sources
Where the numbers come from
6 publications, 27 figures. Every one links to the page it was read from.
LMArena Text 1496
GDPval-AA 1482
AA-Briefcase 1333
LiveBench 78LiveBench · Reasoning 90LiveBench · Coding 77.5LiveBench · Agentic Coding 57.6LiveBench · Mathematics 91.2LiveBench · Data Analysis 76.5LiveBench · Language 78.6LiveBench · Instruction Following 74.3
SimpleBench 74.5%
Vals · Legal Research Bench 43.75%Vals · LegalBench 85.26%Vals · Harvey Legal Agent Benchmark 25.42%Vals · Finance Agent 60.6%Vals · CorpFin 70.94%Vals · TaxEval 80.38%Vals · MortgageTax 65.42%Vals · MedCode 49.35%Vals · MedScribe 90.06%Vals · BioMysteryBench 64.81%Vals · SWE-bench Verified 86.6% (Mini-SWE-agent)Vals · Vibe Code Bench 79.1% (OpenHands)Vals · Code Migration 29.95%Vals · MMLU Pro 88.28%Vals · ProofBench 43%
Badge
[](https://publicai.io/model-index/m/muse-spark-1-2)