‹ PublicAI Index
The LLM benchmark aggregator.
Muse Spark 1.3
Meta
Strongest in Legal (#1 of 151), weakest in Mathematics (#26 of 140). Above par in 19 of 19 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5.1 and Claude Fable 5 and ahead of GPT-6 Astra and GPT-5.5.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Professional60.2−1.4#3/1681/1
Legal67.1leads#1/1511/1
Cybersecurity57.3−2.4#4/91/1
Finance58.8−4.1#9/1511/1
Core abilities63.3−4.1#5/2042/3
Instruction following67.7−6#4/571/1
General intelligence65.8−4.3#6/2042/3
Data analysis60.4−5.7#8/571/1
Language56.6−14.9#19/571/1
Agents64.7−3.4#7/2682/5
Knowledge work69.1−4.5#6/1782/2
Human preference66.1−1.4#7/3421/1
Human preference66.1−1.4#7/3421/1
Coding61.1−8#8/1652/5
Agentic coding62.4−6.3#7/1572/4
Code generation57.6−12.4#18/771/2
Reasoning60−6.9#11/1783/4
Reasoning63.8−5.3#9/1402/3
Mathematics57.2−8.5#26/1402/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Muse Spark 1.3, left for the other.
§ 3 · Sources
Where the numbers come from
7 publications, 20 figures. Every one links to the page it was read from.
LMArena Text 1494
Artificial Analysis Intelligence Index 48
GDPval-AA 1674
AA-Briefcase 1587
LiveBench 81.6LiveBench · Reasoning 89.7LiveBench · Coding 81.1LiveBench · Agentic Coding 64.1LiveBench · Mathematics 95.9LiveBench · Data Analysis 79.6LiveBench · Language 82.8LiveBench · Instruction Following 78
SimpleBench 81.8%
Vals · Legal Research Bench 55.29%Vals · Harvey Legal Agent Benchmark 23.75%Vals · Finance Agent 59.96%Vals · CyberBench 72.74%Vals · Vibe Code Bench 85.86% (OpenHands)Vals · Code Migration 47.41%Vals · ProofBench 58%
Badge
[](https://publicai.io/model-index/m/muse-spark-1-3)