Request a pilot
‹ PublicAI Index

The LLM benchmark aggregator.

Muse Spark 1.2

Meta

Strongest in Finance (#2 of 151), weakest in Mathematics (#79 of 140). Above par in 20 of 22 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5 and Claude Fable 5.1 and ahead of Qwen3.8 Max and Claude Opus 4.6.

§ 1 · Profile

What it is good at

Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.

Professional59.7#6/168
Finance62.6#2/151
Legal62.5#5/151
Medical60.1#9/140
Biology research42.1#11/16
Human preference66.2#6/342
Human preference66.2#6/342
Knowledge57.3#23/138
Academic knowledge58.8#20/123
Agents60.7#24/268
Knowledge work64#21/178
Coding55.8#29/165
Agentic coding57.4#23/157
Code generation50.2#45/77
Reasoning55.8#33/178
Reasoning62.1#14/140
Mathematics50.1#79/140
Core abilities55.1#44/204
Instruction following61.2#11/57
Data analysis55.2#25/57
Language48.7#33/57
General intelligence55.1#60/204

§ 2 · Head to head

What it beats, and what beats it

The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Muse Spark 1.2, left for the other.

GPT-4owon 14 · lost 0
Claude 3.7 Sonnetwon 14 · lost 0
Mimo V2.5won 14 · lost 0
Llama 4 Scout Instructwon 15 · lost 0
Command Awon 15 · lost 0
Gemini 2.0 Flashwon 15 · lost 0
Inkling Smallwon 15 · lost 0
Qwen3.6 27Bwon 16 · lost 0
GPT-4.1won 16 · lost 0
Kimi K2won 16 · lost 0
Mistral Medium 3.5won 16 · lost 0
Mimo V2.5 Prowon 16 · lost 0
GPT-4o Miniwon 17 · lost 0
GPT-4.1 Nanowon 17 · lost 0
GPT-4.1 Miniwon 17 · lost 0
Mistral Large 3won 17 · lost 0
DeepSeek V3won 18 · lost 0
GPT-5.4 Miniwon 20 · lost 0
MiniMax M3won 21 · lost 0
Inklingwon 21 · lost 0
Grok 4.3won 21 · lost 0
Gemini 3.5 Flash Litewon 21 · lost 0
GPT-5.4 Nanowon 20 · lost 1
Qwen3.6 Pluswon 20 · lost 1
Kimi K2.6won 20 · lost 1
GPT OSS 120Bwon 16 · lost 1
Gemini 2.5 Flashwon 16 · lost 1
Claude Haiku 4.5won 16 · lost 1
DeepSeek V3.2won 16 · lost 1
Claude Sonnet 4.5won 16 · lost 1
GLM-4.6won 15 · lost 1
Nemotron 3 Ultrawon 14 · lost 1
MiniMax M2.5won 14 · lost 1
GLM-4.7won 14 · lost 1
GLM-5.1won 14 · lost 1
GPT OSS 20Bwon 13 · lost 1
GPT-5 Nanowon 13 · lost 1
MiniMax M2.7won 13 · lost 1
DeepSeek V4 Flashwon 20 · lost 2
GLM-5.2won 19 · lost 2
GPT-6 Lunawon 18 · lost 2
O3 Miniwon 14 · lost 2
Gemini 3.1 Flash Litewon 14 · lost 2
Grok 3 Miniwon 13 · lost 2
Grok 4 Fastwon 13 · lost 2
GPT-5.6 Lunawon 19 · lost 3
Grok 4.1 Fastwon 12 · lost 2
Gemma 4 31Bwon 12 · lost 2
Muse Sparkwon 12 · lost 2
Qwen3.8 27Bwon 18 · lost 3
Claude Sonnet 4.6won 18 · lost 3
GLM-5.3 Flashwon 18 · lost 3
Qwen3.7 Maxwon 18 · lost 3
Muse Spark 1.1won 18 · lost 3
Claude Sonnet 4won 15 · lost 3
GPT-5 Miniwon 14 · lost 3
GPT-5won 14 · lost 3
Kimi K2.5won 14 · lost 3
Grok 4won 13 · lost 3
Claude Opus 4won 13 · lost 3
GPT-5.1won 12 · lost 3
GLM-5won 12 · lost 3
GPT-5.2won 16 · lost 4
Qwen3 Maxwon 11 · lost 3
Gemini 2.5 Prowon 14 · lost 4
Gemini 3.6 Flashwon 17 · lost 5
O4 Miniwon 13 · lost 4
Claude Sonnet 5won 16 · lost 5
DeepSeek V4 Prowon 16 · lost 5
Grok 4.5won 16 · lost 5
Gemini 3.5 Flashwon 16 · lost 5
Qwen3 235B A22B Instructwon 12 · lost 4
Claude Opus 4.5won 15 · lost 5
DeepSeek R1won 13 · lost 5
O3won 12 · lost 5
GPT-5.4won 14 · lost 7
GLM-5.3won 14 · lost 7
Gemini 3.8 Flashwon 13 · lost 9
GPT-5.6 Terrawon 12 · lost 9
DeepSeek V4.1 Flashwon 12 · lost 9
Claude Opus 4.8won 12 · lost 9
GPT-5.5won 12 · lost 9
Gemini 3.1 Prowon 12 · lost 9
Gemini 3.7 Flashwon 12 · lost 9
Gemini 3 Prowon 9 · lost 7
Claude Opus 4.6won 10 · lost 9
Qwen3.8 Maxwon 11 · lost 10
Kimi K3won 10 · lost 12
Grok 4.7won 9 · lost 11
Claude Opus 4.7won 9 · lost 12
GPT-6 Solwon 8 · lost 12
Grok 4.6won 8 · lost 14
GPT-5.6 Solwon 7 · lost 15
GPT-6 Astrawon 6 · lost 14
Claude Opus 5won 5 · lost 17
Claude Opus 5.5won 4 · lost 16
Muse Spark 1.3won 3 · lost 15
Claude Fable 5.1won 3 · lost 18
Claude Fable 5won 1 · lost 20

§ 3 · Sources

Where the numbers come from

6 publications, 27 figures. Every one links to the page it was read from.

LMArena Text ↗1 measureread 2026-09-26
LMArena Text 1496
GDPval-AA ↗1 measureread 2026-09-26
GDPval-AA 1482
AA-Briefcase ↗1 measureread 2026-09-26
AA-Briefcase 1333
LiveBench ↗8 measuresread 2026-09-26
LiveBench 78LiveBench · Reasoning 90LiveBench · Coding 77.5LiveBench · Agentic Coding 57.6LiveBench · Mathematics 91.2LiveBench · Data Analysis 76.5LiveBench · Language 78.6LiveBench · Instruction Following 74.3
SimpleBench ↗1 measureread 2026-09-26
SimpleBench 74.5%
Vals.ai ↗15 measuresread 2026-09-26
Vals · Legal Research Bench 43.75%Vals · LegalBench 85.26%Vals · Harvey Legal Agent Benchmark 25.42%Vals · Finance Agent 60.6%Vals · CorpFin 70.94%Vals · TaxEval 80.38%Vals · MortgageTax 65.42%Vals · MedCode 49.35%Vals · MedScribe 90.06%Vals · BioMysteryBench 64.81%Vals · SWE-bench Verified 86.6% (Mini-SWE-agent)Vals · Vibe Code Bench 79.1% (OpenHands)Vals · Code Migration 29.95%Vals · MMLU Pro 88.28%Vals · ProofBench 43%

Badge

PublicAI Index badge for Muse Spark 1.2[![PublicAI Index](https://publicai.io/model-index/badge?model=muse-spark-1-2)](https://publicai.io/model-index/m/muse-spark-1-2)