‹ PublicAI Index
The LLM benchmark aggregator.
Muse Spark 1.1
Meta
Strongest in Finance (#3 of 151), weakest in Toxicity avoidance (#156 of 272). Above par in 24 of 29 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5 and Muse Spark 1.3 and ahead of Qwen3.8 Max and GPT-5.4.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Professional60.2−1.4#4/1681/1
Finance61.7−1.2#3/1511/1
Legal61.3−5.8#6/1511/1
Medical57.9−8.9#21/1401/1
Safety59.8−1.7#9/3371/3
Jailbreak resistance63.5−1.4#6/2721/1
Harm refusal60−1.6#9/3001/2
Fairness67.8−2.2#17/3001/2
Secure code57.1−9.4#95/2741/1
Toxicity avoidance52.7−6.3#156/2721/1
Human preference65.8−1.7#9/3421/1
Human preference65.8−1.7#9/3421/1
Knowledge57.8−5.7#20/1381/2
Academic knowledge59.4−4.5#17/1231/1
Coding54.9−14.2#31/1652/5
Agentic coding56.4−12.3#31/1572/4
Code generation49.6−20.4#49/771/2
Agents58.6−9.5#42/2683/5
Computer use68.8−2.9#3/371/1
Knowledge work55.1−18.5#52/1782/2
Reasoning53.5−13.4#56/1782/4
Science60.5−3#23/1221/1
Reasoning55.3−13.8#40/1401/3
Mathematics46.9−18.8#90/1401/2
Core abilities48.3−19.1#115/2041/3
Instruction following53−20.7#27/571/1
Data analysis48.5−17.6#38/571/1
Language40.6−30.9#45/571/1
General intelligence50.8−19.3#89/2041/3
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Muse Spark 1.1, left for the other.
§ 3 · Sources
Where the numbers come from
7 publications, 31 figures. Every one links to the page it was read from.
LMArena Text 1491
GDPval-AA 1208
AA-Briefcase 848
LiveBench 75.3LiveBench · Reasoning 87.7LiveBench · Coding 77.2LiveBench · Agentic Coding 58.5LiveBench · Mathematics 87.1LiveBench · Data Analysis 72.5LiveBench · Language 74.3LiveBench · Instruction Following 69.6
OSWorld 80.7%
Vals · Legal Research Bench 37.98%Vals · LegalBench 84.98%Vals · Harvey Legal Agent Benchmark 20%Vals · Finance Agent 57.21%Vals · CorpFin 71.29%Vals · TaxEval 79.72%Vals · MortgageTax 66.14%Vals · MedScribe 88.89%Vals · SWE-bench Verified 82% (Mini-SWE-agent)Vals · Vibe Code Bench 72.16% (OpenHands)Vals · Code Migration 31.11%Vals · GPQA Diamond 91.16%Vals · MMLU Pro 88.73%
Enkrypt · Jailbreak risk 1.4%Enkrypt · Harmful content risk 0%Enkrypt · CBRN risk 6.3%Enkrypt · Toxicity risk 4.4%Enkrypt · Bias risk 50.1%Enkrypt · Insecure code risk 19.1%
Badge
[](https://publicai.io/model-index/m/muse-spark-1-1)