‹ PublicAI Index
The LLM benchmark aggregator.
Claude Fable 5.1
Anthropic
Strongest in Academic knowledge (#1 of 123), weakest in Instruction following (#15 of 57). Above par in 24 of 24 scopes. Among the models it meets almost everywhere, it finishes behind Claude Opus 5.5 and Claude Fable 5 and ahead of GPT-6 Astra and Claude Opus 5.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Knowledge61.6−1.9#2/1381/2
Academic knowledge63.9leads#1/1231/1
Coding66.5−2.6#2/1653/5
Code generation68.4−1.6#2/771/2
Agentic coding65.9−2.8#3/1573/4
Reasoning66.2−0.7#3/1784/4
Mathematics65.7leads#2/1402/2
Reasoning68.2−0.9#3/1403/3
Science62.1−1.4#9/1221/1
Core abilities65.8−1.6#3/2042/3
Language69.2−2.3#2/571/1
General intelligence69.7−0.4#3/2042/3
Data analysis61.6−4.5#5/571/1
Instruction following58.9−14.8#15/571/1
Agents66−2.1#4/2682/5
Knowledge work70.8−2.8#2/1782/2
Professional60.1−1.5#5/1681/1
Medical63.3−3.5#3/1401/1
IT operations55.1−18.7#4/111/1
Finance60.1−2.8#5/1511/1
Cybersecurity55.3−4.4#5/91/1
Legal61.1−6#8/1511/1
Human preference66.7−0.8#5/3421/1
Human preference66.7−0.8#5/3421/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Claude Fable 5.1, left for the other.
§ 3 · Sources
Where the numbers come from
9 publications, 30 figures. Every one links to the page it was read from.
LMArena Text 1501
Artificial Analysis Intelligence Index 53
GDPval-AA 1735
AA-Briefcase 1678
Terminal-Bench 57.9% (Claude Code)
ARC-AGI-2 90%
LiveBench 83.4LiveBench · Reasoning 91.7LiveBench · Coding 86.4LiveBench · Agentic Coding 66.1LiveBench · Mathematics 97LiveBench · Data Analysis 80.3LiveBench · Language 89.5LiveBench · Instruction Following 73
SimpleBench 86.6%
Vals · Legal Research Bench 55.29%Vals · LegalBench 88.51%Vals · Harvey Legal Agent Benchmark 6.67%Vals · Finance Agent 58.88%Vals · TaxEval 75.96%Vals · MortgageTax 70.79%Vals · MedCode 53.51%Vals · MedScribe 91.29%Vals · CyberBench 70.42%Vals · SRE Bench 22.9%Vals · Vibe Code Bench 90.26% (OpenHands)Vals · Code Migration 54.61%Vals · GPQA Diamond 93.43%Vals · MMLU Pro 92.38%Vals · ProofBench 100%
Badge
[](https://publicai.io/model-index/m/claude-fable-5-1)