‹ PublicAI Index
The LLM benchmark aggregator.
DeepSeek V4.1 Flash
DeepSeek · 763B · open weights
Strongest in Decisions (#1 of 85), weakest in Medical (#62 of 140). Above par in 28 of 31 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5 and Claude Fable 5.1 and ahead of Gemini 3.5 Flash and Claude Sonnet 5.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Decisions73.3leads#1/851/1
Calibration72.7leads#1/851/1
Routing & classification74leads#2/851/1
Coding62.2−6.9#6/1652/5
Agentic coding64.6−4.1#6/1572/4
Code generation55.3−14.7#26/771/2
Repository Q&A53.7−13—/0✱✱
Core abilities58.1−9.3#17/2042/3
Data analysis59.9−6.2#10/571/1
General intelligence61.6−8.5#14/2042/3
Language53.6−17.9#24/571/1
Instruction following53.7−20#26/571/1
Human preference64.4−3.1#25/3421/1
Human preference64.4−3.1#25/3421/1
Agents60.4−7.7#27/2682/5
Knowledge work63.5−10.1#24/1782/2
Workflow automation62leads—/0✱✱
Reasoning55.3−11.6#38/1783/4
Reasoning57.1−12#32/1402/3
Mathematics54−11.7#60/1402/2
Science46.4−17.1—/122✱0/1
Expert reasoning59.5−1.4—/0✱✱
Professional52.7−8.9#53/1681/1
Cybersecurity58.2−1.5#2/91/1
IT operations40.1−33.7#9/111/1
Biology research47.5−21#10/161/1
Legal55.9−11.2#34/1511/1
Finance54.6−8.3#51/1511/1
Medical52.9−13.9#62/1401/1
Knowledge50.6−12.9—/138✱0/2
Multimodal understanding50.8−21.3—/19✱0/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for DeepSeek V4.1 Flash, left for the other.
§ 3 · Sources
Where the numbers come from
9 publications, 45 figures. Every one links to the page it was read from.
LMArena Text 1477
Artificial Analysis Intelligence Index 39
GDPval-AA 1328
AA-Briefcase 1425
LiveBench 81.1LiveBench · Reasoning 86.7LiveBench · Coding 80LiveBench · Agentic Coding 77.3LiveBench · Mathematics 93.3LiveBench · Data Analysis 79.3LiveBench · Language 81.2LiveBench · Instruction Following 70
SimpleBench 66.7%
Vals · Legal Research Bench 41.35%Vals · LegalBench 83.28%Vals · Harvey Legal Agent Benchmark 6.67%Vals · Finance Agent 53.48%Vals · MedCode 41.17%Vals · MedScribe 85.5%Vals · BioMysteryBench 67.78%Vals · CyberBench 73.69%Vals · SRE Bench 0.76%Vals · Vibe Code Bench 84.74% (OpenHands)Vals · Code Migration 45.62%Vals · ProofBench 54%
JevBench · Intelligence 94%JevBench · Calibration 95.5%
GPQA Diamond (Pass@1) 90.9%Codeforces (Rating) 3471MathArena Apex (Pass@1) 65.6%Terminal-Bench 2.1 (Pass@1) 90.6%Terminal-Bench 3.0 (Pass@1) 30%Terminal-Bench 4.0 (Pass@1) 31.2%DeepSWE v1.1 (Resolved) 74.2%ProgramBench (Almost@1) 20.3%NL2Repo-Bench (Score) 64%CyberGym (Pass@1) 88.1%SEC-Bench Pro (Pass@1) 62.8%ExploitGym (Pass@1) 15.3%HLE w/ tools (Pass@1) 63.9%AutomationBench (Pass@1) 54.8%Agent's Last Exam (Pass@1) 31.8%Chartography w/ tools (Pass@1) 78.9%BabyVision w/ tools (Pass@1) 89.6%ZeroBench-main w/ tools (Pass@5) 49%
Badge
[](https://publicai.io/model-index/m/deepseek-v4-1-flash)