Request a pilot
‹ PublicAI Index

The LLM benchmark aggregator.

DeepSeek V4.1 Flash

DeepSeek · 763B · open weights

Strongest in Decisions (#1 of 85), weakest in Medical (#62 of 140). Above par in 28 of 31 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5 and Claude Fable 5.1 and ahead of Gemini 3.5 Flash and Claude Sonnet 5.

§ 1 · Profile

What it is good at

Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.

Decisions73.3#1/85
Calibration72.7#1/85
Routing & classification74#2/85
Coding62.2#6/165
Agentic coding64.6#6/157
Code generation55.3#26/77
Repository Q&A53.7—/0✱
Core abilities58.1#17/204
Data analysis59.9#10/57
General intelligence61.6#14/204
Language53.6#24/57
Instruction following53.7#26/57
Human preference64.4#25/342
Human preference64.4#25/342
Agents60.4#27/268
Knowledge work63.5#24/178
Workflow automation62—/0✱
Reasoning55.3#38/178
Reasoning57.1#32/140
Mathematics54#60/140
Science46.4—/122✱
Expert reasoning59.5—/0✱
Professional52.7#53/168
Cybersecurity58.2#2/9
IT operations40.1#9/11
Biology research47.5#10/16
Legal55.9#34/151
Finance54.6#51/151
Medical52.9#62/140
Knowledge50.6—/138✱
Multimodal understanding50.8—/19✱

§ 2 · Head to head

What it beats, and what beats it

The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for DeepSeek V4.1 Flash, left for the other.

DeepSeek V4 Flashwon 24 · lost 2
GPT-6 Lunawon 21 · lost 2
GPT-5.4 Nanowon 19 · lost 2
Grok 4.3won 19 · lost 2
Qwen3.6 Pluswon 19 · lost 2
Gemini 3.5 Flash Litewon 19 · lost 2
GPT-5.4 Miniwon 18 · lost 2
Kimi K2.6won 18 · lost 3
Gemini 3.1 Flash Litewon 16 · lost 3
GLM-5.2won 21 · lost 4
Qwen3.7 Maxwon 17 · lost 4
MiniMax M3won 19 · lost 5
Inklingwon 19 · lost 5
Claude Sonnet 4.6won 16 · lost 5
Qwen3.8 27Bwon 16 · lost 6
GPT-5.6 Lunawon 19 · lost 8
DeepSeek V4 Prowon 17 · lost 8
GPT-5.2won 13 · lost 7
Grok 4.5won 14 · lost 8
GLM-5.3 Flashwon 13 · lost 8
Muse Spark 1.1won 13 · lost 8
GPT-5.6 Terrawon 14 · lost 9
Grok 4.7won 12 · lost 8
Claude Opus 4.5won 12 · lost 8
GLM-5.3won 15 · lost 10
Gemini 3.6 Flashwon 13 · lost 9
GPT-5.4won 12 · lost 9
Claude Sonnet 5won 13 · lost 10
Gemini 3.5 Flashwon 11 · lost 10
Gemini 3.8 Flashwon 11 · lost 12
Gemini 3.1 Prowon 10 · lost 11
Claude Opus 4.7won 10 · lost 11
Grok 4.6won 10 · lost 13
Qwen3.8 Maxwon 9 · lost 12
Muse Spark 1.2won 9 · lost 12
Claude Opus 4.6won 8 · lost 11
GPT-6 Solwon 8 · lost 12
Gemini 3.7 Flashwon 8 · lost 14
Claude Opus 4.8won 7 · lost 14
Kimi K3won 9 · lost 18
Claude Opus 5won 7 · lost 21
GPT-5.6 Solwon 6 · lost 22
GPT-5.5won 4 · lost 18
Muse Spark 1.3won 3 · lost 16
GPT-6 Astrawon 3 · lost 19
Claude Opus 5.5won 2 · lost 19
Claude Fable 5.1won 1 · lost 22
Claude Fable 5won 0 · lost 21

§ 3 · Sources

Where the numbers come from

9 publications, 45 figures. Every one links to the page it was read from.

LMArena Text ↗1 measureread 2026-09-26
LMArena Text 1477
Artificial Analysis ↗1 measureread 2026-09-26
Artificial Analysis Intelligence Index 39
GDPval-AA ↗1 measureread 2026-09-26
GDPval-AA 1328
AA-Briefcase ↗1 measureread 2026-09-26
AA-Briefcase 1425
LiveBench ↗8 measuresread 2026-09-26
LiveBench 81.1LiveBench · Reasoning 86.7LiveBench · Coding 80LiveBench · Agentic Coding 77.3LiveBench · Mathematics 93.3LiveBench · Data Analysis 79.3LiveBench · Language 81.2LiveBench · Instruction Following 70
SimpleBench ↗1 measureread 2026-09-26
SimpleBench 66.7%
Vals.ai ↗12 measuresread 2026-09-26
Vals · Legal Research Bench 41.35%Vals · LegalBench 83.28%Vals · Harvey Legal Agent Benchmark 6.67%Vals · Finance Agent 53.48%Vals · MedCode 41.17%Vals · MedScribe 85.5%Vals · BioMysteryBench 67.78%Vals · CyberBench 73.69%Vals · SRE Bench 0.76%Vals · Vibe Code Bench 84.74% (OpenHands)Vals · Code Migration 45.62%Vals · ProofBench 54%
JevBench ↗2 measuresread 2026-09-26
JevBench · Intelligence 94%JevBench · Calibration 95.5%
GPQA Diamond (Pass@1) 90.9%Codeforces (Rating) 3471MathArena Apex (Pass@1) 65.6%Terminal-Bench 2.1 (Pass@1) 90.6%Terminal-Bench 3.0 (Pass@1) 30%Terminal-Bench 4.0 (Pass@1) 31.2%DeepSWE v1.1 (Resolved) 74.2%ProgramBench (Almost@1) 20.3%NL2Repo-Bench (Score) 64%CyberGym (Pass@1) 88.1%SEC-Bench Pro (Pass@1) 62.8%ExploitGym (Pass@1) 15.3%HLE w/ tools (Pass@1) 63.9%AutomationBench (Pass@1) 54.8%Agent's Last Exam (Pass@1) 31.8%Chartography w/ tools (Pass@1) 78.9%BabyVision w/ tools (Pass@1) 89.6%ZeroBench-main w/ tools (Pass@5) 49%

Badge

PublicAI Index badge for DeepSeek V4.1 Flash[![PublicAI Index](https://publicai.io/model-index/badge?model=deepseek-v4-1-flash)](https://publicai.io/model-index/m/deepseek-v4-1-flash)

✱ Placed by a one-off publication, not a board that re-ran the model. ✱✱ marks a figure the model’s own publisher printed, which counts for less again. Scores are 0–100 on the PublicAI Index scale, 50 = the average of the models each source lists. Scores belong to their publishers.

Snapshot 2026-09-26 · the full index · JSON API