Request a pilot
‹ PublicAI Index

The LLM benchmark aggregator.

DeepSeek V4 Flash

DeepSeek

Strongest in Biology research (#12 of 16), weakest in Safety (#275 of 337). Above par in 17 of 33 scopes. Among the models it meets almost everywhere, it finishes behind Claude Opus 5.5 and Claude Fable 5.1 and ahead of MiniMax M3 and Gemini 3.6 Flash.

§ 1 · Profile

What it is good at

Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.

Agents59.6#34/268
Knowledge work62.4#28/178
Workflow automation36.5—/0✱
Coding53.2#46/165
Agentic coding55.5#36/157
Code generation45.1#56/77
Repository Q&A41.1—/0✱
Reasoning54#49/178
Science59.6#27/122
Reasoning56.6#33/140
Mathematics48.2#85/140
Expert reasoning39.2—/0✱
Knowledge55.2#50/138
Academic knowledge56.2#46/123
Professional50.9#72/168
Biology research41.4#12/16
Legal52.1#63/151
Finance52.8#69/151
Medical50.7#75/140
Cybersecurity35.9—/9✱
Human preference60.8#77/342
Human preference60.8#77/342
Core abilities50.8#83/204
Data analysis59.9#13/57
Language49.8#32/57
Instruction following45.8#36/57
General intelligence49.2#108/204
Safety46#275/337
Secure code49.4#161/274
Fairness45.2#192/300
Harm refusal46.4#217/300
Jailbreak resistance40.3#219/272
Toxicity avoidance47#220/272

§ 2 · Head to head

What it beats, and what beats it

The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for DeepSeek V4 Flash, left for the other.

DeepSeek V3won 22 · lost 3
GPT-5.4 Nanowon 20 · lost 3
GPT-5.4 Miniwon 19 · lost 3
Llama 4 Scout Instructwon 17 · lost 5
Command Awon 17 · lost 5
Gemini 2.0 Flashwon 17 · lost 5
Gemini 3.5 Flash Litewon 17 · lost 5
GPT OSS 20Bwon 16 · lost 5
Grok 4.1 Fastwon 16 · lost 5
Claude 3.5 Haikuwon 15 · lost 5
GPT-4o Miniwon 18 · lost 6
GPT-4.1 Nanowon 18 · lost 6
Claude Haiku 4.5won 18 · lost 6
Grok 4.3won 21 · lost 7
GLM-4.6won 17 · lost 6
GPT OSS 120Bwon 17 · lost 7
GPT-4.1 Miniwon 17 · lost 7
DeepSeek V3.2won 17 · lost 7
Grok 3won 14 · lost 6
Kimi K2won 16 · lost 7
Mimo V2.5 Prowon 16 · lost 7
Qwen3.6 Pluswon 15 · lost 7
Claude 3.5 Sonnetwon 13 · lost 7
DeepSeek R1won 15 · lost 9
MiniMax M2.7won 13 · lost 8
Grok 4won 14 · lost 9
Qwen3.8 27Bwon 14 · lost 9
Kimi K2.6won 17 · lost 11
Gemini 2.5 Flashwon 14 · lost 10
GPT-4owon 12 · lost 9
GPT-5 Nanowon 12 · lost 9
O3 Miniwon 13 · lost 10
Claude Sonnet 4won 14 · lost 11
Claude 3.7 Sonnetwon 11 · lost 10
GPT-4.1won 12 · lost 11
Gemini 3.6 Flashwon 15 · lost 14
MiniMax M3won 16 · lost 15
Gemma 4 31Bwon 11 · lost 11
GLM-5.3 Flashwon 14 · lost 14
Qwen3.7 Maxwon 14 · lost 14
Inklingwon 15 · lost 16
Qwen3 235B A22B Instructwon 11 · lost 12
GPT-6 Lunawon 12 · lost 14
Claude Sonnet 4.6won 12 · lost 14
GPT-5 Miniwon 11 · lost 13
O4 Miniwon 11 · lost 13
Kimi K2.5won 11 · lost 13
GLM-5.1won 10 · lost 12
O1won 9 · lost 11
Gemini 2.5 Prowon 11 · lost 14
GLM-5.2won 13 · lost 18
Kimi K2 Thinkingwon 8 · lost 12
Claude Opus 4won 9 · lost 14
GPT-5won 9 · lost 15
Gemini 3.5 Flashwon 10 · lost 18
GPT-5.6 Lunawon 11 · lost 20
O3won 8 · lost 16
Claude Sonnet 4.5won 8 · lost 16
GPT-5.1won 7 · lost 15
Claude Opus 4.5won 7 · lost 15
GPT-5.2won 8 · lost 19
Muse Spark 1.1won 8 · lost 20
Gemini 3.8 Flashwon 8 · lost 22
Gemini 3.1 Prowon 6 · lost 17
Claude Opus 4.6won 6 · lost 18
Grok 4.5won 7 · lost 21
Grok 4.7won 5 · lost 21
GLM-5.3won 5 · lost 21
Gemini 3.7 Flashwon 4 · lost 18
Claude Opus 4.7won 5 · lost 23
Claude Sonnet 5won 5 · lost 25
Gemini 3 Prowon 3 · lost 17
GPT-5.4won 4 · lost 24
GPT-6 Solwon 3 · lost 23
Qwen3.8 Maxwon 3 · lost 25
Muse Spark 1.2won 2 · lost 20
DeepSeek V4.1 Flashwon 2 · lost 24
GPT-5.5won 2 · lost 26
DeepSeek V4 Prowon 2 · lost 30
Claude Opus 5won 2 · lost 31
Mimo V2.6 Prowon 1 · lost 19
GPT-6 Astrawon 1 · lost 21
Kimi K3won 1 · lost 26
Claude Opus 4.8won 1 · lost 27
Grok 4.6won 1 · lost 29
GPT-5.6 Terrawon 1 · lost 29
GPT-5.6 Solwon 1 · lost 32
Claude Fable 5won 0 · lost 22
Claude Fable 5.1won 0 · lost 23
Claude Opus 5.5won 0 · lost 26

§ 3 · Sources

Where the numbers come from

10 publications, 49 figures. Every one links to the page it was read from.

LMArena Text ↗1 measureread 2026-09-26
LMArena Text 1439
GDPval-AA ↗1 measureread 2026-09-26
GDPval-AA 1427
AA-Briefcase ↗1 measureread 2026-09-26
AA-Briefcase 1257
ARC-AGI-2 ↗1 measureread 2026-09-26
ARC-AGI-2 61.4%
LiveBench ↗8 measuresread 2026-09-26
LiveBench 74.2LiveBench · Reasoning 86.6LiveBench · Coding 75LiveBench · Agentic Coding 46.8LiveBench · Mathematics 86.8LiveBench · Data Analysis 79.3LiveBench · Language 79.2LiveBench · Instruction Following 65.5
Kagi LLM Benchmark ↗1 measureread 2026-09-26
Kagi LLM Benchmark 52.2%
SimpleBench ↗1 measureread 2026-09-26
SimpleBench 61.1%
Vals.ai ↗15 measuresread 2026-09-26
Vals · Legal Research Bench 30.29%Vals · LegalBench 77.71%Vals · Harvey Legal Agent Benchmark 8.33%Vals · Finance Agent 49.52%Vals · CorpFin 61.85%Vals · TaxEval 70.69%Vals · MedCode 41.41%Vals · MedScribe 80.36%Vals · BioMysteryBench 64.44%Vals · SWE-bench Verified 88.8% (Mini-SWE-agent)Vals · Vibe Code Bench 74.74% (OpenHands)Vals · Code Migration 38.63%Vals · GPQA Diamond 89.9%Vals · MMLU Pro 86.21%Vals · ProofBench 56%
Enkrypt AI Safety Leaderboard ↗6 measuresread 2026-09-26
Enkrypt · Jailbreak risk 21%Enkrypt · Harmful content risk 23.3%Enkrypt · CBRN risk 29.7%Enkrypt · Toxicity risk 8.4%Enkrypt · Bias risk 85.3%Enkrypt · Insecure code risk 34.7%
GPQA Diamond (Pass@1) 89.9%Codeforces (Rating) 3289MathArena Apex (Pass@1) 58.6%Terminal-Bench 2.1 (Pass@1) 82.7%Terminal-Bench 3.0 (Pass@1) 7.6%Terminal-Bench 4.0 (Pass@1) 7%DeepSWE v1.1 (Resolved) 54.4%NL2Repo-Bench (Score) 54.2%CyberGym (Pass@1) 76.7%SEC-Bench Pro (Pass@1) 30.9%ExploitGym (Pass@1) 1.8%HLE w/ tools (Pass@1) 51.5%AutomationBench (Pass@1) 37.7%Agent's Last Exam (Pass@1) 25.2%

Badge

PublicAI Index badge for DeepSeek V4 Flash[![PublicAI Index](https://publicai.io/model-index/badge?model=deepseek-v4-flash)](https://publicai.io/model-index/m/deepseek-v4-flash)

✱ Placed by a one-off publication, not a board that re-ran the model. ✱✱ marks a figure the model’s own publisher printed, which counts for less again. Scores are 0–100 on the PublicAI Index scale, 50 = the average of the models each source lists. Scores belong to their publishers.

Snapshot 2026-09-26 · the full index · JSON API