‹ PublicAI Index
The LLM benchmark aggregator.
DeepSeek V4 Flash
DeepSeek
Strongest in Biology research (#12 of 16), weakest in Safety (#275 of 337). Above par in 17 of 33 scopes. Among the models it meets almost everywhere, it finishes behind Claude Opus 5.5 and Claude Fable 5.1 and ahead of MiniMax M3 and Gemini 3.6 Flash.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Agents59.6−8.5#34/2682/5
Knowledge work62.4−11.2#28/1782/2
Workflow automation36.5−25.5—/0✱✱
Coding53.2−15.9#46/1652/5
Agentic coding55.5−13.2#36/1572/4
Code generation45.1−24.9#56/771/2
Repository Q&A41.1−25.6—/0✱✱
Reasoning54−12.9#49/1784/4
Science59.6−3.9#27/1221/1
Reasoning56.6−12.5#33/1403/3
Mathematics48.2−17.5#85/1402/2
Expert reasoning39.2−21.7—/0✱✱
Knowledge55.2−8.3#50/1381/2
Academic knowledge56.2−7.7#46/1231/1
Professional50.9−10.7#72/1681/1
Biology research41.4−27.1#12/161/1
Legal52.1−15#63/1511/1
Finance52.8−10.1#69/1511/1
Medical50.7−16.1#75/1401/1
Cybersecurity35.9−23.8—/9✱0/1
Human preference60.8−6.7#77/3421/1
Human preference60.8−6.7#77/3421/1
Core abilities50.8−16.6#83/2042/3
Data analysis59.9−6.2#13/571/1
Language49.8−21.7#32/571/1
Instruction following45.8−27.9#36/571/1
General intelligence49.2−20.9#108/2042/3
Safety46−15.5#275/3371/3
Secure code49.4−17.1#161/2741/1
Fairness45.2−24.8#192/3001/2
Harm refusal46.4−15.2#217/3001/2
Jailbreak resistance40.3−24.6#219/2721/1
Toxicity avoidance47−12#220/2721/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for DeepSeek V4 Flash, left for the other.
§ 3 · Sources
Where the numbers come from
10 publications, 49 figures. Every one links to the page it was read from.
LMArena Text 1439
GDPval-AA 1427
AA-Briefcase 1257
ARC-AGI-2 61.4%
LiveBench 74.2LiveBench · Reasoning 86.6LiveBench · Coding 75LiveBench · Agentic Coding 46.8LiveBench · Mathematics 86.8LiveBench · Data Analysis 79.3LiveBench · Language 79.2LiveBench · Instruction Following 65.5
Kagi LLM Benchmark 52.2%
SimpleBench 61.1%
Vals · Legal Research Bench 30.29%Vals · LegalBench 77.71%Vals · Harvey Legal Agent Benchmark 8.33%Vals · Finance Agent 49.52%Vals · CorpFin 61.85%Vals · TaxEval 70.69%Vals · MedCode 41.41%Vals · MedScribe 80.36%Vals · BioMysteryBench 64.44%Vals · SWE-bench Verified 88.8% (Mini-SWE-agent)Vals · Vibe Code Bench 74.74% (OpenHands)Vals · Code Migration 38.63%Vals · GPQA Diamond 89.9%Vals · MMLU Pro 86.21%Vals · ProofBench 56%
Enkrypt · Jailbreak risk 21%Enkrypt · Harmful content risk 23.3%Enkrypt · CBRN risk 29.7%Enkrypt · Toxicity risk 8.4%Enkrypt · Bias risk 85.3%Enkrypt · Insecure code risk 34.7%
GPQA Diamond (Pass@1) 89.9%Codeforces (Rating) 3289MathArena Apex (Pass@1) 58.6%Terminal-Bench 2.1 (Pass@1) 82.7%Terminal-Bench 3.0 (Pass@1) 7.6%Terminal-Bench 4.0 (Pass@1) 7%DeepSWE v1.1 (Resolved) 54.4%NL2Repo-Bench (Score) 54.2%CyberGym (Pass@1) 76.7%SEC-Bench Pro (Pass@1) 30.9%ExploitGym (Pass@1) 1.8%HLE w/ tools (Pass@1) 51.5%AutomationBench (Pass@1) 37.7%Agent's Last Exam (Pass@1) 25.2%
Badge
[](https://publicai.io/model-index/m/deepseek-v4-flash)