‹ PublicAI Index
The LLM benchmark aggregator.
DeepSeek V4
DeepSeek
Strongest in Science (#30 of 122), weakest in Mathematics (#112 of 140). Above par in 6 of 11 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5.1 and Claude Fable 5 and ahead of Kimi K2 and Claude Sonnet 4.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Knowledge56.3−7.2#35/1381/2
Academic knowledge57.5−6.4#32/1231/1
Coding51.1−18#67/1651/5
Agentic coding51.2−17.5#64/1571/4
Professional49.5−12.1#87/1681/1
Finance51.3−11.6#79/1511/1
Legal49−18.1#85/1511/1
Medical47.6−19.2#93/1401/1
Reasoning48.2−18.7#100/1781/4
Science59.3−4.2#30/1221/1
Mathematics39.6−26.1#112/1401/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for DeepSeek V4, left for the other.
§ 3 · Sources
Where the numbers come from
1 publication, 15 figures. Every one links to the page it was read from.
Vals · Legal Research Bench 23.08%Vals · CaseLaw 59.38%Vals · LegalBench 80.32%Vals · Harvey Legal Agent Benchmark 3.75%Vals · Finance Agent 44.08%Vals · CorpFin 61.38%Vals · TaxEval 72.08%Vals · MedCode 40.45%Vals · MedScribe 75.14%Vals · SWE-bench Verified 77.4% (Mini-SWE-agent)Vals · Vibe Code Bench 49.93% (OpenHands)Vals · Code Migration 26.2%Vals · GPQA Diamond 89.39%Vals · MMLU Pro 87.25%Vals · ProofBench 16%
Badge
[](https://publicai.io/model-index/m/deepseek-v4)