‹ PublicAI Index
The LLM benchmark aggregator.
DeepSeek V3.1
DeepSeek · 685B · open weights
Strongest in Factual grounding (#19 of 101), weakest in Professional (#123 of 168). Above par in 6 of 11 scopes. Among the models it meets almost everywhere, it finishes behind Claude Opus 5.5 and Claude Opus 5 and ahead of Kimi K2 and Grok 3.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Safety53.5−8#84/3371/3
Factual grounding61.3−9.3#19/1011/1
Core abilities50.2−17.2#90/2041/3
General intelligence50.2−19.9#91/2041/3
Reasoning47.2−19.7#106/1781/4
Reasoning45.4−23.7#79/1401/3
Human preference58.5−9#106/3421/1
Human preference58.5−9#106/3421/1
Professional45.9−15.7#123/1681/1
Legal44.6−22.5#105/1511/1
Finance44.1−18.8#118/1511/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for DeepSeek V3.1, left for the other.
§ 3 · Sources
Where the numbers come from
5 publications, 6 figures. Every one links to the page it was read from.
LMArena Text 1416
Kagi LLM Benchmark 53.2%
SimpleBench 40%
Vals · CaseLaw 53.91%Vals · CorpFin 51.48%
Vectara · Factual consistency 94.5%
Badge
[](https://publicai.io/model-index/m/deepseek-v3-1)