‹ PublicAI Index
The LLM benchmark aggregator.
DeepSeek V3.2
DeepSeek · 685B · open weights
Strongest in Tool use (#11 of 81), weakest in Harm refusal (#254 of 300). Above par in 14 of 26 scopes. Among the models it meets almost everywhere, it finishes behind Claude Opus 5 and Claude Opus 5.5 and ahead of Kimi K2 and Gemini 2.5 Flash.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Agents55−13.1#64/2682/5
Tool use63.7−10.3#11/811/1
Knowledge work49.4−24.2#86/1781/2
Knowledge53.8−9.7#64/1381/2
Academic knowledge54.6−9.3#59/1231/1
Reasoning51−15.9#82/1782/4
Mathematics54.9−10.8#53/1401/2
Science52.9−10.6#68/1221/1
Reasoning43.6−25.5#90/1401/3
Coding48.8−20.3#83/1652/5
Agentic coding48.6−20.1#79/1572/4
Human preference59.4−8.1#97/3421/1
Human preference59.4−8.1#97/3421/1
Core abilities49.8−17.6#99/2041/3
General intelligence49.7−20.4#100/2041/3
Professional48.2−13.4#102/1681/1
Medical55.6−11.2#34/1401/1
Legal46.4−20.7#97/1511/1
Finance45.3−17.6#113/1511/1
Safety48.9−12.6#213/3372/3
Factual grounding59.3−11.3#26/1011/1
Secure code57.6−8.9#91/2741/1
Fairness50.5−19.5#105/3001/2
Toxicity avoidance51.1−7.9#187/2721/1
Jailbreak resistance36.8−28.1#230/2721/1
Harm refusal44.6−17#254/3001/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for DeepSeek V3.2, left for the other.
§ 3 · Sources
Where the numbers come from
9 publications, 23 figures. Every one links to the page it was read from.
LMArena Text 1425
GDPval-AA 681
ARC-AGI-2 4%
Aider polyglot 74.2%
BFCL v4 56.73%
Kagi LLM Benchmark 52.2%
Vals · CaseLaw 55.41%Vals · LegalBench 76.08%Vals · CorpFin 50.97%Vals · TaxEval 68.15%Vals · MedQA 93.92%Vals · SWE-bench Verified 67.6% (Mini-SWE-agent)Vals · Vibe Code Bench 5.11% (OpenHands)Vals · GPQA Diamond 80.3%Vals · MMLU Pro 84.92%Vals · AIME 84.58%
Enkrypt · Jailbreak risk 23.9%Enkrypt · Harmful content risk 53.3%Enkrypt · CBRN risk 18.5%Enkrypt · Toxicity risk 5.5%Enkrypt · Bias risk 77%Enkrypt · Insecure code risk 18.2%
Vectara · Factual consistency 93.7%
Badge
[](https://publicai.io/model-index/m/deepseek-v3-2)