‹ PublicAI Index
The LLM benchmark aggregator.
Kimi K2.5
Moonshot · 1T · open weights
Strongest in Computer use (#9 of 37), weakest in Safety (#218 of 337). Above par in 16 of 26 scopes. Among the models it meets almost everywhere, it finishes behind Claude Opus 5.5 and Gemini 3 Pro and ahead of GPT-4o and Claude 3.5 Sonnet.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Core abilities58.8−8.6#12/2041/3
General intelligence62.6−7.5#10/2041/3
Knowledge54.9−8.6#55/1381/2
Academic knowledge55.9−8#51/1231/1
Human preference61.8−5.7#62/3421/1
Human preference61.8−5.7#62/3421/1
Agents54.5−13.6#70/2682/5
Computer use59.4−12.3#9/371/1
Knowledge work51.8−21.8#67/1781/2
Reasoning52.3−14.6#74/1783/4
Mathematics57.8−7.9#19/1401/2
Science55.6−7.9#55/1221/1
Reasoning46.2−22.9#74/1402/3
Professional49.6−12#85/1681/1
Finance53.4−9.5#65/1511/1
Medical50.9−15.9#73/1401/1
Legal43.4−23.7#113/1511/1
Coding43.1−26#132/1651/5
Agentic coding42.3−26.4#127/1571/4
Safety48.7−12.8#218/3372/3
Secure code63.7−2.8#25/2741/1
Factual grounding39.3−31.3#88/1011/1
Toxicity avoidance53.9−5.1#133/2721/1
Fairness46.2−23.8#165/3001/2
Harm refusal48.6−13#177/3001/2
Jailbreak resistance40.6−24.3#215/2721/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Kimi K2.5, left for the other.
§ 3 · Sources
Where the numbers come from
9 publications, 29 figures. Every one links to the page it was read from.
LMArena Text 1450
GDPval-AA 824
ARC-AGI-2 11.8%
OSWorld 63.3%
Kagi LLM Benchmark 78.5%
SimpleBench 46.8%
Vals · Legal Research Bench 15.87%Vals · CaseLaw 58.73%Vals · Harvey Legal Agent Benchmark 0%Vals · Finance Agent 35.79%Vals · CorpFin 68.26%Vals · TaxEval 74.2%Vals · MortgageTax 66.53%Vals · MedQA 94.37%Vals · MedCode 39.32%Vals · MedScribe 76.44%Vals · SWE-bench Verified 70% (Mini-SWE-agent)Vals · Vibe Code Bench 17.54% (OpenHands)Vals · Code Migration 6.96%Vals · GPQA Diamond 84.09%Vals · MMLU Pro 85.91%Vals · AIME 95.63%
Enkrypt · Jailbreak risk 20.7%Enkrypt · Harmful content risk 5.6%Enkrypt · CBRN risk 33.2%Enkrypt · Toxicity risk 3.6%Enkrypt · Bias risk 83.7%Enkrypt · Insecure code risk 5.8%
Vectara · Factual consistency 85.8%
Badge
[](https://publicai.io/model-index/m/kimi-k2-5)