‹ PublicAI Index
The LLM benchmark aggregator.
Kimi K2
Moonshot · 1T · open weights
Strongest in Tool use (#9 of 81), weakest in Toxicity avoidance (#244 of 272). Above par in 13 of 26 scopes. Among the models it meets almost everywhere, it finishes behind Claude Opus 5.5 and Claude Fable 5 and ahead of GPT OSS 120B and Qwen3.7 Max.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Agents58.5−9.6#43/2681/5
Tool use65.4−8.6#9/811/1
Coding53.1−16#47/1651/5
Agentic coding53.8−14.9#50/1571/4
Professional49.1−12.5#91/1681/1
Legal52.5−14.6#59/1511/1
Medical50.2−16.6#79/1401/1
Finance46−16.9#109/1511/1
Knowledge48.1−15.4#98/1381/2
Academic knowledge47.7−16.2#90/1231/1
Core abilities49.8−17.6#100/2041/3
General intelligence49.7−20.4#101/2041/3
Human preference58.8−8.7#103/3421/1
Human preference58.8−8.7#103/3421/1
Reasoning45−21.9#129/1782/4
Mathematics49.1−16.6#84/1401/2
Science46.8−16.7#87/1221/1
Reasoning39.3−29.8#115/1401/3
Safety51.9−9.6#135/3373/3
Safe-prompt compliance59.1−1.3#11/821/1
Harm refusal57.4−4.2#28/3002/2
Secure code63.2−3.3#30/2741/1
Fairness54.3−15.7#73/3002/2
Factual grounding30−40.6#92/1011/1
Jailbreak resistance37.7−27.2#228/2721/1
Toxicity avoidance37.3−21.7#244/2721/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Kimi K2, left for the other.
§ 3 · Sources
Where the numbers come from
9 publications, 24 figures. Every one links to the page it was read from.
LMArena Text 1419
Aider polyglot 59.1%
BFCL v4 59.06%
Kagi LLM Benchmark 52.2%
SimpleBench 26.3%
Vals · LegalBench 81.45%Vals · CorpFin 50.39%Vals · TaxEval 70.2%Vals · MedQA 83.97%Vals · GPQA Diamond 71.46%Vals · MMLU Pro 79.39%Vals · AIME 62.71%
HELM Safety · HarmBench 97.4%HELM Safety · SimpleSafetyTests 100%HELM Safety · Anthropic Red Team 99.3%HELM Safety · BBQ 94.9%HELM Safety · XSTest 98.2%
Enkrypt · Jailbreak risk 23.2%Enkrypt · Harmful content risk 15.6%Enkrypt · CBRN risk 12.8%Enkrypt · Toxicity risk 15.2%Enkrypt · Bias risk 74.9%Enkrypt · Insecure code risk 6.7%
Vectara · Factual consistency 82.1%
Badge
[](https://publicai.io/model-index/m/kimi-k2)