‹ PublicAI Index
The LLM benchmark aggregator.
Kimi K2 Thinking
Moonshot · 1T · open weights
Strongest in Legal (#37 of 151), weakest in Harm refusal (#191 of 300). Above par in 14 of 20 scopes. Among the models it meets almost everywhere, it finishes behind GPT-5 and Claude Opus 5 and ahead of GPT-6 Luna and Mimo V2.6 Flash.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Professional53.8−7.8#44/1681/1
Legal55.6−11.5#37/1511/1
Medical54.8−12#40/1401/1
Finance52.3−10.6#74/1511/1
Core abilities54−13.4#55/2041/3
General intelligence55.7−14.4#57/2041/3
Reasoning50.7−16.2#84/1782/4
Mathematics55.1−10.6#47/1401/2
Science51.7−11.8#72/1221/1
Reasoning45.2−23.9#80/1401/3
Knowledge49.8−13.7#87/1381/2
Academic knowledge49.8−14.1#80/1231/1
Coding45.2−23.9#112/1651/5
Agentic coding44.2−24.5#109/1571/4
Safety52.2−9.3#128/3371/3
Fairness56.7−13.3#54/3001/2
Secure code56.9−9.6#97/2741/1
Toxicity avoidance54.8−4.2#114/2721/1
Jailbreak resistance51.4−13.5#168/2721/1
Harm refusal47.9−13.7#191/3001/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Kimi K2 Thinking, left for the other.
§ 3 · Sources
Where the numbers come from
4 publications, 17 figures. Every one links to the page it was read from.
Kagi LLM Benchmark 64.4%
SimpleBench 39.6%
Vals · CaseLaw 65.7%Vals · LegalBench 80.2%Vals · CorpFin 60.57%Vals · TaxEval 71.71%Vals · MedQA 92.59%Vals · SWE-bench Verified 60.2% (Mini-SWE-agent)Vals · GPQA Diamond 78.54%Vals · MMLU Pro 81.07%Vals · AIME 85.42%
Enkrypt · Jailbreak risk 11.6%Enkrypt · Harmful content risk 26.1%Enkrypt · CBRN risk 24.3%Enkrypt · Toxicity risk 2.9%Enkrypt · Bias risk 67.4%Enkrypt · Insecure code risk 19.6%
Badge
[](https://publicai.io/model-index/m/kimi-k2-thinking)