‹ PublicAI Index
The LLM benchmark aggregator.
Kimi K2.6
Moonshot · 1T · open weights
Strongest in Computer use (#7 of 37), weakest in Harm refusal (#248 of 300). Above par in 14 of 30 scopes. Among the models it meets almost everywhere, it finishes behind Claude Opus 5.5 and Claude Fable 5 and ahead of O4 Mini and GPT-5 Mini.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Knowledge56.6−6.9#28/1381/2
Academic knowledge57.9−6#25/1231/1
Human preference62.9−4.6#47/3421/1
Human preference62.9−4.6#47/3421/1
Agents56−12.1#56/2683/5
Computer use64.7−7#7/371/1
Knowledge work52.9−20.7#64/1782/2
Professional51.2−10.4#68/1681/1
Finance55.2−7.7#43/1511/1
Legal48.7−18.4#88/1511/1
Medical48.8−18#90/1401/1
Coding49.2−19.9#80/1652/5
Code generation52.5−17.5#37/771/2
Agentic coding48.2−20.5#81/1572/4
Reasoning47.5−19.4#104/1782/4
Science59.1−4.4#31/1221/1
Reasoning43.9−25.2#83/1401/3
Mathematics42.8−22.9#106/1401/2
Core abilities41.4−26#181/2041/3
Instruction following43.8−29.9#39/571/1
Language42.1−29.4#40/571/1
Data analysis36−30.1#50/571/1
General intelligence43.1−27#148/2041/3
Safety48.2−13.3#232/3372/3
Secure code60.4−6.1#63/2741/1
Factual grounding47.9−22.7#67/1011/1
Fairness54.2−15.8#74/3001/2
Toxicity avoidance52.3−6.7#175/2721/1
Jailbreak resistance33−31.9#242/2721/1
Harm refusal44.9−16.7#248/3001/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Kimi K2.6, left for the other.
§ 3 · Sources
Where the numbers come from
8 publications, 34 figures. Every one links to the page it was read from.
LMArena Text 1461
GDPval-AA 1026
AA-Briefcase 818
LiveBench 70.5LiveBench · Reasoning 79.4LiveBench · Coding 78.6LiveBench · Agentic Coding 46.9LiveBench · Mathematics 84.3LiveBench · Data Analysis 65.1LiveBench · Language 75.1LiveBench · Instruction Following 64.4
OSWorld 73.1%
Vals · Legal Research Bench 15.87%Vals · CaseLaw 61.2%Vals · LegalBench 84.74%Vals · Harvey Legal Agent Benchmark 1.67%Vals · Finance Agent 44.9%Vals · CorpFin 66.74%Vals · TaxEval 74.65%Vals · MortgageTax 65.82%Vals · MedCode 40.14%Vals · MedScribe 78.15%Vals · SWE-bench Verified 76.2% (Mini-SWE-agent)Vals · Vibe Code Bench 37.89% (OpenHands)Vals · Code Migration 27.77%Vals · GPQA Diamond 89.14%Vals · MMLU Pro 87.57%
Enkrypt · Jailbreak risk 27.1%Enkrypt · Harmful content risk 9.4%Enkrypt · CBRN risk 48.5%Enkrypt · Toxicity risk 4.7%Enkrypt · Bias risk 71.3%Enkrypt · Insecure code risk 12.4%
Vectara · Factual consistency 89.2%
Badge
[](https://publicai.io/model-index/m/kimi-k2-6)