‹ PublicAI Index
The LLM benchmark aggregator.
Kimi K3
Moonshot · 2.8T · open weights
Strongest in Finance (#6 of 151), weakest in Mathematics (#74 of 140). Above par in 23 of 28 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5.1 and Claude Fable 5 and ahead of Muse Spark 1.2 and Gemini 3 Pro.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Professional58.9−2.7#9/1681/1
Finance59.8−3.1#6/1511/1
Biology research54.2−14.3#6/161/1
Legal60.2−6.9#11/1511/1
Medical58.8−8#14/1401/1
Cybersecurity42.8−16.9—/9✱0/1
Core abilities60.1−7.3#9/2042/3
Language61.7−9.8#10/571/1
General intelligence61.9−8.2#12/2042/3
Data analysis58.9−7.2#15/571/1
Instruction following56.1−17.6#21/571/1
Human preference65.5−2#12/3421/1
Human preference65.5−2#12/3421/1
Coding58.5−10.6#14/1652/5
Code generation58.2−11.8#17/771/2
Agentic coding58.6−10.1#19/1572/4
Repository Q&A46−20.7—/0✱✱
Agents62.7−5.4#17/2682/5
Knowledge work66.5−7.1#15/1782/2
Workflow automation49.9−12.1—/0✱✱
Knowledge57−6.5#25/1381/2
Academic knowledge58.4−5.5#22/1231/1
Multimodal understanding35.9−36.2—/19✱0/1
Reasoning56.7−10.2#27/1784/4
Science61.8−1.7#13/1221/1
Reasoning59.3−9.8#24/1403/3
Mathematics51.5−14.2#74/1402/2
Expert reasoning47.9−13—/0✱✱
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Kimi K3, left for the other.
§ 3 · Sources
Where the numbers come from
9 publications, 46 figures. Every one links to the page it was read from.
LMArena Text 1488
Artificial Analysis Intelligence Index 44
GDPval-AA 1524
AA-Briefcase 1505
ARC-AGI-2 60.4%
LiveBench 79.2LiveBench · Reasoning 90.7LiveBench · Coding 81.4LiveBench · Agentic Coding 62.2LiveBench · Mathematics 84.4LiveBench · Data Analysis 78.7LiveBench · Language 85.5LiveBench · Instruction Following 71.4
SimpleBench 60.7%
Vals · Legal Research Bench 44.23%Vals · LegalBench 86.02%Vals · Harvey Legal Agent Benchmark 10.83%Vals · Finance Agent 54.36%Vals · CorpFin 71.56%Vals · TaxEval 75.72%Vals · MortgageTax 66.34%Vals · MedCode 48.88%Vals · MedScribe 87.96%Vals · BioMysteryBench 71.48%Vals · SWE-bench Verified 93.4% (Mini-SWE-agent)Vals · Vibe Code Bench 84.96% (OpenHands)Vals · Code Migration 16.1%Vals · GPQA Diamond 92.93%Vals · MMLU Pro 87.97%Vals · ProofBench 87%
GPQA Diamond (Pass@1) 92.9%HLE (Pass@1) 43.5%MathArena Apex (Pass@1) 65.6%Terminal-Bench 2.1 (Pass@1) 88.3%Terminal-Bench 3.0 (Pass@1) 17.7%Terminal-Bench 4.0 (Pass@1) 12.6%DeepSWE v1.1 (Resolved) 67.5%ProgramBench (Almost@1) 17.5%NL2Repo-Bench (Score) 58%CyberGym (Pass@1) 80%HLE w/ tools (Pass@1) 59.8%AutomationBench (Pass@1) 46.7%Agent's Last Exam (Pass@1) 27.6%Chartography w/ tools (Pass@1) 68.1%BabyVision w/ tools (Pass@1) 85.7%ZeroBench-main w/ tools (Pass@5) 41%
Badge
[](https://publicai.io/model-index/m/kimi-k3)