‹ PublicAI Index
The LLM benchmark aggregator.
MiniMax M3
MiniMax · 427B · open weights
Strongest in Computer use (#5 of 37), weakest in Safety (#240 of 337). Above par in 20 of 36 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5 and Claude Fable 5.1 and ahead of Kimi K2.6 and GPT-5.4 Mini.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Agents59.6−8.5#33/2683/5
Computer use65.8−5.9#5/371/1
Knowledge work58.4−15.2#39/1782/2
Tool use43.9−30.1—/81✱0/1
Web research57.5−6.7—/4✱0/1
Workflow automation50.8−11.2—/0✱✱
Professional55−6.6#34/1681/1
Medical56.9−9.9#25/1401/1
Finance56.7−6.2#29/1511/1
Legal52.5−14.6#60/1511/1
Knowledge53.1−10.4#69/1381/2
Academic knowledge53.7−10.2#63/1231/1
Factuality51.9−7.1—/0✱✱
Human preference60.9−6.6#76/3421/1
Human preference60.9−6.6#76/3421/1
Coding42.6−26.5#139/1652/5
Code generation31.2−38.8#75/771/2
Agentic coding45.8−22.9#101/1572/4
Repository Q&A42.9−23.8—/0✱✱
Reasoning41.4−25.5#148/1783/4
Science61.6−1.9#14/1221/1
Reasoning40.8−28.3#111/1402/3
Mathematics32.8−32.9#140/1402/2
Expert reasoning57.6−3.3—/0✱✱
Core abilities44−23.4#159/2042/3
Data analysis54.7−11.4#27/571/1
Language45.3−26.2#37/571/1
Instruction following31.7−42#53/571/1
General intelligence44.1−26#139/2042/3
Long context64.6leads—/0✱✱
Safety47.7−13.8#240/3371/3
Toxicity avoidance55.6−3.4#94/2721/1
Secure code56.7−9.8#100/2741/1
Fairness48.9−21.1#129/3001/2
Harm refusal45.6−16#233/3001/2
Jailbreak resistance34−30.9#239/2721/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for MiniMax M3, left for the other.
§ 3 · Sources
Where the numbers come from
11 publications, 59 figures. Every one links to the page it was read from.
LMArena Text 1440
Artificial Analysis Intelligence Index 29
GDPval-AA 1230
AA-Briefcase 1090
LiveBench 67.3LiveBench · Reasoning 74.5LiveBench · Coding 68.2LiveBench · Agentic Coding 40.7LiveBench · Mathematics 76.9LiveBench · Data Analysis 76.2LiveBench · Language 76.8LiveBench · Instruction Following 57.5
OSWorld 75.2%
SimpleBench 45.8%
Vals · Legal Research Bench 29.81%Vals · LegalBench 85.42%Vals · Harvey Legal Agent Benchmark 4.17%Vals · Finance Agent 48.27%Vals · CorpFin 68.1%Vals · TaxEval 72.73%Vals · MortgageTax 68.36%Vals · MedCode 46.29%Vals · MedScribe 87.25%Vals · SWE-bench Verified 75% (Mini-SWE-agent)Vals · Vibe Code Bench 47.57% (OpenHands)Vals · Code Migration 19.93%Vals · GPQA Diamond 92.68%Vals · MMLU Pro 84.22%Vals · ProofBench 18%
Enkrypt · Jailbreak risk 26.3%Enkrypt · Harmful content risk 5.6%Enkrypt · CBRN risk 49.2%Enkrypt · Toxicity risk 2.4%Enkrypt · Bias risk 79.6%Enkrypt · Insecure code risk 20%
GDPVal-AA 1380tau3-Banking 15.3%Toolathlon Verified 53.7%Automation Bench Public 20.5%Apex-Agents (pass@1) 23.8%MCPMark 48.8%BrowseComp 83.5%WildClawBench 56.4%Terminal-Bench 2.1 65.2%SciCode 45.4%SWE-Atlas-QnA 42.3%SWE Bench Pro 43.8%Humanity's Last Exam (without tools) 39%GPQA Diamond 92.9%CritPt 3.7%AA-LCR 80.3%AA-Omniscience Accuracy 17%AA-Omniscience Non-Hallucination 82%
GDPVal-AA (Elo) 1380Toolathlon Verified 53.7%Terminal-Bench 2.1 65.2%SWE-bench Pro (strict) 43.8%MCPMark 48.8%SWE-Atlas-QnA (strict) 42.3%
Badge
[](https://publicai.io/model-index/m/minimax-m3)