Request a pilot
‹ PublicAI Index

The LLM benchmark aggregator.

MiniMax M3

MiniMax · 427B · open weights

Strongest in Computer use (#5 of 37), weakest in Safety (#240 of 337). Above par in 20 of 36 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5 and Claude Fable 5.1 and ahead of Kimi K2.6 and GPT-5.4 Mini.

§ 1 · Profile

What it is good at

Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.

Agents59.6#33/268
Computer use65.8#5/37
Knowledge work58.4#39/178
Tool use43.9—/81✱
Web research57.5—/4✱
Workflow automation50.8—/0✱
Professional55#34/168
Medical56.9#25/140
Finance56.7#29/151
Legal52.5#60/151
Knowledge53.1#69/138
Academic knowledge53.7#63/123
Factuality51.9—/0✱
Human preference60.9#76/342
Human preference60.9#76/342
Coding42.6#139/165
Code generation31.2#75/77
Agentic coding45.8#101/157
Repository Q&A42.9—/0✱
Reasoning41.4#148/178
Science61.6#14/122
Reasoning40.8#111/140
Mathematics32.8#140/140
Expert reasoning57.6—/0✱
Core abilities44#159/204
Data analysis54.7#27/57
Language45.3#37/57
Instruction following31.7#53/57
General intelligence44.1#139/204
Long context64.6—/0✱
Safety47.7#240/337
Toxicity avoidance55.6#94/272
Secure code56.7#100/274
Fairness48.9#129/300
Harm refusal45.6#233/300
Jailbreak resistance34#239/272

§ 2 · Head to head

What it beats, and what beats it

The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for MiniMax M3, left for the other.

Command Awon 19 · lost 4
GPT-5.4 Nanowon 17 · lost 6
Grok 4.1 Fastwon 16 · lost 6
GPT-4.1 Nanowon 18 · lost 7
GPT-4o Miniwon 17 · lost 7
GLM-4.6won 17 · lost 7
Llama 4 Scout Instructwon 16 · lost 7
Gemini 2.0 Flashwon 15 · lost 7
DeepSeek V3won 17 · lost 8
Nemotron 3 Ultrawon 14 · lost 8
Gemini 3.5 Flash Litewon 14 · lost 8
GPT OSS 120Bwon 15 · lost 9
Grok 4.3won 17 · lost 11
Qwen3.6 Pluswon 13 · lost 9
DeepSeek R1won 14 · lost 10
O3 Miniwon 13 · lost 10
Claude Haiku 4.5won 14 · lost 11
GPT-5.4 Miniwon 12 · lost 10
Kimi K2.6won 15 · lost 14
GPT-5 Nanowon 11 · lost 11
Claude 3.7 Sonnetwon 11 · lost 11
Grok 4won 12 · lost 12
Qwen3 235B A22B Instructwon 12 · lost 12
DeepSeek V4 Flashwon 15 · lost 16
GPT-4.1 Miniwon 12 · lost 13
Gemma 4 31Bwon 12 · lost 13
GPT-4.1won 11 · lost 13
Kimi K2won 11 · lost 13
GPT-5 Miniwon 11 · lost 14
O4 Miniwon 11 · lost 14
Gemini 2.5 Flashwon 11 · lost 14
DeepSeek V3.2won 11 · lost 14
Kimi K2.5won 11 · lost 14
Qwen3.7 Maxwon 12 · lost 16
Inklingwon 15 · lost 20
Claude Sonnet 4won 11 · lost 15
Gemini 2.5 Prowon 10 · lost 15
Gemini 3.6 Flashwon 11 · lost 17
Qwen3.8 27Bwon 10 · lost 16
GLM-5.2won 13 · lost 21
Gemini 3.5 Flashwon 11 · lost 18
GLM-5.1won 8 · lost 14
GPT-6 Lunawon 9 · lost 16
O3won 9 · lost 17
Claude Sonnet 4.5won 8 · lost 18
Claude Opus 4won 7 · lost 16
Mimo V2.5 Prowon 7 · lost 16
Claude Sonnet 4.6won 8 · lost 19
GPT-5won 7 · lost 17
DeepSeek V4 Prowon 9 · lost 22
GPT-5.6 Lunawon 9 · lost 25
GPT-5.4won 7 · lost 21
Grok 4.5won 7 · lost 22
GPT-5.1won 5 · lost 17
Gemini 3.1 Prowon 5 · lost 18
DeepSeek V4.1 Flashwon 5 · lost 19
Muse Spark 1.1won 6 · lost 23
GPT-6 Solwon 5 · lost 20
GLM-5.3won 5 · lost 20
GPT-5.6 Terrawon 6 · lost 27
GPT-5.2won 5 · lost 23
Gemini 3.8 Flashwon 5 · lost 23
GPT-5.6 Solwon 5 · lost 27
Claude Sonnet 5won 5 · lost 29
Grok 4.6won 4 · lost 24
Qwen3.8 Maxwon 4 · lost 24
Claude Opus 4.5won 3 · lost 20
Claude Opus 4.6won 3 · lost 21
Grok 4.7won 3 · lost 22
GLM-5.3 Flashwon 3 · lost 25
Claude Opus 4.8won 3 · lost 25
Gemini 3.7 Flashwon 2 · lost 20
Kimi K3won 2 · lost 23
GPT-5.5won 2 · lost 26
Claude Opus 4.7won 2 · lost 26
Claude Opus 5won 2 · lost 30
Claude Opus 5.5won 1 · lost 24
Claude Fable 5.1won 0 · lost 22
Claude Fable 5won 0 · lost 24

§ 3 · Sources

Where the numbers come from

11 publications, 59 figures. Every one links to the page it was read from.

LMArena Text ↗1 measureread 2026-09-26
LMArena Text 1440
Artificial Analysis ↗1 measureread 2026-09-26
Artificial Analysis Intelligence Index 29
GDPval-AA ↗1 measureread 2026-09-26
GDPval-AA 1230
AA-Briefcase ↗1 measureread 2026-09-26
AA-Briefcase 1090
LiveBench ↗8 measuresread 2026-09-26
LiveBench 67.3LiveBench · Reasoning 74.5LiveBench · Coding 68.2LiveBench · Agentic Coding 40.7LiveBench · Mathematics 76.9LiveBench · Data Analysis 76.2LiveBench · Language 76.8LiveBench · Instruction Following 57.5
OSWorld ↗1 measureread 2026-09-26
OSWorld 75.2%
SimpleBench ↗1 measureread 2026-09-26
SimpleBench 45.8%
Vals.ai ↗15 measuresread 2026-09-26
Vals · Legal Research Bench 29.81%Vals · LegalBench 85.42%Vals · Harvey Legal Agent Benchmark 4.17%Vals · Finance Agent 48.27%Vals · CorpFin 68.1%Vals · TaxEval 72.73%Vals · MortgageTax 68.36%Vals · MedCode 46.29%Vals · MedScribe 87.25%Vals · SWE-bench Verified 75% (Mini-SWE-agent)Vals · Vibe Code Bench 47.57% (OpenHands)Vals · Code Migration 19.93%Vals · GPQA Diamond 92.68%Vals · MMLU Pro 84.22%Vals · ProofBench 18%
Enkrypt AI Safety Leaderboard ↗6 measuresread 2026-09-26
Enkrypt · Jailbreak risk 26.3%Enkrypt · Harmful content risk 5.6%Enkrypt · CBRN risk 49.2%Enkrypt · Toxicity risk 2.4%Enkrypt · Bias risk 79.6%Enkrypt · Insecure code risk 20%
GDPVal-AA 1380tau3-Banking 15.3%Toolathlon Verified 53.7%Automation Bench Public 20.5%Apex-Agents (pass@1) 23.8%MCPMark 48.8%BrowseComp 83.5%WildClawBench 56.4%Terminal-Bench 2.1 65.2%SciCode 45.4%SWE-Atlas-QnA 42.3%SWE Bench Pro 43.8%Humanity's Last Exam (without tools) 39%GPQA Diamond 92.9%CritPt 3.7%AA-LCR 80.3%AA-Omniscience Accuracy 17%AA-Omniscience Non-Hallucination 82%
GDPVal-AA (Elo) 1380Toolathlon Verified 53.7%Terminal-Bench 2.1 65.2%SWE-bench Pro (strict) 43.8%MCPMark 48.8%SWE-Atlas-QnA (strict) 42.3%

Badge

PublicAI Index badge for MiniMax M3[![PublicAI Index](https://publicai.io/model-index/badge?model=minimax-m3)](https://publicai.io/model-index/m/minimax-m3)

✱ Placed by a one-off publication, not a board that re-ran the model. ✱✱ marks a figure the model’s own publisher printed, which counts for less again. Scores are 0–100 on the PublicAI Index scale, 50 = the average of the models each source lists. Scores belong to their publishers.

Snapshot 2026-09-26 · the full index · JSON API