Request a pilot
‹ PublicAI Index

The LLM benchmark aggregator.

GPT-5.2

OpenAI

Strongest in Secure code (#1 of 274), weakest in Toxicity avoidance (#126 of 272). Above par in 24 of 29 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5 and Claude Fable 5.1 and ahead of Grok 4.5 and Gemini 3.5 Flash.

§ 1 · Profile

What it is good at

Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.

Professional57.3#16/168
Medical58.3#19/140
Finance57.7#23/151
Legal57.2#24/151
Safety57.7#19/337
Secure code66.5#1/274
Fairness62.5#29/300
Harm refusal56.8#36/300
Jailbreak resistance60.4#37/272
Factual grounding47.9#66/101
Toxicity avoidance54.3#126/272
Reasoning55.3#39/178
Mathematics58.9#14/140
Science60.9#20/122
Reasoning50#61/140
Agents57.3#49/268
Tool use63.1#13/81
Knowledge55.2#49/138
Academic knowledge56.3#45/123
Core abilities52.3#68/204
Data analysis58.1#19/57
Language51#28/57
Instruction following39.3#47/57
General intelligence56.7#50/204
Coding49.9#76/165
Code generation47.3#52/77
Agentic coding50.8#67/157
Human preference60.6#81/342
Human preference60.6#81/342

§ 2 · Head to head

What it beats, and what beats it

The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for GPT-5.2, left for the other.

Gemini 1.5 Prowon 18 · lost 0
GLM-4.5won 19 · lost 0
Mistral Large 3won 19 · lost 0
GPT OSS 20Bwon 20 · lost 0
Gemini 2.0 Flashwon 22 · lost 0
GPT-4.1 Nanowon 24 · lost 0
GPT-4.1 Miniwon 24 · lost 0
GPT OSS 120Bwon 23 · lost 1
Command Awon 23 · lost 1
GPT-4o Miniwon 22 · lost 1
Mistral Smallwon 20 · lost 1
Gemini 2.5 Flash Litewon 18 · lost 1
Grok Build 0.1won 17 · lost 1
Grok 4 Fastwon 17 · lost 1
DeepSeek V3won 23 · lost 2
DeepSeek V3.2won 23 · lost 2
Llama 4 Scout Instructwon 21 · lost 2
GPT-5.4 Nanowon 21 · lost 2
Grok 3won 19 · lost 2
Kimi K2 Thinkingwon 18 · lost 2
Mistral Largewon 17 · lost 2
Kimi K2.7 Codewon 16 · lost 2
Grok 4.3won 24 · lost 3
GLM-4.6won 21 · lost 3
Grok 4.1 Fastwon 20 · lost 3
GPT-4owon 19 · lost 3
GPT-5.4 Miniwon 19 · lost 3
Mimo V2.5 Prowon 19 · lost 3
Gemini 3.5 Flash Litewon 18 · lost 3
GPT-4.1won 21 · lost 4
Kimi K2.5won 20 · lost 4
MiniMax M3won 23 · lost 5
O3 Miniwon 18 · lost 4
O4 Miniwon 21 · lost 5
Claude Haiku 4.5won 20 · lost 5
Kimi K2won 20 · lost 5
Qwen3 235B A22B Instructwon 20 · lost 5
Kimi K2.6won 22 · lost 6
Qwen3.7 Maxwon 21 · lost 6
Gemini 2.5 Flashwon 20 · lost 6
Claude 3.7 Sonnetwon 16 · lost 5
Qwen3.6 Pluswon 16 · lost 5
GLM-5.1won 16 · lost 5
GPT-5 Miniwon 19 · lost 6
MiniMax M2.7won 15 · lost 5
DeepSeek R1won 18 · lost 6
GPT-5 Nanowon 17 · lost 6
Claude 3.5 Haikuwon 14 · lost 5
Gemma 4 31Bwon 16 · lost 6
Gemini 3.1 Flash Litewon 13 · lost 5
GLM-5won 13 · lost 5
Gemini 2.5 Prowon 18 · lost 7
GPT-6 Lunawon 17 · lost 7
DeepSeek V4 Flashwon 19 · lost 8
O1won 14 · lost 6
Claude Sonnet 4won 17 · lost 8
O3won 17 · lost 8
Inklingwon 19 · lost 9
GPT-5.6 Lunawon 19 · lost 9
Claude Opus 4won 16 · lost 8
GLM-5.3 Flashwon 18 · lost 9
Claude 3.5 Sonnetwon 13 · lost 7
GLM-5.2won 18 · lost 10
Claude Sonnet 4.5won 16 · lost 9
Qwen3.8 27Bwon 14 · lost 8
GPT-5.1won 14 · lost 8
Grok 4won 15 · lost 9
Gemini 3.6 Flashwon 16 · lost 11
Claude Sonnet 4.6won 15 · lost 11
Gemini 3.5 Flashwon 15 · lost 12
Grok 4.5won 14 · lost 13
Gemini 3 Prowon 11 · lost 11
Claude Opus 4.5won 12 · lost 12
GPT-5won 11 · lost 13
Mimo V2.6 Prowon 8 · lost 10
Qwen3.8 Maxwon 12 · lost 15
DeepSeek V4 Prowon 12 · lost 16
GPT-5.6 Terrawon 12 · lost 16
Gemini 3.8 Flashwon 11 · lost 16
GPT-6 Solwon 10 · lost 15
Claude Sonnet 5won 11 · lost 17
GPT-5.4won 11 · lost 17
GPT-5.6 Solwon 11 · lost 17
Gemini 3.1 Prowon 9 · lost 14
Muse Spark 1.1won 10 · lost 17
DeepSeek V4.1 Flashwon 7 · lost 13
GLM-5.3won 7 · lost 14
Claude Opus 4.6won 8 · lost 17
Grok 4.7won 7 · lost 17
GPT-5.5won 8 · lost 20
Grok 4.6won 7 · lost 20
Claude Opus 4.7won 7 · lost 21
Gemini 3.7 Flashwon 5 · lost 16
GPT-6 Astrawon 4 · lost 16
Muse Spark 1.2won 4 · lost 16
Claude Opus 5won 5 · lost 22
Claude Opus 4.8won 4 · lost 23
Claude Opus 5.5won 3 · lost 21
Kimi K3won 1 · lost 20
Claude Fable 5.1won 0 · lost 21
Claude Fable 5won 0 · lost 21

§ 3 · Sources

Where the numbers come from

9 publications, 33 figures. Every one links to the page it was read from.

LMArena Text ↗1 measureread 2026-09-26
LMArena Text 1437
ARC-AGI-2 ↗1 measureread 2026-09-26
ARC-AGI-2 52.9%
LiveBench ↗8 measuresread 2026-09-26
LiveBench 74.6LiveBench · Reasoning 83.2LiveBench · Coding 76.1LiveBench · Agentic Coding 50.3LiveBench · Mathematics 93.2LiveBench · Data Analysis 78.2LiveBench · Language 79.8LiveBench · Instruction Following 61.8
BFCL ↗1 measureread 2026-09-26
BFCL v4 55.87%
Kagi LLM Benchmark ↗1 measureread 2026-09-26
Kagi LLM Benchmark 73.3%
SimpleBench ↗1 measureread 2026-09-26
SimpleBench 45.8%
Vals.ai ↗13 measuresread 2026-09-26
Vals · CaseLaw 66.02%Vals · LegalBench 82.76%Vals · CorpFin 65.89%Vals · TaxEval 75.76%Vals · MortgageTax 67.13%Vals · MedQA 94.13%Vals · MedCode 49.75%Vals · MedScribe 84.39%Vals · SWE-bench Verified 75.8% (Mini-SWE-agent)Vals · Vibe Code Bench 53.5% (OpenHands)Vals · GPQA Diamond 91.67%Vals · MMLU Pro 86.23%Vals · AIME 96.88%
Enkrypt AI Safety Leaderboard ↗6 measuresread 2026-09-26
Enkrypt · Jailbreak risk 4%Enkrypt · Harmful content risk 8.3%Enkrypt · CBRN risk 10.2%Enkrypt · Toxicity risk 3.3%Enkrypt · Bias risk 58.4%Enkrypt · Insecure code risk 0%
Vectara · Factual consistency 89.2%

Badge

PublicAI Index badge for GPT-5.2[![PublicAI Index](https://publicai.io/model-index/badge?model=gpt-5-2)](https://publicai.io/model-index/m/gpt-5-2)