‹ PublicAI Index
The LLM benchmark aggregator.
GPT-5.2
OpenAI
Strongest in Secure code (#1 of 274), weakest in Toxicity avoidance (#126 of 272). Above par in 24 of 29 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5 and Claude Fable 5.1 and ahead of Grok 4.5 and Gemini 3.5 Flash.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Professional57.3−4.3#16/1681/1
Medical58.3−8.5#19/1401/1
Finance57.7−5.2#23/1511/1
Legal57.2−9.9#24/1511/1
Safety57.7−3.8#19/3372/3
Secure code66.5leads#1/2741/1
Fairness62.5−7.5#29/3001/2
Harm refusal56.8−4.8#36/3001/2
Jailbreak resistance60.4−4.5#37/2721/1
Factual grounding47.9−22.7#66/1011/1
Toxicity avoidance54.3−4.7#126/2721/1
Reasoning55.3−11.6#39/1784/4
Mathematics58.9−6.8#14/1402/2
Science60.9−2.6#20/1221/1
Reasoning50−19.1#61/1403/3
Agents57.3−10.8#49/2681/5
Tool use63.1−10.9#13/811/1
Knowledge55.2−8.3#49/1381/2
Academic knowledge56.3−7.6#45/1231/1
Core abilities52.3−15.1#68/2042/3
Data analysis58.1−8#19/571/1
Language51−20.5#28/571/1
Instruction following39.3−34.4#47/571/1
General intelligence56.7−13.4#50/2042/3
Coding49.9−19.2#76/1652/5
Code generation47.3−22.7#52/771/2
Agentic coding50.8−17.9#67/1572/4
Human preference60.6−6.9#81/3421/1
Human preference60.6−6.9#81/3421/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for GPT-5.2, left for the other.
§ 3 · Sources
Where the numbers come from
9 publications, 33 figures. Every one links to the page it was read from.
LMArena Text 1437
ARC-AGI-2 52.9%
LiveBench 74.6LiveBench · Reasoning 83.2LiveBench · Coding 76.1LiveBench · Agentic Coding 50.3LiveBench · Mathematics 93.2LiveBench · Data Analysis 78.2LiveBench · Language 79.8LiveBench · Instruction Following 61.8
BFCL v4 55.87%
Kagi LLM Benchmark 73.3%
SimpleBench 45.8%
Vals · CaseLaw 66.02%Vals · LegalBench 82.76%Vals · CorpFin 65.89%Vals · TaxEval 75.76%Vals · MortgageTax 67.13%Vals · MedQA 94.13%Vals · MedCode 49.75%Vals · MedScribe 84.39%Vals · SWE-bench Verified 75.8% (Mini-SWE-agent)Vals · Vibe Code Bench 53.5% (OpenHands)Vals · GPQA Diamond 91.67%Vals · MMLU Pro 86.23%Vals · AIME 96.88%
Enkrypt · Jailbreak risk 4%Enkrypt · Harmful content risk 8.3%Enkrypt · CBRN risk 10.2%Enkrypt · Toxicity risk 3.3%Enkrypt · Bias risk 58.4%Enkrypt · Insecure code risk 0%
Vectara · Factual consistency 89.2%
Badge
[](https://publicai.io/model-index/m/gpt-5-2)