‹ PublicAI Index
The LLM benchmark aggregator.
GPT-5.2 Codex
OpenAI
Strongest in Code generation (#6 of 77, on 1 of its 2 boards), weakest in Reasoning (#120 of 178). Above par in 3 of 11 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5 and Claude Fable 5.1 and ahead of Qwen3.7 Max and Claude Sonnet 4.6.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Coding51.8−17.3#61/1652/5
Code generation62.7−7.3#6/771/2
Agentic coding48−20.7#83/1572/4
Core abilities48.4−19#114/2041/3
Data analysis58.1−8#20/571/1
Instruction following47.4−26.3#34/571/1
Language39.5−32#47/571/1
General intelligence48.7−21.4#109/2041/3
Reasoning45.9−21#120/1781/4
Mathematics49.4−16.3#81/1401/2
Reasoning41.6−27.5#105/1401/3
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for GPT-5.2 Codex, left for the other.
§ 3 · Sources
Where the numbers come from
2 publications, 10 figures. Every one links to the page it was read from.
LiveBench 74LiveBench · Reasoning 77.7LiveBench · Coding 83.6LiveBench · Agentic Coding 49.4LiveBench · Mathematics 88.8LiveBench · Data Analysis 78.2LiveBench · Language 73.7LiveBench · Instruction Following 66.4
Vals · SWE-bench Verified 72.4% (Mini-SWE-agent)Vals · Vibe Code Bench 37.91% (OpenHands)
Badge
[](https://publicai.io/model-index/m/gpt-5-2-codex)