‹ PublicAI Index
The LLM benchmark aggregator.
GPT-5.6 Luna
OpenAI
Strongest in Decisions (#3 of 85), weakest in Toxicity avoidance (#198 of 272). Above par in 31 of 38 scopes. Among the models it meets almost everywhere, it finishes behind Claude Opus 5.5 and Claude Fable 5 and ahead of Gemini 3.6 Flash and GPT-4.1.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Decisions70.6−2.7#3/851/1
Routing & classification74leads#3/851/1
Calibration67.2−5.5#4/851/1
Coding57−12.1#23/1653/5
Code generation61.3−8.7#7/771/2
Agentic coding55.9−12.8#34/1573/4
Agents60.5−7.6#26/2682/5
Knowledge work63.7−9.9#23/1782/2
Tool use57.1−16.9—/81✱0/1
Web research57.4−6.8—/4✱0/1
Workflow automation54.4−7.6—/0✱✱
Professional52.8−8.8#52/1681/1
Biology research35.9−32.6#15/161/1
Finance58−4.9#19/1511/1
Medical53.2−13.6#60/1401/1
Legal51.5−15.6#67/1511/1
Knowledge55−8.5#53/1381/2
Academic knowledge56−7.9#49/1231/1
Factuality50.6−8.4—/0✱✱
Human preference62.2−5.3#57/3421/1
Human preference62.2−5.3#57/3421/1
Reasoning52.7−14.2#70/1784/4
Science60.9−2.6#19/1221/1
Reasoning52.5−16.6#51/1403/3
Mathematics49.3−16.4#82/1402/2
Expert reasoning58−2.9—/0✱✱
Safety53.7−7.8#77/3371/3
Secure code62.6−3.9#39/2741/1
Harm refusal54−7.6#87/3001/2
Fairness49.9−20.1#118/3001/2
Jailbreak resistance53.8−11.1#133/2721/1
Toxicity avoidance50.1−8.9#198/2721/1
Core abilities45.3−22.1#143/2042/3
Data analysis57.7−8.4#21/571/1
Instruction following36.3−37.4#49/571/1
Language37.4−34.1#50/571/1
General intelligence47.6−22.5#118/2042/3
Long context62.2−2.4—/0✱✱
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for GPT-5.6 Luna, left for the other.
§ 3 · Sources
Where the numbers come from
12 publications, 56 figures. Every one links to the page it was read from.
LMArena Text 1454
GDPval-AA 1443
AA-Briefcase 1342
Terminal-Bench 17.3% (Codex)
ARC-AGI-2 59.6%
LiveBench 73.6LiveBench · Reasoning 85.6LiveBench · Coding 82.9LiveBench · Agentic Coding 48.4LiveBench · Mathematics 87.2LiveBench · Data Analysis 78LiveBench · Language 72.6LiveBench · Instruction Following 60.1
Kagi LLM Benchmark 49.1%
SimpleBench 46.8%
Vals · Legal Research Bench 36.54%Vals · LegalBench 84.03%Vals · Harvey Legal Agent Benchmark 1.25%Vals · Finance Agent 55.04%Vals · CorpFin 64.22%Vals · TaxEval 76.17%Vals · MortgageTax 67.29%Vals · MedCode 42.39%Vals · MedScribe 84.39%Vals · BioMysteryBench 61.48%Vals · SWE-bench Verified 93% (Mini-SWE-agent)Vals · Vibe Code Bench 77.06% (OpenHands)Vals · Code Migration 44.55%Vals · GPQA Diamond 91.67%Vals · MMLU Pro 86.04%Vals · ProofBench 60%
JevBench · Intelligence 93.1%JevBench · Calibration 87.4%
Enkrypt · Jailbreak risk 9.6%Enkrypt · Harmful content risk 0%Enkrypt · CBRN risk 22.2%Enkrypt · Toxicity risk 6.2%Enkrypt · Bias risk 78%Enkrypt · Insecure code risk 8%
GDPVal-AA 1569tau3-Banking 31.1%Toolathlon Verified 67.5%Automation Bench Public 33.5%Apex-Agents (pass@1) 28.6%MCPMark 66.9%BrowseComp 83.3%WildClawBench 50.4%Terminal-Bench 2.1 80.9%SciCode 52.5%SWE Bench Pro 48.8%Humanity's Last Exam (without tools) 39.5%GPQA Diamond 91.1%CritPt 21%AA-LCR 78.3%AA-Omniscience Accuracy 43%AA-Omniscience Non-Hallucination 7%
Badge
[](https://publicai.io/model-index/m/gpt-5-6-luna)