‹ PublicAI Index
The LLM benchmark aggregator.
GPT-6 Luna
OpenAI
Strongest in Routing & classification (#1 of 85), weakest in Toxicity avoidance (#210 of 272). Above par in 19 of 29 scopes. Among the models it meets almost everywhere, it finishes behind Claude Opus 5.5 and Claude Opus 5 and ahead of Gemini 3.5 Flash and GLM-5.1.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Decisions72.7−0.6#2/851/1
Routing & classification74leads#1/851/1
Calibration71.3−1.4#2/851/1
Coding55.5−13.6#30/1652/5
Agentic coding56.3−12.4#32/1572/4
Code generation53.3−16.7#34/771/2
Agents59.5−8.6#35/2682/5
Knowledge work62.4−11.2#29/1782/2
Human preference61−6.5#74/3421/1
Human preference61−6.5#74/3421/1
Reasoning51.3−15.6#80/1783/4
Reasoning50.8−18.3#58/1402/3
Mathematics51.8−13.9#72/1402/2
Professional49.7−11.9#82/1681/1
Biology research36.5−32#14/161/1
Medical54.3−12.5#48/1401/1
Finance52.1−10.8#76/1511/1
Legal48.5−18.6#90/1511/1
Safety53.4−8.1#88/3371/3
Secure code63.9−2.6#21/2741/1
Fairness60.2−9.8#42/3001/2
Harm refusal49.8−11.8#159/3001/2
Jailbreak resistance50−14.9#181/2721/1
Toxicity avoidance48.9−10.1#210/2721/1
Core abilities44.5−22.9#154/2042/3
Data analysis50−16.1#35/571/1
Language39.7−31.8#46/571/1
Instruction following28.9−44.8#55/571/1
General intelligence52−18.1#80/2042/3
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for GPT-6 Luna, left for the other.
§ 3 · Sources
Where the numbers come from
9 publications, 30 figures. Every one links to the page it was read from.
LMArena Text 1442
Artificial Analysis Intelligence Index 37
GDPval-AA 1367
AA-Briefcase 1299
ARC-AGI-2 59.3%
LiveBench 72LiveBench · Reasoning 81.8LiveBench · Coding 79LiveBench · Agentic Coding 51.2LiveBench · Mathematics 89.1LiveBench · Data Analysis 73.4LiveBench · Language 73.8LiveBench · Instruction Following 55.9
Vals · Legal Research Bench 30.29%Vals · Harvey Legal Agent Benchmark 2.92%Vals · Finance Agent 49.87%Vals · MedCode 44.69%Vals · MedScribe 83.71%Vals · BioMysteryBench 61.48%Vals · Vibe Code Bench 81.65% (OpenHands)Vals · Code Migration 42.55%Vals · ProofBench 64%
JevBench · Intelligence 97.4%JevBench · Calibration 93.5%
Enkrypt · Jailbreak risk 12.8%Enkrypt · Harmful content risk 0%Enkrypt · CBRN risk 33%Enkrypt · Toxicity risk 7.1%Enkrypt · Bias risk 62%Enkrypt · Insecure code risk 5.3%
Badge
[](https://publicai.io/model-index/m/gpt-6-luna)