‹ PublicAI Index
The LLM benchmark aggregator.
GPT-6 Astra
OpenAI
Strongest in Data analysis (#1 of 57), weakest in Safety (#158 of 337). Above par in 23 of 24 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5 and Claude Opus 5.5 and ahead of Claude Opus 4.8 and GLM-5.3.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Core abilities67.1−0.3#2/2042/3
Data analysis66.1leads#1/571/1
Language69−2.5#3/571/1
General intelligence68.5−1.6#4/2042/3
Instruction following63.5−10.2#7/571/1
Reasoning66.4−0.5#2/1784/4
Reasoning68.6−0.5#2/1403/3
Mathematics65.3−0.4#3/1402/2
Coding62.7−6.4#5/1653/5
Agentic coding64.7−4#5/1573/4
Code generation56.1−13.9#22/771/2
Agents63.4−4.7#12/2682/5
Knowledge work67.5−6.1#12/1782/2
Human preference64.5−3#23/3421/1
Human preference64.5−3#23/3421/1
Professional56−5.6#28/1681/1
IT operations73.8leads#1/111/1
Biology research68.5leads#2/161/1
Cybersecurity27.4−32.3#9/91/1
Medical58.6−8.2#16/1401/1
Legal53.5−13.6#49/1511/1
Finance54.5−8.4#52/1511/1
Safety51−10.5#158/3371/3
Factual grounding53.2−17.4#43/1011/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for GPT-6 Astra, left for the other.
§ 3 · Sources
Where the numbers come from
10 publications, 27 figures. Every one links to the page it was read from.
LMArena Text 1478
Artificial Analysis Intelligence Index 53
GDPval-AA 1542
AA-Briefcase 1569
Terminal-Bench 58.2% (Codex)
ARC-AGI-2 95%
LiveBench 82.2LiveBench · Reasoning 92.7LiveBench · Coding 80.4LiveBench · Agentic Coding 57.3LiveBench · Mathematics 96.8LiveBench · Data Analysis 83LiveBench · Language 89.4LiveBench · Instruction Following 75.6
SimpleBench 83.6%
Vals · Legal Research Bench 39.42%Vals · Harvey Legal Agent Benchmark 5.42%Vals · Finance Agent 53.54%Vals · MedCode 48.49%Vals · MedScribe 87.91%Vals · BioMysteryBench 79.26%Vals · CyberBench 41.07%Vals · SRE Bench 56.87%Vals · Vibe Code Bench 89.59% (OpenHands)Vals · Code Migration 67.74%Vals · ProofBench 99%
Vectara · Factual consistency 91.3%
Badge
[](https://publicai.io/model-index/m/gpt-6-astra)