‹ PublicAI IndexGoogle
The LLM benchmark aggregator.
Gemini 4 Argon
Strongest in Human preference (#1 of 345), weakest in Reasoning (#22 of 183). Above par in 17 of 17 scopes. Among the models it meets almost everywhere, it finishes behind Claude Opus 5.5 and Claude Opus 5 and ahead of GPT-6.1 Sol and GPT-6 Astra.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Human preference68.8leads#1/3451/1
Human preference68.8leads#1/3451/1
Professional64.1leads#1/1771/1
Legal66.3−0.4#2/1571/1
Medical64.1−2.5#2/1461/1
Cybersecurity62.4−0.1#2/431/1
Finance61.8−0.8#3/1571/1
IT operations65.9−7.8#3/161/1
Biology research61.2−8.4#6/231/1
Coding63−5.5#5/1681/5
Agentic coding65−3#4/1611/4
Core abilities61.3−5.8#8/2121/3
General intelligence66.3−3.6#7/2121/3
Agents62.6−4.9#18/2772/5
Knowledge work66.4−6.3#15/1872/2
Reasoning57.8−8.4#22/1831/4
Mathematics61.9−3#9/1451/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Gemini 4 Argon, left for the other.
§ 3 · Sources
Where the numbers come from
5 publications, 16 figures. Every one links to the page it was read from.
LMArena Text 1525
Artificial Analysis Intelligence Index 53
GDPval-AA 1627
AA-Briefcase 1490
Vals · Legal Research Bench 54.81%Vals · LegalBench 88.3%Vals · Harvey Legal Agent Benchmark 19.58%Vals · Finance Agent 65.4%Vals · MedCode 58.8%Vals · MedScribe 87.43%Vals · BioMysteryBench 76.3%Vals · CyberBench 77.86%Vals · SRE Bench 44.27%Vals · Vibe Code Bench 91.91% (OpenHands)Vals · Code Migration 68.17%Vals · ProofBench 99%
Badge
[](https://publicai.io/model-index/m/gemini-4-argon)