‹ PublicAI IndexGoogle
The LLM benchmark aggregator.
Gemini 3.1 Pro
Strongest in Science (#1 of 122), weakest in Safety (#190 of 337). Above par in 17 of 24 scopes. Among the models it meets almost everywhere, it finishes behind Claude Opus 5.5 and Claude Fable 5 and ahead of Muse Spark 1.1 and DeepSeek V4.1 Flash.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Knowledge60.2−3.3#5/1381/2
Academic knowledge62.2−1.7#4/1231/1
Core abilities59.5−7.9#10/2042/3
Instruction following69.7−4#3/571/1
Language61.5−10#12/571/1
Data analysis58.6−7.5#16/571/1
General intelligence53.9−16.2#68/2042/3
Human preference65.4−2.1#14/3421/1
Human preference65.4−2.1#14/3421/1
Reasoning56.9−10#26/1784/4
Science63.5leads#1/1221/1
Reasoning60.2−8.9#18/1403/3
Mathematics51.7−14#73/1402/2
Professional54.5−7.1#38/1681/1
Medical60.2−6.6#7/1401/1
Finance54.4−8.5#55/1511/1
Legal51.1−16#70/1511/1
Coding46.4−22.7#107/1652/5
Code generation48.2−21.8#51/771/2
Agentic coding45.8−22.9#100/1572/4
Agents46.8−21.3#155/2682/5
Knowledge work45.8−27.8#102/1782/2
Safety49.7−11.8#190/3371/3
Factual grounding48.9−21.7#59/1011/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Gemini 3.1 Pro, left for the other.
§ 3 · Sources
Where the numbers come from
9 publications, 33 figures. Every one links to the page it was read from.
LMArena Text 1487
Artificial Analysis Intelligence Index 30
GDPval-AA 776
AA-Briefcase 453
ARC-AGI-2 77.1%
LiveBench 77LiveBench · Reasoning 84LiveBench · Coding 76.5LiveBench · Agentic Coding 44.1LiveBench · Mathematics 91LiveBench · Data Analysis 78.5LiveBench · Language 85.4LiveBench · Instruction Following 79.1
SimpleBench 79.6%
Vals · Legal Research Bench 20.67%Vals · CaseLaw 64.84%Vals · LegalBench 87.4%Vals · Harvey Legal Agent Benchmark 0%Vals · Finance Agent 42.98%Vals · CorpFin 64.49%Vals · TaxEval 72.88%Vals · MortgageTax 69.4%Vals · MedQA 96.37%Vals · MedCode 59.06%Vals · MedScribe 76.11%Vals · SWE-bench Verified 78.8% (Mini-SWE-agent)Vals · Vibe Code Bench 32.03% (OpenHands)Vals · Code Migration 17.31%Vals · GPQA Diamond 95.45%Vals · MMLU Pro 90.99%Vals · AIME 98.13%Vals · ProofBench 26%
Vectara · Factual consistency 89.6%
Badge
[](https://publicai.io/model-index/m/gemini-3-1-pro)