‹ PublicAI IndexGoogle
The LLM benchmark aggregator.
Gemini 3 Pro
Strongest in Tool use (#3 of 81), weakest in Safety (#155 of 337). Above par in 20 of 23 scopes. Among the models it meets almost everywhere, it finishes behind Claude Opus 5.5 and Claude Fable 5 and ahead of Qwen3.8 Max and GPT-5.6 Terra.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Knowledge59.2−4.3#8/1381/2
Academic knowledge61.1−2.8#7/1231/1
Core abilities59.3−8.1#11/2041/3
General intelligence63.4−6.7#8/2041/3
Reasoning59.9−7#12/1783/4
Reasoning60.9−8.2#15/1402/3
Mathematics58.1−7.6#17/1401/2
Science60.9−2.6#18/1221/1
Agents63.3−4.8#14/2681/5
Tool use74leads#3/811/1
Human preference65.2−2.3#15/3421/1
Human preference65.2−2.3#15/3421/1
Professional54.1−7.5#41/1681/1
Medical55.6−11.2#32/1401/1
Finance56.2−6.7#33/1511/1
Legal50.3−16.8#75/1511/1
Coding47.2−21.9#95/1651/5
Agentic coding46.8−21.9#89/1571/4
Safety51.2−10.3#155/3372/3
Safe-prompt compliance57.1−3.3#20/821/1
Fairness58.4−11.6#44/3001/2
Factual grounding40.9−29.7#86/1011/1
Harm refusal50.2−11.4#154/3001/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Gemini 3 Pro, left for the other.
§ 3 · Sources
Where the numbers come from
8 publications, 24 figures. Every one links to the page it was read from.
LMArena Text 1485
ARC-AGI-2 54%
BFCL v4 72.51%
Kagi LLM Benchmark 80.1%
SimpleBench 76.4%
Vals · CaseLaw 53.05%Vals · LegalBench 87.03%Vals · CorpFin 63.67%Vals · TaxEval 72.57%Vals · MortgageTax 69.08%Vals · MedQA 96.03%Vals · MedCode 52.2%Vals · MedScribe 72.04%Vals · SWE-bench Verified 76.4% (Mini-SWE-agent)Vals · Vibe Code Bench 14.3% (OpenHands)Vals · GPQA Diamond 91.67%Vals · MMLU Pro 90.1%Vals · AIME 96.68%
HELM Safety · HarmBench 72.5%HELM Safety · SimpleSafetyTests 97.5%HELM Safety · Anthropic Red Team 97.1%HELM Safety · BBQ 98.4%HELM Safety · XSTest 97.3%
Vectara · Factual consistency 86.4%
Badge
[](https://publicai.io/model-index/m/gemini-3-pro)