‹ PublicAI IndexGoogle
The LLM benchmark aggregator.
Gemini 3 Flash
Strongest in Academic knowledge (#19 of 123), weakest in Safety (#249 of 337). Above par in 11 of 16 scopes. Among the models it meets almost everywhere, it finishes behind Claude Opus 4.7 and Claude Opus 5 and ahead of GLM-5.2 and GPT-5.6 Luna.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Knowledge57.7−5.8#22/1381/2
Academic knowledge59.2−4.7#19/1231/1
Reasoning56−10.9#28/1783/4
Mathematics57.8−7.9#20/1401/2
Science58.2−5.3#37/1221/1
Reasoning53.4−15.7#48/1402/3
Human preference64−3.5#33/3421/1
Human preference64−3.5#33/3421/1
Professional52.3−9.3#58/1681/1
Medical56.5−10.3#27/1401/1
Finance55−7.9#46/1511/1
Legal46.7−20.4#95/1511/1
Coding44.3−24.8#120/1651/5
Agentic coding43.7−25#116/1571/4
Safety47.2−14.3#249/3371/3
Factual grounding41.1−29.5#85/1011/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Gemini 3 Flash, left for the other.
§ 3 · Sources
Where the numbers come from
5 publications, 21 figures. Every one links to the page it was read from.
LMArena Text 1473
ARC-AGI-2 33.6%
SimpleBench 61.1%
Vals · Legal Research Bench 18.27%Vals · CaseLaw 55.84%Vals · LegalBench 86.86%Vals · Harvey Legal Agent Benchmark 0%Vals · Finance Agent 42.55%Vals · CorpFin 66.43%Vals · TaxEval 73.88%Vals · MortgageTax 68.72%Vals · MedQA 95.81%Vals · MedCode 55.92%Vals · MedScribe 69.92%Vals · SWE-bench Verified 75% (Mini-SWE-agent)Vals · Vibe Code Bench 20.2% (OpenHands)Vals · Code Migration 6.39%Vals · GPQA Diamond 87.88%Vals · MMLU Pro 88.59%Vals · AIME 95.63%
Vectara · Factual consistency 86.5%
Badge
[](https://publicai.io/model-index/m/gemini-3-flash)