‹ PublicAI IndexGoogle
The LLM benchmark aggregator.
Gemini 3.5 Flash
Strongest in Web research (#3 of 4), weakest in Harm refusal (#262 of 300). Above par in 21 of 29 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5 and Claude Opus 5 and ahead of GPT-5.6 Luna and Kimi K2.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Knowledge58.6−4.9#12/1381/2
Academic knowledge60.4−3.5#10/1231/1
Human preference64.4−3.1#24/3421/1
Human preference64.4−3.1#24/3421/1
Professional55.4−6.2#31/1681/1
Finance58.6−4.3#12/1511/1
Medical57.7−9.1#22/1401/1
Legal50.8−16.3#73/1511/1
Reasoning53.7−13.2#52/1784/4
Science61.6−1.9#15/1221/1
Reasoning57.8−11.3#30/1403/3
Mathematics45.2−20.5#96/1402/2
Core abilities52.1−15.3#70/2041/3
Instruction following63.5−10.2#8/571/1
Language60−11.5#14/571/1
Data analysis35.7−30.4#51/571/1
General intelligence49.7−20.4#102/2041/3
Coding50.5−18.6#71/1652/5
Code generation51.6−18.4#40/771/2
Agentic coding50.2−18.5#72/1572/4
Agents52−16.1#97/2683/5
Web research44−20.2#3/41/1
Knowledge work55.2−18.4#51/1782/2
Safety46.7−14.8#261/3371/3
Secure code60.8−5.7#59/2741/1
Toxicity avoidance55.1−3.9#104/2721/1
Fairness48.3−21.7#137/3001/2
Jailbreak resistance28.9−36#255/2721/1
Harm refusal43.7−17.9#262/3001/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Gemini 3.5 Flash, left for the other.
§ 3 · Sources
Where the numbers come from
8 publications, 35 figures. Every one links to the page it was read from.
LMArena Text 1477
GDPval-AA 1185
AA-Briefcase 872
ARC-AGI-2 72.1%
LiveBench 74.6LiveBench · Reasoning 82LiveBench · Coding 78.2LiveBench · Agentic Coding 49LiveBench · Mathematics 88.2LiveBench · Data Analysis 64.9LiveBench · Language 84.6LiveBench · Instruction Following 75.6
SimpleBench 76.7%
Vals · Legal Research Bench 30.77%Vals · LegalBench 83.6%Vals · Harvey Legal Agent Benchmark 2.5%Vals · Finance Agent 57.86%Vals · CorpFin 64.69%Vals · TaxEval 74.37%Vals · MortgageTax 68.12%Vals · MedCode 55.83%Vals · MedScribe 76.57%Vals · Web Search Index 41.36% (Exa)Vals · SWE-bench Verified 78.8% (Mini-SWE-agent)Vals · Vibe Code Bench 48.68% (OpenHands)Vals · Code Migration 26.75%Vals · GPQA Diamond 92.68%Vals · MMLU Pro 89.52%Vals · ProofBench 31%
Enkrypt · Jailbreak risk 30.6%Enkrypt · Harmful content risk 15%Enkrypt · CBRN risk 52%Enkrypt · Toxicity risk 2.7%Enkrypt · Bias risk 80.4%Enkrypt · Insecure code risk 11.6%
Badge
[](https://publicai.io/model-index/m/gemini-3-5-flash)