‹ PublicAI IndexGoogle
The LLM benchmark aggregator.
Gemini 3.6 Flash
Strongest in Science (#8 of 122), weakest in Safety (#280 of 337). Above par in 21 of 29 scopes. Among the models it meets almost everywhere, it finishes behind Gemini 3.7 Flash and Claude Opus 5 and ahead of Claude Sonnet 4 and Kimi K2.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Knowledge58.4−5.1#15/1381/2
Academic knowledge60.1−3.8#13/1231/1
Human preference64.9−2.6#17/3421/1
Human preference64.9−2.6#17/3421/1
Professional53.1−8.5#50/1681/1
Biology research30.8−37.7#16/161/1
Finance57.8−5.1#21/1511/1
Medical57.5−9.3#23/1401/1
Legal51.3−15.8#68/1511/1
Reasoning53.3−13.6#59/1783/4
Science62.1−1.4#8/1221/1
Reasoning54.3−14.8#45/1402/3
Mathematics45.9−19.8#94/1401/2
Agents55.4−12.7#61/2682/5
Knowledge work57−16.6#42/1782/2
Coding50.7−18.4#70/1652/5
Code generation51−19#42/771/2
Agentic coding50.6−18.1#68/1572/4
Core abilities50.5−16.9#89/2041/3
Instruction following63.2−10.5#9/571/1
Language58.7−12.8#15/571/1
Data analysis32.5−33.6#52/571/1
General intelligence48.1−22#115/2041/3
Safety45.7−15.8#280/3371/3
Secure code52.9−13.6#130/2741/1
Fairness46.2−23.8#164/3001/2
Toxicity avoidance52.4−6.6#165/2721/1
Harm refusal45−16.6#247/3001/2
Jailbreak resistance31.4−33.5#251/2721/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Gemini 3.6 Flash, left for the other.
§ 3 · Sources
Where the numbers come from
7 publications, 33 figures. Every one links to the page it was read from.
LMArena Text 1482
GDPval-AA 1265
AA-Briefcase 950
ARC-AGI-2 60.4%
LiveBench 73.6LiveBench · Reasoning 85.1LiveBench · Coding 77.9LiveBench · Agentic Coding 43.4LiveBench · Mathematics 86.4LiveBench · Data Analysis 63LiveBench · Language 83.9LiveBench · Instruction Following 75.4
Vals · Legal Research Bench 25%Vals · LegalBench 86.7%Vals · Harvey Legal Agent Benchmark 3.33%Vals · Finance Agent 56.3%Vals · CorpFin 63.33%Vals · TaxEval 74.86%Vals · MortgageTax 67.61%Vals · MedCode 53.15%Vals · MedScribe 79.66%Vals · BioMysteryBench 58.52%Vals · SWE-bench Verified 79.6% (Mini-SWE-agent)Vals · Vibe Code Bench 64% (OpenHands)Vals · Code Migration 30.93%Vals · GPQA Diamond 93.43%Vals · MMLU Pro 89.28%
Enkrypt · Jailbreak risk 28.5%Enkrypt · Harmful content risk 8.9%Enkrypt · CBRN risk 52.8%Enkrypt · Toxicity risk 4.6%Enkrypt · Bias risk 83.7%Enkrypt · Insecure code risk 27.6%
Badge
[](https://publicai.io/model-index/m/gemini-3-6-flash)