‹ PublicAI IndexGoogle
The LLM benchmark aggregator.
Gemini 2.0 Flash
Strongest in Safe-prompt compliance (#48 of 82), weakest in Toxicity avoidance (#201 of 272). Above par in 8 of 23 scopes. Among the models it meets almost everywhere, it finishes behind GPT-5 and O3 and ahead of GPT-5.4 Nano and Gemini 3.5 Flash Lite.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Knowledge46−17.5#104/1381/2
Academic knowledge45.2−18.7#96/1231/1
Coding45.2−23.9#113/1651/5
Agentic coding44.2−24.5#110/1571/4
Professional45.5−16.1#126/1681/1
Legal50.1−17#76/1511/1
Medical48.9−17.9#88/1401/1
Finance40.7−22.2#126/1511/1
Safety51.5−10#145/3372/3
Safe-prompt compliance52.6−7.8#48/821/1
Jailbreak resistance57.2−7.7#79/2721/1
Fairness50.2−19.8#112/3002/2
Harm refusal52.2−9.4#123/3002/2
Secure code45.9−20.6#189/2741/1
Toxicity avoidance49.9−9.1#201/2721/1
Core abilities44.9−22.5#147/2041/3
General intelligence42.7−27.4#152/2041/3
Reasoning38.9−28#155/1783/4
Science42.4−21.1#95/1221/1
Mathematics40.3−25.4#110/1401/2
Reasoning35.8−33.3#131/1402/3
Human preference53.2−14.3#156/3421/1
Human preference53.2−14.3#156/3421/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Gemini 2.0 Flash, left for the other.
§ 3 · Sources
Where the numbers come from
8 publications, 24 figures. Every one links to the page it was read from.
LMArena Text 1360
ARC-AGI-2 1.3%
Aider polyglot 22.2%
Kagi LLM Benchmark 37.8%
SimpleBench 18.9%
Vals · LegalBench 78.36%Vals · CorpFin 33.72%Vals · TaxEval 65.25%Vals · MortgageTax 59.66%Vals · MedQA 81.47%Vals · GPQA Diamond 65.15%Vals · MMLU Pro 77.38%Vals · AIME 29.79%
HELM Safety · HarmBench 66.2%HELM Safety · SimpleSafetyTests 98.5%HELM Safety · Anthropic Red Team 99.4%HELM Safety · BBQ 95.4%HELM Safety · XSTest 95.3%
Enkrypt · Jailbreak risk 6.7%Enkrypt · Harmful content risk 36.1%Enkrypt · CBRN risk 6.7%Enkrypt · Toxicity risk 6.4%Enkrypt · Bias risk 86.3%Enkrypt · Insecure code risk 41.8%
Badge
[](https://publicai.io/model-index/m/gemini-2-0-flash)