‹ PublicAI IndexGoogle
The LLM benchmark aggregator.
Gemini 2.5 Flash
Strongest in Safe-prompt compliance (#1 of 82), weakest in Jailbreak resistance (#243 of 272). Above par in 21 of 27 scopes. Among the models it meets almost everywhere, it finishes behind Claude Opus 5 and Claude Opus 5.5 and ahead of O4 Mini and GPT-4.1.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Agents57.4−10.7#48/2681/5
Tool use63.3−10.7#12/811/1
Coding52.9−16.2#51/1652/5
Code generation53.1−16.9#35/771/2
Agentic coding52.7−16#51/1571/4
Professional51.9−9.7#61/1681/1
Legal53.4−13.7#50/1511/1
Medical51.3−15.5#70/1401/1
Finance52.2−10.7#75/1511/1
Knowledge52.5−11#74/1381/2
Academic knowledge53−10.9#68/1231/1
Core abilities51.4−16#78/2041/3
General intelligence52−18.1#81/2041/3
Human preference57.9−9.6#115/3421/1
Human preference57.9−9.6#115/3421/1
Reasoning45.9−21#121/1783/4
Science50.3−13.2#75/1221/1
Reasoning43.3−25.8#91/1402/3
Mathematics46.1−19.6#93/1401/2
Safety50.7−10.8#162/3373/3
Safe-prompt compliance60.4leads#1/821/1
Factual grounding55.5−15.1#38/1011/1
Fairness55.7−14.3#63/3002/2
Secure code58.9−7.6#75/2741/1
Harm refusal48.5−13.1#181/3002/2
Toxicity avoidance46.9−12.1#221/2721/1
Jailbreak resistance33−31.9#243/2721/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Gemini 2.5 Flash, left for the other.
§ 3 · Sources
Where the numbers come from
11 publications, 29 figures. Every one links to the page it was read from.
LMArena Text 1409
ARC-AGI-2 2.5%
LiveCodeBench 61.9%
Aider polyglot 55.1%
BFCL v4 56.24%
Kagi LLM Benchmark 56.8%
SimpleBench 41.2%
Vals · LegalBench 82.63%Vals · CorpFin 59.75%Vals · TaxEval 72.4%Vals · MortgageTax 61.92%Vals · MedQA 91.17%Vals · MedCode 40.33%Vals · MedScribe 78.5%Vals · GPQA Diamond 76.52%Vals · MMLU Pro 83.66%Vals · AIME 51.46%
HELM Safety · HarmBench 62.6%HELM Safety · SimpleSafetyTests 98%HELM Safety · Anthropic Red Team 98.8%HELM Safety · BBQ 97.7%HELM Safety · XSTest 98.8%
Enkrypt · Jailbreak risk 27.1%Enkrypt · Harmful content risk 45.4%Enkrypt · CBRN risk 15.8%Enkrypt · Toxicity risk 8.5%Enkrypt · Bias risk 75.2%Enkrypt · Insecure code risk 15.4%
Vectara · Factual consistency 92.2%
Badge
[](https://publicai.io/model-index/m/gemini-2-5-flash)