‹ PublicAI IndexGoogle
The LLM benchmark aggregator.
Gemini 1.5 Flash
Strongest in Safe-prompt compliance (#68 of 82), weakest in Toxicity avoidance (#229 of 272). Above par in 4 of 17 scopes. Among the models it meets almost everywhere, it finishes behind GPT-5.1 and GPT-5 and ahead of Mistral Small and Jamba 1.5 Mini.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Knowledge33.7−29.8#131/1381/2
Academic knowledge30.5−33.4#115/1231/1
Safety50.8−10.7#161/3372/3
Safe-prompt compliance45.5−14.9#68/821/1
Secure code58.8−7.7#77/2741/1
Fairness49.2−20.8#126/3002/2
Harm refusal51.6−10#132/3002/2
Jailbreak resistance52.9−12#144/2721/1
Toxicity avoidance45.6−13.4#229/2721/1
Reasoning36.6−30.3#166/1781/4
Science29−34.5#113/1221/1
Mathematics37−28.7#119/1401/2
Professional35.9−25.7#168/1681/1
Legal38.6−28.5#135/1511/1
Finance29.2−33.7#151/1511/1
Human preference48.2−19.3#212/3421/1
Human preference48.2−19.3#212/3421/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Gemini 1.5 Flash, left for the other.
§ 3 · Sources
Where the numbers come from
4 publications, 19 figures. Every one links to the page it was read from.
LMArena Text 1309
Vals · LegalBench 63.24%Vals · CorpFin 38.19%Vals · TaxEval 48.2%Vals · MortgageTax 42.77%Vals · GPQA Diamond 45.96%Vals · MMLU Pro 65.61%Vals · AIME 17.29%
HELM Safety · HarmBench 80%HELM Safety · SimpleSafetyTests 97%HELM Safety · Anthropic Red Team 99.9%HELM Safety · BBQ 94.7%HELM Safety · XSTest 92.1%
Enkrypt · Jailbreak risk 10.3%Enkrypt · Harmful content risk 44.4%Enkrypt · CBRN risk 13.3%Enkrypt · Toxicity risk 9.4%Enkrypt · Bias risk 87.9%Enkrypt · Insecure code risk 15.6%
Badge
[](https://publicai.io/model-index/m/gemini-1-5-flash)