‹ PublicAI IndexGoogle
The LLM benchmark aggregator.
Gemini 2.5 Flash Lite
Strongest in Factual grounding (#3 of 101), weakest in Harm refusal (#156 of 300). Above par in 7 of 20 scopes. Among the models it meets almost everywhere, it finishes behind Claude Sonnet 5 and Muse Spark 1.1 and ahead of Ling 3.0 Flash and O3 Mini.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Safety54.4−7.1#64/3372/3
Factual grounding66.8−3.8#3/1011/1
Safe-prompt compliance58.2−2.2#14/821/1
Fairness55.3−14.7#65/3001/2
Harm refusal50−11.6#156/3001/2
Knowledge47.8−15.7#100/1381/2
Academic knowledge47.4−16.5#92/1231/1
Professional47.2−14.4#111/1681/1
Legal52.9−14.2#55/1511/1
Finance46.9−16#103/1511/1
Medical43.9−22.9#116/1401/1
Agents49.6−18.5#125/2681/5
Tool use49.3−24.7#38/811/1
Reasoning45.5−21.4#125/1781/4
Science45.9−17.6#89/1221/1
Mathematics43.6−22.1#103/1401/2
Core abilities45.8−21.6#139/2041/3
General intelligence44−26.1#142/2041/3
Human preference54.6−12.9#145/3421/1
Human preference54.6−12.9#145/3421/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Gemini 2.5 Flash Lite, left for the other.
§ 3 · Sources
Where the numbers come from
6 publications, 19 figures. Every one links to the page it was read from.
LMArena Text 1375
BFCL v4 36.87%
Kagi LLM Benchmark 40.5%
Vals · LegalBench 82.04%Vals · CorpFin 57.58%Vals · TaxEval 64.72%Vals · MortgageTax 57.55%Vals · MedQA 88.87%Vals · MedCode 34.19%Vals · MedScribe 66.88%Vals · GPQA Diamond 70.2%Vals · MMLU Pro 79.12%Vals · AIME 42.08%
HELM Safety · HarmBench 67%HELM Safety · SimpleSafetyTests 96.5%HELM Safety · Anthropic Red Team 98.7%HELM Safety · BBQ 94.9%HELM Safety · XSTest 97.8%
Vectara · Factual consistency 96.7%
Badge
[](https://publicai.io/model-index/m/gemini-2-5-flash-lite)