‹ PublicAI IndexGoogle
The LLM benchmark aggregator.
Gemini 3.5 Flash Lite
Strongest in Instruction following (#32 of 57), weakest in Core abilities (#202 of 204). Above par in 8 of 22 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5 and Claude Fable 5.1 and ahead of Claude Sonnet 4 and DeepSeek V3.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Human preference62.4−5.1#56/3421/1
Human preference62.4−5.1#56/3421/1
Knowledge54.8−8.7#57/1381/2
Academic knowledge55.8−8.1#53/1231/1
Professional49.6−12#84/1681/1
Finance54.1−8.8#59/1511/1
Medical47.5−19.3#94/1401/1
Legal45.4−21.7#100/1511/1
Agents50.2−17.9#115/2682/5
Knowledge work50.2−23.4#81/1782/2
Coding44.9−24.2#117/1652/5
Code generation47.3−22.7#53/771/2
Agentic coding44.2−24.5#111/1572/4
Reasoning38.6−28.3#159/1783/4
Science55.4−8.1#56/1221/1
Reasoning33.8−35.3#136/1402/3
Mathematics32.9−32.8#139/1401/2
Core abilities37.7−29.7#202/2042/3
Instruction following48.8−24.9#32/571/1
Language35.9−35.6#52/571/1
Data analysis26−40.1#57/571/1
General intelligence38.9−31.2#180/2042/3
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Gemini 3.5 Flash Lite, left for the other.
§ 3 · Sources
Where the numbers come from
7 publications, 27 figures. Every one links to the page it was read from.
LMArena Text 1456
Artificial Analysis Intelligence Index 22
GDPval-AA 970
AA-Briefcase 648
ARC-AGI-2 10.3%
LiveBench 63.9LiveBench · Reasoning 60.2LiveBench · Coding 76.1LiveBench · Agentic Coding 45.3LiveBench · Mathematics 73.7LiveBench · Data Analysis 53.2LiveBench · Language 71.8LiveBench · Instruction Following 67.2
Vals · Legal Research Bench 13.94%Vals · LegalBench 84.07%Vals · Harvey Legal Agent Benchmark 0%Vals · Finance Agent 47.44%Vals · CorpFin 60.68%Vals · TaxEval 72.61%Vals · MortgageTax 68.68%Vals · MedCode 43.49%Vals · MedScribe 70.89%Vals · SWE-bench Verified 75% (Mini-SWE-agent)Vals · Vibe Code Bench 37.16% (OpenHands)Vals · Code Migration 6.15%Vals · GPQA Diamond 83.84%Vals · MMLU Pro 85.84%
Badge
[](https://publicai.io/model-index/m/gemini-3-5-flash-lite)