‹ PublicAI IndexGoogle
The LLM benchmark aggregator.
Gemini 3.1 Flash Lite
Strongest in Routing & classification (#7 of 85), weakest in Agents (#223 of 268). Above par in 13 of 22 scopes. Among the models it meets almost everywhere, it finishes behind Claude Opus 5 and Claude Fable 5 and ahead of DeepSeek V3.2 and GPT OSS 120B.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Decisions54.1−19.2#25/851/1
Routing & classification59.8−14.2#7/851/1
Calibration48.3−24.4#55/851/1
Knowledge55.2−8.3#48/1381/2
Academic knowledge56.3−7.6#44/1231/1
Core abilities54.9−12.5#48/2041/3
General intelligence57.1−13#45/2041/3
Reasoning53.4−13.5#58/1781/4
Mathematics54.5−11.2#58/1401/2
Science53.5−10#65/1221/1
Human preference60.2−7.3#85/3421/1
Human preference60.2−7.3#85/3421/1
Professional46−15.6#122/1681/1
Finance48.5−14.4#98/1511/1
Medical46.7−20.1#101/1401/1
Legal42.5−24.6#120/1511/1
Safety51.4−10.1#146/3371/3
Factual grounding54.5−16.1#39/1011/1
Coding39.4−29.7#157/1651/5
Agentic coding38.2−30.5#151/1571/4
Agents41.6−26.5#223/2682/5
Knowledge work39.1−34.5#137/1782/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Gemini 3.1 Flash Lite, left for the other.
§ 3 · Sources
Where the numbers come from
7 publications, 23 figures. Every one links to the page it was read from.
LMArena Text 1433
GDPval-AA 416
AA-Briefcase 207
Kagi LLM Benchmark 67.2%
Vals · Legal Research Bench 3.37%Vals · CaseLaw 54.98%Vals · LegalBench 83.76%Vals · Harvey Legal Agent Benchmark 0%Vals · Finance Agent 29.99%Vals · CorpFin 59.36%Vals · TaxEval 71.79%Vals · MortgageTax 68.04%Vals · MedCode 47.6%Vals · MedScribe 63.9%Vals · SWE-bench Verified 62.8% (Mini-SWE-agent)Vals · Vibe Code Bench 0% (OpenHands)Vals · Code Migration 4.61%Vals · GPQA Diamond 81.06%Vals · MMLU Pro 86.24%Vals · AIME 83.33%
JevBench · Intelligence 54.5%JevBench · Calibration 59.3%
Vectara · Factual consistency 91.8%
Badge
[](https://publicai.io/model-index/m/gemini-3-1-flash-lite)