‹ PublicAI IndexGoogle
The LLM benchmark aggregator.
Gemini 3.7 Flash
Strongest in Instruction following (#2 of 57), weakest in General intelligence (#53 of 204). Above par in 21 of 23 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5.1 and Claude Fable 5 and ahead of Muse Spark 1.1 and GLM-5.3.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Knowledge59.3−4.2#7/1381/2
Academic knowledge61.1−2.8#6/1231/1
Human preference65.5−2#13/3421/1
Human preference65.5−2#13/3421/1
Reasoning58.2−8.7#20/1783/4
Science62.5−1#5/1221/1
Reasoning59.8−9.3#22/1402/3
Mathematics54.9−10.8#52/1402/2
Core abilities57.5−9.9#23/2041/3
Instruction following71.1−2.6#2/571/1
Language61.7−9.8#11/571/1
Data analysis40.9−25.2#47/571/1
General intelligence56.4−13.7#53/2041/3
Professional56.4−5.2#23/1681/1
IT operations42.7−31.1#6/111/1
Medical59.7−7.1#11/1401/1
Finance58.4−4.5#14/1511/1
Legal57.2−9.9#25/1511/1
Coding53.8−15.3#39/1653/5
Code generation53.1−16.9#36/771/2
Agentic coding54−14.7#47/1573/4
Agents57.7−10.4#47/2682/5
Knowledge work60−13.6#36/1782/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Gemini 3.7 Flash, left for the other.
§ 3 · Sources
Where the numbers come from
7 publications, 28 figures. Every one links to the page it was read from.
LMArena Text 1488
GDPval-AA 1371
AA-Briefcase 1107
Terminal-Bench 11.2% (mini-SWE-agent)
ARC-AGI-2 84.6%
LiveBench 78.8LiveBench · Reasoning 87.8LiveBench · Coding 78.9LiveBench · Agentic Coding 58.3LiveBench · Mathematics 93.5LiveBench · Data Analysis 68LiveBench · Language 85.5LiveBench · Instruction Following 79.9
Vals · Legal Research Bench 34.62%Vals · LegalBench 87.26%Vals · Harvey Legal Agent Benchmark 8.75%Vals · Finance Agent 59.04%Vals · TaxEval 74.73%Vals · MortgageTax 66.65%Vals · MedCode 53.39%Vals · MedScribe 83.94%Vals · SRE Bench 4.58%Vals · SWE-bench Verified 80.8% (Mini-SWE-agent)Vals · Vibe Code Bench 70.39% (OpenHands)Vals · Code Migration 34.8%Vals · GPQA Diamond 93.94%Vals · MMLU Pro 90.12%Vals · ProofBench 58%
Badge
[](https://publicai.io/model-index/m/gemini-3-7-flash)