‹ PublicAI IndexGoogle
The LLM benchmark aggregator.
Gemini 3.8 Flash
Strongest in Instruction following (#1 of 57), weakest in Jailbreak resistance (#200 of 272). Above par in 24 of 30 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5.1 and Claude Fable 5 and ahead of DeepSeek V4.1 Flash and GPT-5.6 Terra.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Knowledge59.4−4.1#6/1381/2
Academic knowledge61.2−2.7#5/1231/1
Human preference65.9−1.6#8/3421/1
Human preference65.9−1.6#8/3421/1
Reasoning59.7−7.2#13/1784/4
Science62.8−0.7#4/1221/1
Reasoning65.5−3.6#6/1403/3
Mathematics51.4−14.3#77/1402/2
Core abilities56.1−11.3#32/2042/3
Instruction following73.7leads#1/571/1
Language66−5.5#5/571/1
General intelligence57.3−12.8#41/2042/3
Data analysis26−40.1#56/571/1
Agents58.9−9.2#38/2682/5
Knowledge work61.6−12#32/1782/2
Professional53.7−7.9#47/1681/1
Cybersecurity29.9−29.8#8/91/1
Finance58.7−4.2#11/1511/1
Biology research37.5−31#13/161/1
Legal58.8−8.3#18/1511/1
Medical56.8−10#26/1401/1
Coding51.5−17.6#63/1653/5
Agentic coding54.4−14.3#43/1573/4
Code generation40−30#61/771/2
Safety52.2−9.3#126/3371/3
Secure code63−3.5#33/2741/1
Fairness56.5−13.5#55/3001/2
Toxicity avoidance54−5#129/2721/1
Harm refusal47.8−13.8#193/3001/2
Jailbreak resistance46.8−18.1#200/2721/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Gemini 3.8 Flash, left for the other.
§ 3 · Sources
Where the numbers come from
10 publications, 37 figures. Every one links to the page it was read from.
LMArena Text 1492
Artificial Analysis Intelligence Index 41
GDPval-AA 1412
AA-Briefcase 1202
Terminal-Bench 19.1% (mini-SWE-agent)
ARC-AGI-2 89.2%
LiveBench 75.8LiveBench · Reasoning 89.3LiveBench · Coding 72.5LiveBench · Agentic Coding 54.2LiveBench · Mathematics 91.6LiveBench · Data Analysis 54LiveBench · Language 87.8LiveBench · Instruction Following 81.4
SimpleBench 82.4%
Vals · Legal Research Bench 38.94%Vals · LegalBench 86.99%Vals · Harvey Legal Agent Benchmark 10%Vals · Finance Agent 61.44%Vals · TaxEval 74.45%Vals · MortgageTax 65.34%Vals · MedCode 48.13%Vals · MedScribe 84.5%Vals · BioMysteryBench 62.22%Vals · CyberBench 43.75%Vals · SWE-bench Verified 80% (Mini-SWE-agent)Vals · Vibe Code Bench 78.65% (OpenHands)Vals · Code Migration 36.55%Vals · GPQA Diamond 94.44%Vals · MMLU Pro 90.22%Vals · ProofBench 48%
Enkrypt · Jailbreak risk 15.5%Enkrypt · Harmful content risk 3.9%Enkrypt · CBRN risk 36.2%Enkrypt · Toxicity risk 3.5%Enkrypt · Bias risk 67.7%Enkrypt · Insecure code risk 7.1%
Badge
[](https://publicai.io/model-index/m/gemini-3-8-flash)