‹ PublicAI IndexGoogle
The LLM benchmark aggregator.
Gemini 1.5 Pro
Strongest in Multimodal understanding (#3 of 19), weakest in Secure code (#252 of 274). Above par in 8 of 20 scopes. Among the models it meets almost everywhere, it finishes behind Claude Sonnet 4.5 and GPT-5.1 and ahead of Command A and GPT-4.1 Nano.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Knowledge54.7−8.8#58/1382/2
Multimodal understanding66.9−5.2#3/191/1
Academic knowledge42.6−21.3#103/1231/1
Professional41.8−19.8#151/1681/1
Medical46.3−20.5#103/1401/1
Legal43−24.1#118/1511/1
Finance37.6−25.3#135/1511/1
Reasoning37.9−29#161/1783/4
Science37.6−25.9#103/1221/1
Mathematics37.4−28.3#118/1401/2
Reasoning38.4−30.7#120/1402/3
Human preference52.3−15.2#165/3421/1
Human preference52.3−15.2#165/3421/1
Safety49.9−11.6#181/3372/3
Safe-prompt compliance41.7−18.7#71/821/1
Harm refusal53.7−7.9#95/3002/2
Jailbreak resistance56.1−8.8#100/2721/1
Fairness50.1−19.9#115/3002/2
Toxicity avoidance52.4−6.6#170/2721/1
Secure code30.2−36.3#252/2741/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Gemini 1.5 Pro, left for the other.
§ 3 · Sources
Where the numbers come from
7 publications, 23 figures. Every one links to the page it was read from.
LMArena Text 1351
ARC-AGI-2 0.8%
MMMU-Pro 46.9%
SimpleBench 27.1%
Vals · LegalBench 69.08%Vals · CorpFin 40.52%Vals · TaxEval 59.48%Vals · MortgageTax 56.76%Vals · MedQA 76.53%Vals · GPQA Diamond 58.33%Vals · MMLU Pro 75.29%Vals · AIME 18.75%
HELM Safety · HarmBench 79.9%HELM Safety · SimpleSafetyTests 97.5%HELM Safety · Anthropic Red Team 99.9%HELM Safety · BBQ 94.5%HELM Safety · XSTest 90.4%
Enkrypt · Jailbreak risk 7.6%Enkrypt · Harmful content risk 31.7%Enkrypt · CBRN risk 9.8%Enkrypt · Toxicity risk 4.6%Enkrypt · Bias risk 85.3%Enkrypt · Insecure code risk 73.8%
Badge
[](https://publicai.io/model-index/m/gemini-1-5-pro)