‹ PublicAI IndexGoogle
The LLM benchmark aggregator.
Gemini 2.5 Pro
Strongest in Safe-prompt compliance (#4 of 82), weakest in Jailbreak resistance (#258 of 272). Above par in 20 of 27 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5 and Claude Fable 5.1 and ahead of Inkling and O3.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Professional55−6.6#33/1681/1
Legal56.5−10.6#29/1511/1
Medical54.7−12.1#43/1401/1
Finance55.1−7.8#44/1511/1
Core abilities56−11.4#36/2041/3
General intelligence58.6−11.5#36/2041/3
Human preference61.4−6.1#67/3421/1
Human preference61.4−6.1#67/3421/1
Reasoning52.7−14.2#69/1783/4
Mathematics55.2−10.5#46/1401/2
Reasoning50.5−18.6#59/1402/3
Science53.3−10.2#66/1221/1
Knowledge53−10.5#71/1381/2
Academic knowledge53.5−10.4#65/1231/1
Coding49.6−19.5#79/1653/5
Code generation59.3−10.7#12/771/2
Agentic coding46.3−22.4#94/1572/4
Safety50.2−11.3#177/3373/3
Safe-prompt compliance60.2−0.2#4/821/1
Factual grounding57.5−13.1#32/1011/1
Fairness60.4−9.6#41/3002/2
Toxicity avoidance53.3−5.7#146/2721/1
Secure code47.6−18.9#181/2741/1
Harm refusal47−14.6#209/3002/2
Jailbreak resistance28.3−36.6#258/2721/1
Agents42.7−25.4#213/2682/5
Knowledge work40.6−33#130/1782/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Gemini 2.5 Pro, left for the other.
§ 3 · Sources
Where the numbers come from
12 publications, 33 figures. Every one links to the page it was read from.
LMArena Text 1446
GDPval-AA 444
AA-Briefcase 302
ARC-AGI-2 4.9%
LiveCodeBench 73.6%
Aider polyglot 83.1%
Kagi LLM Benchmark 70.3%
SimpleBench 62.4%
Vals · CaseLaw 63.88%Vals · LegalBench 84.32%Vals · CorpFin 60.8%Vals · TaxEval 72.89%Vals · MortgageTax 68.92%Vals · MedQA 93.14%Vals · MedCode 50.59%Vals · MedScribe 73.55%Vals · SWE-bench Verified 54.4% (Mini-SWE-agent)Vals · Vibe Code Bench 0.4% (OpenHands)Vals · GPQA Diamond 80.81%Vals · MMLU Pro 84.06%Vals · AIME 85.83%
HELM Safety · HarmBench 65.4%HELM Safety · SimpleSafetyTests 97%HELM Safety · Anthropic Red Team 99.5%HELM Safety · BBQ 96.4%HELM Safety · XSTest 98.7%
Enkrypt · Jailbreak risk 31.1%Enkrypt · Harmful content risk 54.6%Enkrypt · CBRN risk 20.8%Enkrypt · Toxicity risk 4%Enkrypt · Bias risk 61.2%Enkrypt · Insecure code risk 38.5%
Vectara · Factual consistency 93%
Badge
[](https://publicai.io/model-index/m/gemini-2-5-pro)