‹ PublicAI Index
The LLM benchmark aggregator.
Mercury 2.5
Inception
Strongest in Legal (#119 of 151), weakest in Agents (#207 of 268). Above par in 0 of 12 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5 and GPT-6 Astra and ahead of Command A and Grok Build 0.1.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Reasoning40.7−26.2#151/1781/4
Mathematics35.9−29.8#125/1401/2
Coding40−29.1#153/1651/5
Agentic coding38.5−30.2#149/1571/4
Core abilities44.2−23.2#157/2041/3
General intelligence41.6−28.5#161/2041/3
Professional38.7−22.9#160/1681/1
Legal42.7−24.4#119/1511/1
Medical32.5−34.3#137/1401/1
Finance35−27.9#144/1511/1
Agents43.6−24.5#207/2682/5
Knowledge work41.7−31.9#125/1782/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Mercury 2.5, left for the other.
§ 3 · Sources
Where the numbers come from
4 publications, 12 figures. Every one links to the page it was read from.
Artificial Analysis Intelligence Index 12
GDPval-AA 462
AA-Briefcase 377
Vals · Legal Research Bench 4.33%Vals · LegalBench 83.06%Vals · Harvey Legal Agent Benchmark 0%Vals · Finance Agent 18.56%Vals · MedCode 31.33%Vals · MedScribe 55.09%Vals · Vibe Code Bench 3.68% (OpenHands)Vals · Code Migration 4.49%Vals · ProofBench 3%
Badge
[](https://publicai.io/model-index/m/mercury-2-5)