‹ PublicAI Index
The LLM benchmark aggregator.
Mercury 2
Inception AI
Strongest in Factual grounding (#81 of 101), weakest in Safety (#233 of 336). Above par in 2 of 6 scopes. Among the models it meets almost everywhere, it finishes behind Claude Opus 4.7 and Gemini 3.1 Pro and ahead of Gemma 3 4B It and Ministral 8B.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Human preference51.5−16#174/3411/1
Human preference51.5−16#174/3411/1
Agents42.7−25.4#213/2672/5
Knowledge work40.5−33.1#130/1772/2
Safety48.2−13.3#233/3361/3
Factual grounding44.1−26.5#81/1011/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Mercury 2, left for the other.
§ 3 · Sources
Where the numbers come from
4 publications, 4 figures. Every one links to the page it was read from.
LMArena Text 1343
GDPval-AA 481
AA-Briefcase 264
Vectara · Factual consistency 87.7%
Badge
[](https://publicai.io/model-index/m/mercury-2)