‹ PublicAI Index
The LLM benchmark aggregator.
Laguna XS.2
Other
Strongest in Science (#104 of 122), weakest in Professional (#164 of 168). Above par in 0 of 10 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5.1 and Claude Fable 5 and ahead of Mercury 2.5 and Command R Plus.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Knowledge37.3−26.2#127/1381/2
Academic knowledge34.8−29.1#111/1231/1
Reasoning43.2−23.7#138/1781/4
Science35.4−28.1#104/1221/1
Coding40.4−28.7#150/1651/5
Agentic coding39−29.7#147/1571/4
Professional36.8−24.8#164/1681/1
Finance37.5−25.4#137/1511/1
Medical31.4−35.4#139/1401/1
Legal37.2−29.9#140/1511/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Laguna XS.2, left for the other.
§ 3 · Sources
Where the numbers come from
1 publication, 12 figures. Every one links to the page it was read from.
Vals · Legal Research Bench 0.96%Vals · LegalBench 71.03%Vals · Harvey Legal Agent Benchmark 0%Vals · Finance Agent 15.6%Vals · CorpFin 56.33%Vals · TaxEval 58.95%Vals · MedCode 21.25%Vals · MedScribe 61.43%Vals · SWE-bench Verified 55.2% (Mini-SWE-agent)Vals · Vibe Code Bench 5.21% (OpenHands)Vals · GPQA Diamond 55.05%Vals · MMLU Pro 69.05%
Badge
[](https://publicai.io/model-index/m/laguna-xs-2)