‹ PublicAI Index
The LLM benchmark aggregator.
Nemotron 3 Ultra
NVIDIA · 561B · open weights
Strongest in Science (#44 of 122), weakest in Coding (#145 of 165). Above par in 8 of 22 scopes. Among the models it meets almost everywhere, it finishes behind Claude Sonnet 5 and Claude Opus 5 and ahead of Gemma 4 31B and Llama 4 Maverick Instruct.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Knowledge54.7−8.8#59/1381/2
Academic knowledge55.7−8.2#54/1231/1
Factuality52.6−6.4—/0✱✱
Reasoning50.7−16.2#86/1782/4
Science57−6.5#44/1221/1
Reasoning46.2−22.9#75/1401/3
Expert reasoning49.6−11.3—/0✱✱
Professional48.1−13.5#104/1681/1
Finance51.2−11.7#80/1511/1
Medical46.7−20.1#100/1401/1
Legal45.2−21.9#102/1511/1
Core abilities48.9−18.5#109/2041/3
General intelligence48.3−21.8#112/2041/3
Long context53.2−11.4—/0✱✱
Coding41.6−27.5#145/1651/5
Agentic coding40.6−28.1#142/1571/4
Code generation52−18—/77✱0/2
Agents40.3−27.8—/268✱0/5
Knowledge work33.6−40—/178✱0/2
Tool use34.2−39.8—/81✱0/1
Web research31.8−32.4—/4✱0/1
Workflow automation35.6−26.4—/0✱✱
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Nemotron 3 Ultra, left for the other.
§ 3 · Sources
Where the numbers come from
5 publications, 36 figures. Every one links to the page it was read from.
Artificial Analysis Intelligence Index 23
SimpleBench 41.7%
Vals · Legal Research Bench 15.38%Vals · LegalBench 82.07%Vals · Harvey Legal Agent Benchmark 0.42%Vals · Finance Agent 37.67%Vals · CorpFin 65.46%Vals · TaxEval 73.1%Vals · MedCode 38.62%Vals · SWE-bench Verified 69% (Mini-SWE-agent)Vals · Vibe Code Bench 7.64% (OpenHands)Vals · Code Migration 4.91%Vals · GPQA Diamond 86.11%Vals · MMLU Pro 85.76%
GDPVal-AA 1162tau3-Banking 14.2%Toolathlon Verified 34.3%Automation Bench Public 8%Apex-Agents (pass@1) 9%MCPMark 45.7%BrowseComp 44.4%WildClawBench 34.2%Terminal-Bench 2.1 53.9%SciCode 39.9%SWE Bench Pro 38.7%Humanity's Last Exam (without tools) 28.4%GPQA Diamond 86.7%CritPt 3.1%AA-LCR 71%AA-Omniscience Accuracy 23%AA-Omniscience Non-Hallucination 70%
GDPVal-AA (Elo) 1162Toolathlon Verified 34.3%Terminal-Bench 2.1 53.9%SWE-bench Pro (strict) 38.7%MCPMark 45.7%
Badge
[](https://publicai.io/model-index/m/nemotron-3-ultra)