‹ PublicAI Index
The LLM benchmark aggregator.
Inkling Small
Thinking Machines · 266B · open weights
Strongest in Finance (#53 of 151), weakest in Mathematics (#121 of 140). Above par in 11 of 16 scopes. Among the models it meets almost everywhere, it finishes behind Claude Fable 5 and Claude Opus 4.7 and ahead of GLM-4.6 and DeepSeek R1.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Knowledge54.5−9#61/1381/2
Academic knowledge55.4−8.5#56/1231/1
Professional51.5−10.1#65/1681/1
Finance54.5−8.4#53/1511/1
Medical50.3−16.5#78/1401/1
Legal48.8−18.3#87/1511/1
Coding47−22.1#97/1651/5
Agentic coding46.7−22#91/1571/4
Agents51−17.1#107/2681/5
Knowledge work51.5−22.1#70/1781/2
Reasoning46.4−20.5#112/1782/4
Reasoning51.1−18#56/1401/3
Science55.2−8.3#57/1221/1
Mathematics36.8−28.9#121/1401/2
Human preference57.5−10#119/3421/1
Human preference57.5−10#119/3421/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Inkling Small, left for the other.
§ 3 · Sources
Where the numbers come from
4 publications, 18 figures. Every one links to the page it was read from.
LMArena Text 1405
AA-Briefcase 912
ARC-AGI-2 40.1%
Vals · Legal Research Bench 25.48%Vals · LegalBench 82.99%Vals · Harvey Legal Agent Benchmark 1.67%Vals · Finance Agent 41.26%Vals · CorpFin 69.62%Vals · TaxEval 75.51%Vals · MortgageTax 62.28%Vals · MedCode 37.89%Vals · MedScribe 84.11%Vals · SWE-bench Verified 82.2% (Mini-SWE-agent)Vals · Vibe Code Bench 19.05% (OpenHands)Vals · Code Migration 13.68%Vals · GPQA Diamond 83.59%Vals · MMLU Pro 85.57%Vals · ProofBench 6%
Badge
[](https://publicai.io/model-index/m/inkling-small)