‹ PublicAI Index
The LLM benchmark aggregator.
Nova Micro
Amazon
Strongest in Factual grounding (#18 of 101), weakest in Human preference (#256 of 342). Above par in 6 of 11 scopes. Among the models it meets almost everywhere, it finishes behind Phi 4 and Claude Sonnet 5 and ahead of Gemma 4 31B and Grok 4.1 Fast.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Safety53.6−7.9#80/3372/3
Factual grounding61.3−9.3#18/1011/1
Harm refusal54.5−7.1#79/3001/2
Jailbreak resistance56.6−8.3#92/2721/1
Toxicity avoidance54.7−4.3#117/2721/1
Secure code51.8−14.7#138/2741/1
Fairness44−26#220/3001/2
Agents43.7−24.4#201/2681/5
Tool use38.7−35.3#68/811/1
Human preference41.7−25.8#256/3421/1
Human preference41.7−25.8#256/3421/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Nova Micro, left for the other.
§ 3 · Sources
Where the numbers come from
4 publications, 9 figures. Every one links to the page it was read from.
LMArena Text 1241
BFCL v4 22.29%
Enkrypt · Jailbreak risk 7.2%Enkrypt · Harmful content risk 20.6%Enkrypt · CBRN risk 9.7%Enkrypt · Toxicity risk 3%Enkrypt · Bias risk 87.1%Enkrypt · Insecure code risk 29.8%
Vectara · Factual consistency 94.5%
Badge
[](https://publicai.io/model-index/m/nova-micro)