‹ PublicAI Index
The LLM benchmark aggregator.
Llama 4 Scout Instruct
Meta · 109B · open weights
Strongest in Factual grounding (#37 of 101), weakest in Secure code (#254 of 274). Above par in 4 of 25 scopes. Among the models it meets almost everywhere, it finishes behind Claude Sonnet 5 and Inkling and ahead of Command A and Jamba 1.6 Large.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Knowledge37.9−25.6#123/1381/2
Academic knowledge35.5−28.4#108/1231/1
Core abilities44.6−22.8#150/2041/3
General intelligence42.2−27.9#155/2041/3
Reasoning36.7−30.2#164/1782/4
Reasoning42.7−26.4#102/1401/3
Science29.7−33.8#112/1221/1
Mathematics37.4−28.3#117/1401/2
Professional36.3−25.3#166/1681/1
Legal45.3−21.8#101/1511/1
Finance38.8−24.1#134/1511/1
Medical26.1−40.7#140/1401/1
Human preference49.5−18#200/3421/1
Human preference49.5−18#200/3421/1
Safety48.9−12.6#210/3373/3
Factual grounding55.7−14.9#37/1011/1
Safe-prompt compliance53.3−7.1#44/821/1
Toxicity avoidance56.1−2.9#72/2721/1
Jailbreak resistance56.8−8.1#84/2721/1
Fairness46.1−23.9#171/3002/2
Harm refusal48.8−12.8#172/3002/2
Secure code29.3−37.2#254/2741/1
Agents40.2−27.9#237/2682/5
Tool use43−31#52/811/1
Knowledge work34.7−38.9#161/1781/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Llama 4 Scout Instruct, left for the other.
§ 3 · Sources
Where the numbers come from
9 publications, 27 figures. Every one links to the page it was read from.
LMArena Text 1322
GDPval-AA -178
ARC-AGI-2 0%
BFCL v4 28.13%
Kagi LLM Benchmark 36.9%
Vals · LegalBench 72.04%Vals · CorpFin 46.78%Vals · TaxEval 55.19%Vals · MortgageTax 57.75%Vals · MedQA 50.9%Vals · MedCode 23.31%Vals · MedScribe 50.59%Vals · GPQA Diamond 46.97%Vals · MMLU Pro 69.63%Vals · AIME 18.96%
HELM Safety · HarmBench 60%HELM Safety · SimpleSafetyTests 97%HELM Safety · Anthropic Red Team 96.5%HELM Safety · BBQ 87.5%HELM Safety · XSTest 95.6%
Enkrypt · Jailbreak risk 7%Enkrypt · Harmful content risk 31.7%Enkrypt · CBRN risk 12.2%Enkrypt · Toxicity risk 2%Enkrypt · Bias risk 86.1%Enkrypt · Insecure code risk 75.6%
Vectara · Factual consistency 92.3%
Badge
[](https://publicai.io/model-index/m/llama-4-scout-instruct)