‹ PublicAI Index
The LLM benchmark aggregator.
Phi 4
Microsoft · 15B · open weights
Strongest in Factual grounding (#4 of 101), weakest in Human preference (#253 of 342). Above par in 6 of 11 scopes. Among the models it meets almost everywhere, it finishes behind Claude 3 Opus and Claude Sonnet 5 and ahead of Claude Sonnet 4.5 and DeepSeek V3.2.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Safety58.3−3.2#16/3372/3
Factual grounding65.8−4.8#4/1011/1
Harm refusal59.8−1.8#10/3001/2
Jailbreak resistance62.5−2.4#19/2721/1
Toxicity avoidance57.8−1.2#36/2721/1
Secure code58.6−7.9#79/2741/1
Fairness46.5−23.5#158/3001/2
Agents46.4−21.7#160/2681/5
Tool use43.4−30.6#50/811/1
Human preference43.1−24.4#253/3421/1
Human preference43.1−24.4#253/3421/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Phi 4, left for the other.
§ 3 · Sources
Where the numbers come from
4 publications, 9 figures. Every one links to the page it was read from.
LMArena Text 1256
BFCL v4 28.79%
Enkrypt · Jailbreak risk 2.2%Enkrypt · Harmful content risk 3.3%Enkrypt · CBRN risk 5%Enkrypt · Toxicity risk 0.8%Enkrypt · Bias risk 83.2%Enkrypt · Insecure code risk 16%
Vectara · Factual consistency 96.3%
Badge
[](https://publicai.io/model-index/m/phi-4)