‹ PublicAI Index
The LLM benchmark aggregator.
Phi-4 Mini Instruct
Microsoft
Strongest in Factual grounding (#100 of 101), weakest in Agents (#265 of 268). Above par in 3 of 9 scopes. Among the models it meets almost everywhere, it finishes behind Claude Opus 4.7 and GPT-5 and ahead of GLM-4.6 and Qwen3 8B.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Safety47.2−14.3#252/3372/3
Factual grounding26−44.6#100/1011/1
Secure code54.7−11.8#116/2741/1
Toxicity avoidance53.9−5.1#136/2721/1
Harm refusal49.2−12.4#167/3001/2
Jailbreak resistance51.4−13.5#170/2721/1
Fairness44.6−25.4#206/3001/2
Agents37−31.1#265/2681/5
Knowledge work30.5−43.1#177/1781/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Phi-4 Mini Instruct, left for the other.
§ 3 · Sources
Where the numbers come from
3 publications, 8 figures. Every one links to the page it was read from.
GDPval-AA -425
Enkrypt · Jailbreak risk 11.6%Enkrypt · Harmful content risk 39.4%Enkrypt · CBRN risk 13.7%Enkrypt · Toxicity risk 3.6%Enkrypt · Bias risk 86.3%Enkrypt · Insecure code risk 24%
Vectara · Factual consistency 76.5%
Badge
[](https://publicai.io/model-index/m/phi-4-mini-instruct)