‹ PublicAI Index
The LLM benchmark aggregator.
Mixtral (8x7B)
Mistral
Strongest in Safe-prompt compliance (#62 of 82), weakest in Safety (#301 of 336). Above par in 0 of 7 scopes. Among the models it meets almost everywhere, it finishes behind GPT-5.1 and GPT-5.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Professional39.4−22.2#156/1681/1
Medical33.8−33#131/1401/1
Legal35−32.1#145/1511/1
Safety43.3−18.2#301/3361/3
Safe-prompt compliance47.7−12.7#62/821/1
Fairness47.2−22.8#155/2991/2
Harm refusal39−22.6#290/2991/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Mixtral (8x7B), left for the other.
§ 3 · Sources
Where the numbers come from
2 publications, 7 figures. Every one links to the page it was read from.
Vals · LegalBench 55.84%Vals · MedQA 53.22%
HELM Safety · HarmBench 45.1%HELM Safety · SimpleSafetyTests 90.5%HELM Safety · Anthropic Red Team 92.8%HELM Safety · BBQ 85.7%HELM Safety · XSTest 93.1%
Badge
[](https://publicai.io/model-index/m/mixtral-8x7b)