‹ PublicAI Index
The LLM benchmark aggregator.
Command R
Cohere
Strongest in Safe-prompt compliance (#58 of 82), weakest in Safety (#303 of 336). Above par in 0 of 8 scopes. Among the models it meets almost everywhere, it finishes behind Gemini 3 Pro and Claude Sonnet 4.5 and ahead of Nous Hermes 2 Mixtral 8x7b Dpo and Gemma 7B It.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Professional43.3−18.3#138/1681/1
Legal35−32.1#147/1511/1
Human preference42.6−24.9#254/3411/1
Human preference42.6−24.9#254/3411/1
Safety43−18.5#303/3361/3
Safe-prompt compliance49.5−10.9#58/821/1
Harm refusal42.2−19.4#274/2991/2
Fairness35.4−34.6#293/2991/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Command R, left for the other.
§ 3 · Sources
Where the numbers come from
3 publications, 7 figures. Every one links to the page it was read from.
LMArena Text 1250
Vals · LegalBench 32.97%
HELM Safety · HarmBench 50.4%HELM Safety · SimpleSafetyTests 94.3%HELM Safety · Anthropic Red Team 93.7%HELM Safety · BBQ 72.4%HELM Safety · XSTest 93.9%
Badge
[](https://publicai.io/model-index/m/command-r)