‹ PublicAI Index
The LLM benchmark aggregator.
Command R Plus
Cohere
Strongest in Factual grounding (#30 of 101), weakest in Safety (#274 of 337). Above par in 2 of 18 scopes. Among the models it meets almost everywhere, it finishes behind O3 and O1 and ahead of Jamba 1.6 Mini and Jamba 1.5 Mini.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Knowledge30−33.5#137/1381/2
Academic knowledge26−37.9#122/1231/1
Professional41.8−19.8#150/1681/1
Legal42.4−24.7#122/1511/1
Medical32.9−33.9#136/1401/1
Reasoning35.2−31.7#172/1782/4
Science26.1−37.4#119/1221/1
Reasoning35.3−33.8#135/1401/3
Human preference45.1−22.4#240/3421/1
Human preference45.1−22.4#240/3421/1
Safety46−15.5#274/3373/3
Factual grounding57.8−12.8#30/1011/1
Safe-prompt compliance49.3−11.1#59/821/1
Jailbreak resistance57.2−7.7#80/2721/1
Harm refusal46.6−15#213/3002/2
Toxicity avoidance44.6−14.4#232/2721/1
Fairness42.5−27.5#255/3002/2
Secure code26−40.5#265/2741/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Command R Plus, left for the other.
§ 3 · Sources
Where the numbers come from
6 publications, 18 figures. Every one links to the page it was read from.
LMArena Text 1276
SimpleBench 17.4%
Vals · LegalBench 68.29%Vals · MedQA 2.65%Vals · GPQA Diamond 31.06%Vals · MMLU Pro 44%
HELM Safety · HarmBench 48.5%HELM Safety · SimpleSafetyTests 100%HELM Safety · Anthropic Red Team 98%HELM Safety · BBQ 89.9%HELM Safety · XSTest 93.8%
Enkrypt · Jailbreak risk 6.7%Enkrypt · Harmful content risk 60%Enkrypt · CBRN risk 9.2%Enkrypt · Toxicity risk 10.1%Enkrypt · Bias risk 98.7%Enkrypt · Insecure code risk 86.7%
Vectara · Factual consistency 93.1%
Badge
[](https://publicai.io/model-index/m/command-r-plus)