‹ PublicAI Index
The LLM benchmark aggregator.
Command R7B
Cohere
Strongest in Tool use (#43 of 81), weakest in Fairness (#292 of 299). Above par in 3 of 8 scopes. Among the models it meets almost everywhere, it finishes behind Claude Sonnet 5 and GPT-5.2 and ahead of Gemini 3.6 Flash and DeepSeek R1.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Agents47.7−20.4#143/2671/5
Tool use45.8−28.2#43/811/1
Safety45.9−15.6#276/3361/3
Secure code58.8−7.7#76/2741/1
Toxicity avoidance52.6−6.4#163/2721/1
Jailbreak resistance50.8−14.1#177/2721/1
Harm refusal40.3−21.3#287/2991/2
Fairness36.1−33.9#292/2991/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Command R7B, left for the other.
§ 3 · Sources
Where the numbers come from
2 publications, 7 figures. Every one links to the page it was read from.
BFCL v4 32.07%
Enkrypt · Jailbreak risk 12.1%Enkrypt · Harmful content risk 55%Enkrypt · CBRN risk 28.7%Enkrypt · Toxicity risk 4.5%Enkrypt · Bias risk 99.5%Enkrypt · Insecure code risk 15.6%
Badge
[](https://publicai.io/model-index/m/command-r7b)