‹ PublicAI Index
The LLM benchmark aggregator.
Command A
Cohere · 111B · open weights
Strongest in Tool use (#26 of 81), weakest in Fairness (#291 of 300). Above par in 8 of 24 scopes. Among the models it meets almost everywhere, it finishes behind Claude Sonnet 5 and Claude Opus 4.8 and ahead of Mistral Small and GPT-4o Mini.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Agents53.5−14.6#81/2681/5
Tool use56.3−17.7#26/811/1
Professional47.4−14.2#108/1681/1
Legal54.6−12.5#42/1511/1
Medical48.4−18.4#92/1401/1
Finance39−23.9#133/1511/1
Knowledge37.4−26.1#126/1381/2
Academic knowledge34.9−29#110/1231/1
Coding39.7−29.4#155/1651/5
Agentic coding37.5−31.2#154/1571/4
Human preference52.6−14.9#161/3421/1
Human preference52.6−14.9#161/3421/1
Reasoning36.7−30.2#165/1781/4
Science30.8−32.7#110/1221/1
Mathematics35.9−29.8#126/1401/2
Core abilities41.9−25.5#175/2041/3
General intelligence38.3−31.8#182/2041/3
Safety46.1−15.4#273/3372/3
Factual grounding51.7−18.9#49/1011/1
Toxicity avoidance53.3−5.7#147/2721/1
Secure code49.7−16.8#158/2741/1
Jailbreak resistance51.8−13.1#162/2721/1
Harm refusal41.6−20#277/3001/2
Fairness36.9−33.1#291/3001/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Command A, left for the other.
§ 3 · Sources
Where the numbers come from
6 publications, 19 figures. Every one links to the page it was read from.
LMArena Text 1354
BFCL v4 46.49%
Kagi LLM Benchmark 28.8%
Vals · CaseLaw 64.52%Vals · LegalBench 79.7%Vals · CorpFin 45.96%Vals · TaxEval 61.37%Vals · MedQA 80.55%Vals · SWE-bench Verified 7.8% (Mini-SWE-agent)Vals · GPQA Diamond 48.48%Vals · MMLU Pro 69.17%Vals · AIME 13.33%
Enkrypt · Jailbreak risk 11.3%Enkrypt · Harmful content risk 51.1%Enkrypt · CBRN risk 27.3%Enkrypt · Toxicity risk 4%Enkrypt · Bias risk 98.2%Enkrypt · Insecure code risk 34.2%
Vectara · Factual consistency 90.7%
Badge
[](https://publicai.io/model-index/m/command-a)