‹ PublicAI Index
The LLM benchmark aggregator.
C4ai Aya Expanse 32B
Cohere · 32B
Strongest in Secure code (#51 of 274), weakest in Human preference (#246 of 342). Above par in 5 of 9 scopes. Among the models it meets almost everywhere, it finishes behind Claude Sonnet 5 and Grok 4.7 and ahead of Kimi K2.6 and GPT-6 Sol.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Safety53.8−7.7#75/3372/3
Secure code61.7−4.8#51/2741/1
Jailbreak resistance59.6−5.3#54/2721/1
Toxicity avoidance56.7−2.3#54/2721/1
Factual grounding47.7−22.9#69/1011/1
Harm refusal53.9−7.7#88/3001/2
Fairness45.5−24.5#185/3001/2
Human preference44.2−23.3#246/3421/1
Human preference44.2−23.3#246/3421/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for C4ai Aya Expanse 32B, left for the other.
§ 3 · Sources
Where the numbers come from
3 publications, 8 figures. Every one links to the page it was read from.
LMArena Text 1267
Enkrypt · Jailbreak risk 4.7%Enkrypt · Harmful content risk 31.1%Enkrypt · CBRN risk 5.7%Enkrypt · Toxicity risk 1.6%Enkrypt · Bias risk 84.8%Enkrypt · Insecure code risk 9.8%
Vectara · Factual consistency 89.1%
Badge
[](https://publicai.io/model-index/m/c4ai-aya-expanse-32b)