‹ PublicAI Index
The LLM benchmark aggregator.
OLMo 2 7B Instruct
Ai2 · 7B
Strongest in Jailbreak resistance (#55 of 272), weakest in Safety (#300 of 336). Above par in 2 of 7 scopes. Among the models it meets almost everywhere, it finishes behind Claude 3.5 Sonnet and Claude 3 Opus and ahead of GPT-3.5 Turbo and Mistral Instruct v0.3 (7B).
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Safety43.3−18.2#300/3362/3
Jailbreak resistance59.2−5.7#55/2721/1
Safe-prompt compliance38.2−22.2#73/821/1
Toxicity avoidance55.8−3.2#91/2721/1
Secure code47.7−18.8#180/2741/1
Harm refusal41.5−20.1#278/2992/2
Fairness33.5−36.5#298/2992/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for OLMo 2 7B Instruct, left for the other.
§ 3 · Sources
Where the numbers come from
2 publications, 11 figures. Every one links to the page it was read from.
HELM Safety · HarmBench 57.8%HELM Safety · SimpleSafetyTests 77.5%HELM Safety · Anthropic Red Team 88%HELM Safety · BBQ 52.5%HELM Safety · XSTest 88.8%
Enkrypt · Jailbreak risk 5%Enkrypt · Harmful content risk 13.3%Enkrypt · CBRN risk 10.3%Enkrypt · Toxicity risk 2.2%Enkrypt · Bias risk 89.4%Enkrypt · Insecure code risk 38.2%
Badge
[](https://publicai.io/model-index/m/olmo-2-7b-instruct)