‹ PublicAI Index
The LLM benchmark aggregator.
Qwen2 VL 72B
Alibaba · 72B
Strongest in Knowledge (#1 of 138, on 1 of its 2 boards), weakest in Fairness (#194 of 299). Above par in 6 of 8 scopes. Among the models it meets almost everywhere, it finishes behind Grok 4.7 and Claude 3 Opus and ahead of Gemini 3.8 Flash and DeepSeek V4 Pro.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Knowledge63.5leads#1/1381/2
Multimodal understanding66.2−5.9#4/191/1
Safety50.9−10.6#160/3361/3
Secure code57.3−9.2#94/2741/1
Toxicity avoidance55.6−3.4#95/2721/1
Jailbreak resistance53.9−11#132/2721/1
Harm refusal48.1−13.5#187/2991/2
Fairness45.2−24.8#194/2991/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Qwen2 VL 72B, left for the other.
§ 3 · Sources
Where the numbers come from
2 publications, 7 figures. Every one links to the page it was read from.
MMMU-Pro 46.2%
Enkrypt · Jailbreak risk 9.5%Enkrypt · Harmful content risk 41.7%Enkrypt · CBRN risk 15.3%Enkrypt · Toxicity risk 2.4%Enkrypt · Bias risk 85.3%Enkrypt · Insecure code risk 18.7%
Badge
[](https://publicai.io/model-index/m/qwen2-vl-72b)