The LLM benchmark aggregator.
Hy3
Tencent · 299B · open weights
Strongest in Harm refusal (#3 of 300, on 1 of its 2 boards), weakest in Toxicity avoidance (#120 of 272). Above par in 10 of 12 scopes. Among the models it meets almost everywhere, it finishes behind Claude Opus 5.5 and Claude Opus 4.7 and ahead of Gemini 3.6 Flash and Qwen3.7 Max.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Hy3, left for the other.
§ 3 · Sources
Where the numbers come from
5 publications, 10 figures. Every one links to the page it was read from.
Badge
[](https://publicai.io/model-index/m/hy3)