‹ PublicAI Index
The LLM benchmark aggregator.
Hy4
Tencent · 780B · open weights
Strongest in IT operations (#8 of 11), weakest in Medical (#58 of 140). Above par in 9 of 10 scopes. Among the models it meets almost everywhere, it finishes behind Claude Opus 5 and Claude Fable 5.1 and ahead of Grok 4.7 and Gemini 3.8 Flash.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Coding58.1−11#16/1641/5
Agentic coding59.4−9.3#15/1571/4
Professional53.8−7.8#43/1681/1
IT operations41.2−32.6#8/111/1
Biology research50.2−18.3#9/161/1
Legal58.5−8.6#19/1511/1
Finance55.7−7.2#38/1511/1
Medical53.3−13.5#58/1401/1
Reasoning54−12.9#48/1781/4
Mathematics56−9.7#38/1401/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Hy4, left for the other.
§ 3 · Sources
Where the numbers come from
1 publication, 11 figures. Every one links to the page it was read from.
Vals · Legal Research Bench 45.19%Vals · LegalBench 83.76%Vals · Harvey Legal Agent Benchmark 9.17%Vals · Finance Agent 55.06%Vals · MedCode 43.25%Vals · MedScribe 83.6%Vals · BioMysteryBench 69.26%Vals · SRE Bench 2.29%Vals · Vibe Code Bench 77.48% (OpenHands)Vals · Code Migration 47.43%Vals · ProofBench 75%
Badge
[](https://publicai.io/model-index/m/hy4)