‹ PublicAI Index
The LLM benchmark aggregator.
Ling 3.0 Flash
InclusionAI · 128B · open weights
Strongest in Science (#51 of 122), weakest in Safety (#269 of 337). Above par in 6 of 18 scopes. Among the models it meets almost everywhere, it finishes behind Claude Opus 5 and Claude Sonnet 5 and ahead of Claude 3.5 Haiku and GPT-5 Nano.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Reasoning52.9−14#65/1781/4
Science56.1−7.4#51/1221/1
Knowledge50.8−12.7#82/1381/2
Academic knowledge51−12.9#76/1231/1
Agents49.4−18.7#126/2681/5
Knowledge work49.2−24.4#88/1781/2
Professional44.5−17.1#131/1681/1
Finance46.3−16.6#105/1511/1
Medical45.3−21.5#111/1401/1
Legal41.1−26#129/1511/1
Coding39.4−29.7#156/1651/5
Agentic coding38.2−30.5#150/1571/4
Safety46.4−15.1#269/3371/3
Toxicity avoidance55.8−3.2#87/2721/1
Secure code56.7−9.8#101/2741/1
Fairness43.4−26.6#236/3001/2
Harm refusal45.5−16.1#239/3001/2
Jailbreak resistance31.3−33.6#253/2721/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Ling 3.0 Flash, left for the other.
§ 3 · Sources
Where the numbers come from
3 publications, 20 figures. Every one links to the page it was read from.
AA-Briefcase 798
Vals · Legal Research Bench 0%Vals · LegalBench 79.69%Vals · Harvey Legal Agent Benchmark 1.25%Vals · Finance Agent 30.31%Vals · CorpFin 61.93%Vals · TaxEval 70.65%Vals · MedCode 32.27%Vals · MedScribe 80.9%Vals · SWE-bench Verified 65.2% (Mini-SWE-agent)Vals · Vibe Code Bench 2.91% (OpenHands)Vals · Code Migration 0%Vals · GPQA Diamond 84.85%Vals · MMLU Pro 82.01%
Enkrypt · Jailbreak risk 28.6%Enkrypt · Harmful content risk 6.1%Enkrypt · CBRN risk 52.7%Enkrypt · Toxicity risk 2.2%Enkrypt · Bias risk 88.1%Enkrypt · Insecure code risk 20%
Badge
[](https://publicai.io/model-index/m/ling-3-0-flash)