‹ PublicAI Index
The LLM benchmark aggregator.
Qwen3 14B
Alibaba · 15B · open weights
Strongest in Factual grounding (#16 of 101), weakest in Agents (#184 of 267). Above par in 3 of 7 scopes. Among the models it meets almost everywhere, it finishes behind GPT-5.6 Terra and Claude Sonnet 5 and ahead of Gemma 4 31B and Qwen3 32B.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Safety53.6−7.9#81/3361/3
Factual grounding61.5−9.1#16/1011/1
Core abilities48.8−18.6#111/2041/3
General intelligence48.2−21.9#114/2041/3
Agents44.9−23.2#184/2672/5
Tool use52.3−21.7#33/811/1
Knowledge work37.1−36.5#147/1771/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Qwen3 14B, left for the other.
§ 3 · Sources
Where the numbers come from
4 publications, 4 figures. Every one links to the page it was read from.
GDPval-AA -39
BFCL v4 41.03%
Kagi LLM Benchmark 49.1%
Vectara · Factual consistency 94.6%
Badge
[](https://publicai.io/model-index/m/qwen3-14b)