‹ PublicAI Index
The LLM benchmark aggregator.
Qwen3.8 Flash Next
Alibaba
Strongest in Instruction following (#5 of 57), weakest in Mathematics (#97 of 140). Above par in 8 of 13 scopes. Among the models it meets almost everywhere, it finishes behind Muse Spark 1.3 and Claude Fable 5.1 and ahead of DeepSeek V4 Flash Vision and Gemini 3.5 Flash.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Agents64.2−3.9#8/2682/5
Knowledge work68.4−5.2#7/1782/2
Core abilities54.7−12.7#50/2042/3
Instruction following66.1−7.6#5/571/1
Data analysis51.3−14.8#32/571/1
General intelligence57.3−12.8#42/2042/3
Language41.2−30.3#43/571/1
Coding49.9−19.2#77/1651/5
Agentic coding55.9−12.8#33/1571/4
Code generation40.2−29.8#60/771/2
Reasoning49.8−17.1#92/1781/4
Reasoning54.9−14.2#42/1401/3
Mathematics45−20.7#97/1401/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Qwen3.8 Flash Next, left for the other.
§ 3 · Sources
Where the numbers come from
4 publications, 11 figures. Every one links to the page it was read from.
Artificial Analysis Intelligence Index 40
GDPval-AA 1612
AA-Briefcase 1588
LiveBench 76.2LiveBench · Reasoning 87.4LiveBench · Coding 72.6LiveBench · Agentic Coding 61.6LiveBench · Mathematics 85.8LiveBench · Data Analysis 74.2LiveBench · Language 74.6LiveBench · Instruction Following 77.1
Badge
[](https://publicai.io/model-index/m/qwen3-8-flash-next)