‹ PublicAI Index
The LLM benchmark aggregator.
Llama 3.3 70B Instruct
Meta · 71B · open weights
Strongest in Tool use (#44 of 81), weakest in Agents (#231 of 267). Above par in 0 of 7 scopes. Among the models it meets almost everywhere, it finishes behind GLM-5.2 and GPT-5.6 Terra and ahead of Gemma 3 27B It and Gemma 3 12B It.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Reasoning41.5−25.4#146/1781/4
Reasoning36.4−32.7#130/1401/3
Human preference49.1−18.4#204/3411/1
Human preference49.1−18.4#204/3411/1
Agents41.2−26.9#231/2672/5
Tool use45.7−28.3#44/811/1
Knowledge work34.5−39.1#161/1771/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Llama 3.3 70B Instruct, left for the other.
§ 3 · Sources
Where the numbers come from
4 publications, 4 figures. Every one links to the page it was read from.
LMArena Text 1318
GDPval-AA -189
BFCL v4 31.9%
SimpleBench 19.9%
Badge
[](https://publicai.io/model-index/m/llama-3-3-70b-instruct)