‹ PublicAI Index
The LLM benchmark aggregator.
Grok Build 0.1
SpaceXAI
Strongest in Instruction following (#37 of 57), weakest in Safety (#290 of 337). Above par in 3 of 19 scopes. Among the models it meets almost everywhere, it finishes behind GPT-5.5 and Claude Opus 5.5 and ahead of Command A and Qwen3.6 Plus.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Agents53.8−14.3#78/2681/5
Knowledge work55.7−17.9#47/1781/2
Coding38.8−30.3#160/1652/5
Code generation30−40#77/771/2
Agentic coding42.7−26#124/1572/4
Reasoning37.8−29.1#162/1781/4
Reasoning39.8−29.3#113/1401/3
Mathematics34.2−31.5#132/1401/2
Core abilities41.6−25.8#180/2041/3
Instruction following45.2−28.5#37/571/1
Data analysis45.6−20.5#42/571/1
Language37.2−34.3#51/571/1
General intelligence38.8−31.3#181/2041/3
Safety44.8−16.7#290/3371/3
Secure code60.2−6.3#66/2741/1
Fairness45.5−24.5#182/3001/2
Toxicity avoidance41.3−17.7#239/2721/1
Jailbreak resistance32.2−32.7#248/2721/1
Harm refusal44.1−17.5#260/3001/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Grok Build 0.1, left for the other.
§ 3 · Sources
Where the numbers come from
4 publications, 16 figures. Every one links to the page it was read from.
GDPval-AA 1052
LiveBench 67.8LiveBench · Reasoning 76.4LiveBench · Coding 65.4LiveBench · Agentic Coding 45.8LiveBench · Mathematics 78.4LiveBench · Data Analysis 70.8LiveBench · Language 72.5LiveBench · Instruction Following 65.2
Vals · Vibe Code Bench 13.35% (Grok Build)
Enkrypt · Jailbreak risk 27.8%Enkrypt · Harmful content risk 13.3%Enkrypt · CBRN risk 46.5%Enkrypt · Toxicity risk 12.4%Enkrypt · Bias risk 84.8%Enkrypt · Insecure code risk 12.9%
Badge
[](https://publicai.io/model-index/m/grok-build-0-1)