‹ PublicAI Index
The LLM benchmark aggregator.
Grok 4.1 Fast
SpaceXAI
Strongest in Tool use (#5 of 81), weakest in Safety (#323 of 337). Above par in 12 of 23 scopes. Among the models it meets almost everywhere, it finishes behind Claude Opus 5 and Claude Opus 5.5 and ahead of Kimi K2 and Inkling Small.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Agents62.8−5.3#16/2681/5
Tool use73−1#5/811/1
Reasoning55−11.9#41/1782/4
Mathematics56.8−8.9#31/1401/2
Reasoning52.6−16.5#50/1401/3
Science55.8−7.7#53/1221/1
Knowledge53.1−10.4#70/1381/2
Academic knowledge53.7−10.2#64/1231/1
Human preference59.9−7.6#89/3421/1
Human preference59.9−7.6#89/3421/1
Professional49.1−12.5#92/1681/1
Legal53.2−13.9#53/1511/1
Finance48.7−14.2#96/1511/1
Medical46.1−20.7#108/1401/1
Coding37.1−32#164/1651/5
Agentic coding35.2−33.5#157/1571/4
Safety40.7−20.8#323/3372/3
Factual grounding26.7−43.9#94/1011/1
Secure code53.6−12.9#124/2741/1
Fairness43.5−26.5#233/3001/2
Toxicity avoidance36.5−22.5#246/2721/1
Jailbreak resistance31.3−33.6#252/2721/1
Harm refusal44.2−17.4#257/3001/2
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Grok 4.1 Fast, left for the other.
§ 3 · Sources
Where the numbers come from
6 publications, 23 figures. Every one links to the page it was read from.
LMArena Text 1430
BFCL v4 69.57%
SimpleBench 56%
Vals · CaseLaw 60.45%Vals · LegalBench 82.45%Vals · CorpFin 65.97%Vals · TaxEval 73.14%Vals · MortgageTax 42.61%Vals · MedQA 92.08%Vals · MedCode 28.08%Vals · MedScribe 78.73%Vals · SWE-bench Verified 41.4% (Mini-SWE-agent)Vals · Vibe Code Bench 1.2% (OpenHands)Vals · GPQA Diamond 84.34%Vals · MMLU Pro 84.18%Vals · AIME 91.88%
Enkrypt · Jailbreak risk 28.6%Enkrypt · Harmful content risk 36.7%Enkrypt · CBRN risk 28.2%Enkrypt · Toxicity risk 15.8%Enkrypt · Bias risk 87.9%Enkrypt · Insecure code risk 26.2%
Vectara · Factual consistency 80.8%
Badge
[](https://publicai.io/model-index/m/grok-4-1-fast)