‹ PublicAI Index
The LLM benchmark aggregator.
Grok 4.3
SpaceXAI
Strongest in Science (#22 of 122), weakest in Safety (#253 of 337). Above par in 9 of 28 scopes. Among the models it meets almost everywhere, it finishes behind Claude Opus 5 and GPT-5.5 and ahead of GLM-4.6 and Gemini 2.5 Flash Lite.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Knowledge54.8−8.7#56/1381/2
Academic knowledge55.8−8.1#52/1231/1
Human preference61.1−6.4#72/3421/1
Human preference61.1−6.4#72/3421/1
Professional49.5−12.1#88/1681/1
Legal52.8−14.3#56/1511/1
Finance48.1−14.8#100/1511/1
Medical45.8−21#109/1401/1
Agents50.9−17.2#109/2682/5
Knowledge work51.1−22.5#74/1782/2
Reasoning43.8−23.1#135/1782/4
Science60.7−2.8#22/1221/1
Mathematics42.8−22.9#107/1401/2
Reasoning32.1−37#138/1401/3
Coding37.7−31.4#162/1652/5
Code generation34.6−35.4#72/771/2
Agentic coding38.5−30.2#148/1572/4
Core abilities35−32.4#203/2041/3
Instruction following41−32.7#44/571/1
Language39.3−32.2#48/571/1
Data analysis26−40.1#54/571/1
General intelligence34−36.1#198/2041/3
Safety47.1−14.4#253/3371/3
Secure code56−10.5#109/2741/1
Harm refusal48.1−13.5#186/3001/2
Fairness44−26#219/3001/2
Toxicity avoidance46.7−12.3#222/2721/1
Jailbreak resistance38.6−26.3#225/2721/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Grok 4.3, left for the other.
§ 3 · Sources
Where the numbers come from
6 publications, 32 figures. Every one links to the page it was read from.
LMArena Text 1442
GDPval-AA 923
AA-Briefcase 761
LiveBench 62.3LiveBench · Reasoning 70.8LiveBench · Coding 69.9LiveBench · Agentic Coding 18.5LiveBench · Mathematics 84.3LiveBench · Data Analysis 55.8LiveBench · Language 73.6LiveBench · Instruction Following 62.8
Vals · Legal Research Bench 15.38%Vals · CaseLaw 79.31%Vals · LegalBench 84.46%Vals · Harvey Legal Agent Benchmark 0.42%Vals · Finance Agent 37.73%Vals · CorpFin 68.53%Vals · TaxEval 70.81%Vals · MortgageTax 48.25%Vals · MedCode 38.07%Vals · MedScribe 74.4%Vals · SWE-bench Verified 71.4% (Mini-SWE-agent)Vals · Vibe Code Bench 19.4% (OpenHands)Vals · Code Migration 6.79%Vals · GPQA Diamond 91.41%Vals · MMLU Pro 85.84%
Enkrypt · Jailbreak risk 22.4%Enkrypt · Harmful content risk 5%Enkrypt · CBRN risk 34.8%Enkrypt · Toxicity risk 8.6%Enkrypt · Bias risk 87.1%Enkrypt · Insecure code risk 21.3%
Badge
[](https://publicai.io/model-index/m/grok-4-3)