Preference
8,528,723 votes across 409 models
Blind head-to-head votes. Measures what people prefer reading, which tracks style as well as correctness.
§ 2
Leaderboards form the Overall index and every column’s numbers; reports ✱ widen the coverage, and speak only where no board has; catalogs say where a model can be called. Every figure links to its source.
Preference
8,528,723 votes across 409 models
Blind head-to-head votes. Measures what people prefer reading, which tracks style as well as correctness.
General intelligence
A composite of ten evaluations Artificial Analysis runs itself, under one harness. Only fully evaluated models are indexed: rows the site stars as estimates (not every component run independently, or based on the lab’s own claims) are left out. Several components (Terminal-Bench among them) are also indexed on their own boards, so the two overlap.
Knowledge work
Pairwise-graded outputs on economically valuable knowledge-work tasks, run by Artificial Analysis and reported as an Elo with a confidence interval. Same publisher as the Intelligence Index and AA-Briefcase, so the three are one independent voice.
Knowledge work
Agentic knowledge-work tasks run by Artificial Analysis and reported as an Elo with a confidence interval. Same publisher as the Intelligence Index and GDPval-AA, so the three are one independent voice.
Agentic coding
Top entry dated Sep 3, 2026
Scores a model plus its agent scaffold (Codex, Claude Code, Grok Build). The scaffold is part of the result, not a constant.
Reasoning
Reported per reasoning-effort tier. The same model spans a wide range across tiers; the highest published tier is indexed and recorded beside the score.
General intelligence · Reasoning · Code generation · Agentic coding · Mathematics · Data analysis · Language · Instruction following
Questions are refreshed to resist contamination. The overall figure averages the publisher’s own categories, each of which is also indexed on its own.
Code generation
Pass@1 on LeetCode, AtCoder and Codeforces problems published after each model’s training cutoff, so it resists contamination. Reasoning-effort tier is part of the label; the highest published tier is indexed.
Agentic coding
Percent of 225 Exercism exercises solved by editing code through the aider tool, so the harness is part of the number. Effort tier is in the label; the highest published tier is indexed.
Tool use
Overall accuracy across single-turn, multi-turn, agentic and hallucination-measurement function-calling tasks, run by the Gorilla team at UC Berkeley. Function-calling and prompting modes are one model; the better row is indexed.
Computer use
Success rate on 369 real desktop tasks. Entries are submitted runs, sometimes of a model inside an agent framework, which is recorded beside the score; step budgets vary by entry.
Multimodal understanding
MMMU-Pro overall accuracy. Only figures the organisers verified are indexed; starred, self-reported ones are left out. Human-expert baselines are not models and are skipped.
General intelligence
Accuracy on Kagi’s unpublished set of reasoning, coding and instruction-following questions, run by Kagi. Small and private by design, which resists contamination but cannot be inspected.
Reasoning
Average of five runs on a private set of common-sense trick questions. Small and run by one team; human-baseline rows are not models and are skipped.
Legal · Finance · Medical · Biology research · Cybersecurity · IT operations · Web research · Agentic coding · Science · Academic knowledge · Mathematics
Twenty-two evaluations Vals AI runs itself, several built with domain experts and kept private, each reported as accuracy with a confidence interval. Agentic benchmarks score a model inside a scaffold, recorded beside the figure. One publisher, so all of them are one independent voice.
Routing & classification · Calibration
Two of JevBench’s four axes: Intelligence and Calibration. Its headline score also weighs speed and price, which this index does not read as capability. Benchmark Heaven runs a router of its own and says so; that entry is excluded from its ranking. Half the decisions are sealed, and the public-minus-sealed gap it publishes is large for several entrants.
Harm refusal · Fairness · Safe-prompt compliance
Release v1.17.0
Five safety scenarios scored by an LM judge, on a fixed model set Stanford chooses; releases are periodic, so the newest models can lag by months. Higher is safer on every column, XSTest included.
Jailbreak resistance · Harm refusal · Toxicity avoidance · Fairness · Secure code
Automated red-teaming by one vendor of safety tooling; each column is the share of attacks that succeeded, so lower is safer. Method and prompt sets are Enkrypt’s own and not fully public.
Factual grounding
One task — summarising short documents — scored by Vectara’s own hallucination detector. A model that declines to summarise is scored on what it did answer; the answer rate is published beside the figure.
Report by Institute of Foundation Models (MBZUAI) · 27 benchmarks across 5 categories · marked ✱ wherever it appears
Vendor launch post. Competitor figures are as IFM printed them; the post does not say whether they were re-run or taken from the competitors’ own reports. Effort tiers vary by column.
Report by DataCamp · 9 benchmarks across 5 categories · marked ✱ wherever it appears
Third-party write-up. Figures are the vendors’ own as compiled by the author; some rows are one vendor’s internal evaluations with no external replication. Two models only.
Report by CellCog · 6 benchmarks across 2 categories · marked ✱ wherever it appears
Third-party write-up. Figures are from the vendors’ own reports as the author compiled them, not re-run. Effort tiers are not stated.
Report by DeepSeek · 19 benchmarks across 5 categories · marked ✱ wherever it appears
Vendor model card. Competitor figures are as DeepSeek printed them; the card does not say whether they were re-run or taken from the competitors’ own reports.
Report by LargitData · 1 benchmarks across 1 categories · marked ✱ wherever it appears
One hundred routing decisions over twenty conversations the authors wrote. A model is scored on agreeing with their intended route, so the figure is "fits this routing scheme", not a general accuracy.
Catalog · 219 of 815 models matched · ids, context, prices, open weights · never scored
Used as a catalog of where a model can be called and on what terms, not as a source of scores. Prices and availability are the router’s at the date read and move often. Listing here is not an endorsement of the router over the vendor’s own API.
Catalog · 219 of 815 models matched · ids, context, prices, open weights · never scored
Parameter counts are read from the weights files of the model’s repository (safetensors metadata). Only open-weight models have one; a closed model’s size is undisclosed, not estimated.
Excluded — SWE-bench VerifiedRead on 2026-09-05: the newest of its 181 entries is dated 2026-02-17. None of the models above appear on it. Including a stale board would add a column of blanks, not a signal. source
Excluded — LiteLLM — Jev vs Haiku 4.5 on request classificationRead on 2026-09-22: a careful independent test — 80 authored cases, three runs, frozen results and reproduction scripts — but it compares exactly two models. This index drops any measure with fewer than three, because standardizing a pair places the loser at 35 and the winner at 65 whether they differ by two points or thirty. The figures are worth reading at the source; they are not worth a number here. source
§ 3
Six rules, each so that a number here traces to a number there, and so that thin or self-reported evidence cannot buy a rank.
Every figure is z-scored across the models its source lists and mapped to 0–100, so an Elo and a pass rate share one scale. No single measure may place a model more than two standard deviations from the field it was measured against: on a board whose models sit close together, a modest lead standardizes into an enormous figure, and one of those could outweigh four boards. A figure a source prints as a risk or an error rate is read the other way round: the raw stays as printed, and the model with the least of it is the one ahead. Where a publisher prints an error bar, the figure counts for less.
A missing figure is left missing. Thin evidence is pulled toward 50, so one generous board cannot lift a barely-tested model. A source comparing fewer than three models is not standardized at all — two points have no spread to place anything against.
Each board’s share of the Overall index is fixed and printed beside the table, with its reason.
Overall ranks a model once two independent publishers among the boards that build the index have scored it, and at least half the weighting printed beside the table has actually been measured. Below that the prior decides more than the boards do, so the model is listed with its coverage and no number — a floor is not a placing. A domain ranks on its own board evidence.
Rank by offers a domain only when several boards measure it, or one board ranks at least twenty models there. A single board’s sub-score is not offered as a peer of Coding; it stays on each model’s page.
A ✱ figure decides nothing a board has measured: in a column, the boards’ number stands, and a report speaks only where no board has. Those rows are listed with their ✱ score and take no number. One publication is one voice per column, however many rows it printed — and an independent write-up counts for more than a vendor’s own post.
A model with no Overall score is placed among models that have one, on capability measures only, and shown as ~55 — never ranked. A safety figure never stands in for a capability, nor does a figure for work nothing else here measures; those models are listed with a dash, not a guess.
§ 4
An aggregate hides the disagreements that produced it. These are the ones worth knowing before you cite this page.
The highest tier a source publishes is indexed; the exact label is kept beside the score.
Terminal-Bench results are a model plus an agent framework; the framework is named beside the score.
Seven sources spell one model seven ways. Every merge is logged; a wrong one is possible.
Competitor numbers in a launch post may be copied, not re-run. They are marked, not laundered.
Counted from open weights or read from the name; a closed model is undisclosed, not small.
Read from each publisher on the date shown. Check the source before acting on a number.
The LLM benchmark aggregator — the world’s most comprehensive and robust LLM index, built from everyone’s benchmarks and none of our own.
A single benchmark is easy to target, so topping one board says little about the next. PublicAI runs no evaluations of its own: it aggregates the public ones onto one scale — discounted by published uncertainty, shrunk where evidence is thin, weighted by a scheme printed beside the table. Launch posts are indexed too, marked ✱ and kept out of the headline. Agents get the same answers over MCP and a JSON API.
§ 1 · Ranking
Accuracy and calibration only. JevBench also ranks speed and price, which this index does not read as capability, so a large API model can lead this column at a cost a router would not pay.
| # | Model | Decisions | Sources | |
|---|---|---|---|---|
| 1 | DeepSeek V4.1 FlashDeepSeek763BOverall 62.4 | 73.3 | Show details for DeepSeek V4.1 Flash | |
| 2 | GPT-6 LunaOpenAINEWOverall 55.7 | 72.7 | Show details for GPT-6 Luna | |
| 3 | GPT-5.6 LunaOpenAIOverall 52.9 | 70.6 | Show details for GPT-5.6 Luna | |
| 4 | DjevOther | 70.0 | Show details for Djev | |
| 5 | Gemma 4 31B IT (Autoloops)Google31BNEW | 62.7 | Show details for Gemma 4 31B IT (Autoloops) | |
| —✱ | Gemma 4 31BGoogle31BOverall ~52.4report only | 60.2 | Show details for Gemma 4 31B | |
| 6 | InstinctOtherNEW | 59.3 | Show details for Instinct | |
| 7 | Jev 1.13.0TypeSafe AI | 59.3 | Show details for Jev 1.13.0 | |
| 8 | NInfer Qwen3.8-Flash-Next mixedAlibaba | 58.7 | Show details for NInfer Qwen3.8-Flash-Next mixed | |
| 9 | NInfer Qwen3.8-27B NVFP4Alibaba27B | 58.6 | Show details for NInfer Qwen3.8-27B NVFP4 | |
| 10 | HopperOther | 58.3 | Show details for Hopper | |
| 11 | SimpleJev Qwen3.8-27BAlibaba27B | 58.1 | Show details for SimpleJev Qwen3.8-27B | |
| 12 | JevOneOther | 57.7 | Show details for JevOne | |
| 13 | CygnetOtherNEW | 57.5 | Show details for Cygnet | |
| 14 | Decider-4b v2Other4BNEW | 57.5 | Show details for Decider-4b v2 | |
| 15 | swanOneOtherNEW | 57.3 | Show details for swanOne | |
| 16 | JevK5 v0.2.0Other | 57.1 | Show details for JevK5 v0.2.0 | |
| 17 | Reflex 27BOther27B | 57.1 | Show details for Reflex 27B | |
| 18 | LitJevOther | 56.9 | Show details for LitJev | |
| 19 | Openjev SglangOther | 55.5 | Show details for Openjev Sglang | |
| 20 | reflex 4BOther4B | 55.2 | Show details for reflex 4B | |
| 21 | JqvOther | 55.2 | Show details for Jqv | |
| 22 | local-jev Qwen3.5-4BAlibaba4B | 55.1 | Show details for local-jev Qwen3.5-4B | |
| 23 | ZeroEntropy zerank-2Other | 55.1 | Show details for ZeroEntropy zerank-2 | |
| 24 | OpenJevOther | 55.0 | Show details for OpenJev | |
| 25 | Gemini 3.1 Flash LiteGoogleOverall ~51.7 | 54.1 | Show details for Gemini 3.1 Flash Lite | |
| 26 | JEV Qwen3.5-9B Base NVFP4Alibaba9B | 54.1 | Show details for JEV Qwen3.5-9B Base NVFP4 | |
| 27 | Winnow-12B Q8Other12B | 53.6 | Show details for Winnow-12B Q8 | |
| 28 | Decider 35B A3BOther35B · A3B | 53.4 | Show details for Decider 35B A3B | |
| 29 | Decision 2BOther2B | 53.3 | Show details for Decision 2B | |
| 30 | Standard One 8BOther8BNEW | 53.2 | Show details for Standard One 8B | |
| 31 | Metask Jev 4BOther4B | 53.0 | Show details for Metask Jev 4B | |
| 32 | SemIfOther | 52.9 | Show details for SemIf | |
| 33 | Jev OmniTypeSafe AI | 52.8 | Show details for Jev Omni | |
| 34 | Jobe Qwen3.5-4BAlibaba4B | 52.5 | Show details for Jobe Qwen3.5-4B | |
| 35 | Qwen3 Reranker 4BAlibaba4B | 52.4 | Show details for Qwen3 Reranker 4B | |
| 36 | Jev LocalTypeSafe AI | 52.3 | Show details for Jev Local | |
| 37 | Decision Machine 1Other | 52.2 | Show details for Decision Machine 1 | |
| 38 | VonOther | 52.2 | Show details for Von | |
| 39 | Qwen3.5-9B Jev-like data-mix v2Alibaba9B | 52.1 | Show details for Qwen3.5-9B Jev-like data-mix v2 | |
| 40 | Open-Jev 9BOther9B | 51.1 | Show details for Open-Jev 9B | |
| 41 | SimpleJev Qwen3.6-35B-A3BAlibaba35B · A3B | 51.0 | Show details for SimpleJev Qwen3.6-35B-A3B | |
| 42 | Lev 350MOther | 50.6 | Show details for Lev 350M | |
| 43 | Malkuth 4BOther4BNEW | 50.6 | Show details for Malkuth 4B | |
| 44 | JeffOther | 50.4 | Show details for Jeff | |
| 45 | Bespoke Nimble 9BOther9B | 50.1 | Show details for Bespoke Nimble 9B | |
| 46 | OpenSourceJevOther | 49.7 | Show details for OpenSourceJev | |
| 47 | Decision FastOther | 49.7 | Show details for Decision Fast | |
| 48 | Raw Phi-4 mini direct logitsMicrosoft | 49.2 | Show details for Raw Phi-4 mini direct logits | |
| 49 | openJev Verdict 1.4Other | 49.1 | Show details for openJev Verdict 1.4 | |
| 50 | System One OpenOther | 48.8 | Show details for System One Open | |
| 51 | LayaOther | 48.8 | Show details for Laya | |
| 52 | Open-Jev 2BOther2B | 48.2 | Show details for Open-Jev 2B | |
| 53 | Open Alternative JevOther | 48.0 | Show details for Open Alternative Jev | |
| 54 | TypecastlmOtherNEW | 47.5 | Show details for Typecastlm | |
| 55 | Qwen3.5-0.8B Decision ModelAlibaba0.8B | 47.3 | Show details for Qwen3.5-0.8B Decision Model | |
| 56 | Malkuth 2BOther2BNEW | 47.3 | Show details for Malkuth 2B | |
| 57 | Spark S1 4B V6Other4B | 46.7 | Show details for Spark S1 4B V6 | |
| 58 | Open Jev Deberta V3 LargeOther | 45.9 | Show details for Open Jev Deberta V3 Large | |
| 59 | BAAI bge-reranker-v2-m3Other | 45.5 | Show details for BAAI bge-reranker-v2-m3 | |
| 60 | Mixedbread mxbai-rerank-base-v2MetaStone | 45.5 | Show details for Mixedbread mxbai-rerank-base-v2 | |
| 61 | Certo v1Other | 45.1 | Show details for Certo v1 | |
| 62 | OpenDecisionOther | 45.0 | Show details for OpenDecision | |
| 63 | Alibaba GTE Reranker ModernBERT-baseOther | 43.7 | Show details for Alibaba GTE Reranker ModernBERT-base | |
| 64 | kev 0.6BOther0.6B | 43.5 | Show details for kev 0.6B | |
| 65 | Smalljev semantic-v9Other | 43.4 | Show details for Smalljev semantic-v9 | |
| 66 | JevActOtherNEW | 43.4 | Show details for JevAct | |
| 67 | Verdict SmallOtherNEW | 43.2 | Show details for Verdict Small | |
| 68 | kev 8BOther8B | 43.0 | Show details for kev 8B | |
| 69 | kev 4BOther4B | 42.9 | Show details for kev 4B | |
| 70 | Decider 2BOther2B | 42.9 | Show details for Decider 2B | |
| 71 | kev 0.5BOther0.5B | 42.0 | Show details for kev 0.5B | |
| 72 | GLiNER2.5 multiOther | 41.8 | Show details for GLiNER2.5 multi | |
| 73 | System OneOther | 41.1 | Show details for System One | |
| 74 | Raw Qwen3 4B Instruct direct logitsAlibaba4B | 40.9 | Show details for Raw Qwen3 4B Instruct direct logits | |
| 75 | openJev VerdictOther | 40.9 | Show details for openJev Verdict | |
| 76 | Open Jev JSON CanvasTypeSafe AI | 40.6 | Show details for Open Jev JSON Canvas | |
| 77 | Raw Qwen3 8B direct logitsAlibaba8B | 39.7 | Show details for Raw Qwen3 8B direct logits | |
| 78 | GLiNER2.5 smallOther | 38.6 | Show details for GLiNER2.5 small | |
| 79 | SimpleJevOther | 38.5 | Show details for SimpleJev | |
| 80 | CLM 8BOther8BNEW | 35.7 | Show details for CLM 8B | |
| 81 | Raw Qwen3 1.7B direct logitsAlibaba1.7B | 35.1 | Show details for Raw Qwen3 1.7B direct logits | |
| 82 | GLiNER2 largeOther | 34.3 | Show details for GLiNER2 large | |
| 83 | GLiNER2Other | 32.9 | Show details for GLiNER2 | |
| 84 | Raw Qwen3 0.6B direct logitsAlibaba0.6B | 31.2 | Show details for Raw Qwen3 0.6B direct logits | |
| 85 | MirrorOther | 27.8 | Show details for Mirror |
0–100, standardized: 50 is the average of the models each source lists. Provisional = fewer than 2 independent publishers. ✱ = from reports. Open a row for the model card.