Request a pilot

§ 2

Sources

Leaderboards form the Overall index and every column’s numbers; reports ✱ widen the coverage, and speak only where no board has; catalogs say where a model can be called. Every figure links to its source.

LMArena Textread 2026-09-26

Preference

8,528,723 votes across 409 models

Blind head-to-head votes. Measures what people prefer reading, which tracks style as well as correctness.

Artificial Analysisread 2026-09-26

General intelligence

A composite of ten evaluations Artificial Analysis runs itself, under one harness. Only fully evaluated models are indexed: rows the site stars as estimates (not every component run independently, or based on the lab’s own claims) are left out. Several components (Terminal-Bench among them) are also indexed on their own boards, so the two overlap.

GDPval-AAread 2026-09-26

Knowledge work

Pairwise-graded outputs on economically valuable knowledge-work tasks, run by Artificial Analysis and reported as an Elo with a confidence interval. Same publisher as the Intelligence Index and AA-Briefcase, so the three are one independent voice.

AA-Briefcaseread 2026-09-26

Knowledge work

Agentic knowledge-work tasks run by Artificial Analysis and reported as an Elo with a confidence interval. Same publisher as the Intelligence Index and GDPval-AA, so the three are one independent voice.

Terminal-Benchread 2026-09-26

Agentic coding

Top entry dated Sep 3, 2026

Scores a model plus its agent scaffold (Codex, Claude Code, Grok Build). The scaffold is part of the result, not a constant.

ARC-AGI-2read 2026-09-26

Reasoning

Reported per reasoning-effort tier. The same model spans a wide range across tiers; the highest published tier is indexed and recorded beside the score.

LiveBenchread 2026-09-26

General intelligence · Reasoning · Code generation · Agentic coding · Mathematics · Data analysis · Language · Instruction following

Questions are refreshed to resist contamination. The overall figure averages the publisher’s own categories, each of which is also indexed on its own.

LiveCodeBenchread 2026-09-26

Code generation

Pass@1 on LeetCode, AtCoder and Codeforces problems published after each model’s training cutoff, so it resists contamination. Reasoning-effort tier is part of the label; the highest published tier is indexed.

Aider polyglotread 2026-09-26

Agentic coding

Percent of 225 Exercism exercises solved by editing code through the aider tool, so the harness is part of the number. Effort tier is in the label; the highest published tier is indexed.

BFCLread 2026-09-26

Tool use

Overall accuracy across single-turn, multi-turn, agentic and hallucination-measurement function-calling tasks, run by the Gorilla team at UC Berkeley. Function-calling and prompting modes are one model; the better row is indexed.

OSWorldread 2026-09-26

Computer use

Success rate on 369 real desktop tasks. Entries are submitted runs, sometimes of a model inside an agent framework, which is recorded beside the score; step budgets vary by entry.

MMMU-Proread 2026-09-26

Multimodal understanding

MMMU-Pro overall accuracy. Only figures the organisers verified are indexed; starred, self-reported ones are left out. Human-expert baselines are not models and are skipped.

Kagi LLM Benchmarkread 2026-09-26

General intelligence

Accuracy on Kagi’s unpublished set of reasoning, coding and instruction-following questions, run by Kagi. Small and private by design, which resists contamination but cannot be inspected.

SimpleBenchread 2026-09-26

Reasoning

Average of five runs on a private set of common-sense trick questions. Small and run by one team; human-baseline rows are not models and are skipped.

Vals.airead 2026-09-26

Legal · Finance · Medical · Biology research · Cybersecurity · IT operations · Web research · Agentic coding · Science · Academic knowledge · Mathematics

Twenty-two evaluations Vals AI runs itself, several built with domain experts and kept private, each reported as accuracy with a confidence interval. Agentic benchmarks score a model inside a scaffold, recorded beside the figure. One publisher, so all of them are one independent voice.

JevBenchread 2026-09-26

Routing & classification · Calibration

Two of JevBench’s four axes: Intelligence and Calibration. Its headline score also weighs speed and price, which this index does not read as capability. Benchmark Heaven runs a router of its own and says so; that entry is excluded from its ranking. Half the decisions are sealed, and the public-minus-sealed gap it publishes is large for several entrants.

HELM Safetyread 2026-09-26

Harm refusal · Fairness · Safe-prompt compliance

Release v1.17.0

Five safety scenarios scored by an LM judge, on a fixed model set Stanford chooses; releases are periodic, so the newest models can lag by months. Higher is safer on every column, XSTest included.

Jailbreak resistance · Harm refusal · Toxicity avoidance · Fairness · Secure code

Automated red-teaming by one vendor of safety tooling; each column is the share of attacks that succeeded, so lower is safer. Method and prompt sets are Enkrypt’s own and not fully public.

Factual grounding

One task — summarising short documents — scored by Vectara’s own hallucination detector. A model that declines to summarise is scored on what it did answer; the answer rate is published beside the figure.

IFM — K2 Horizon launchpublished 2026-09-03

Report by Institute of Foundation Models (MBZUAI) · 27 benchmarks across 5 categories · marked ✱ wherever it appears

Vendor launch post. Competitor figures are as IFM printed them; the post does not say whether they were re-run or taken from the competitors’ own reports. Effort tiers vary by column.

Report by DataCamp · 9 benchmarks across 5 categories · marked ✱ wherever it appears

Third-party write-up. Figures are the vendors’ own as compiled by the author; some rows are one vendor’s internal evaluations with no external replication. Two models only.

Report by CellCog · 6 benchmarks across 2 categories · marked ✱ wherever it appears

Third-party write-up. Figures are from the vendors’ own reports as the author compiled them, not re-run. Effort tiers are not stated.

✱DeepSeek — V4.1 Flash model cardpublished 2026-09-10

Report by DeepSeek · 19 benchmarks across 5 categories · marked ✱ wherever it appears

Vendor model card. Competitor figures are as DeepSeek printed them; the card does not say whether they were re-run or taken from the competitors’ own reports.

Report by LargitData · 1 benchmarks across 1 categories · marked ✱ wherever it appears

One hundred routing decisions over twenty conversations the authors wrote. A model is scored on agreeing with their intended route, so the figure is "fits this routing scheme", not a general accuracy.

OpenRouterread 2026-09-26

Catalog · 219 of 815 models matched · ids, context, prices, open weights · never scored

Used as a catalog of where a model can be called and on what terms, not as a source of scores. Prices and availability are the router’s at the date read and move often. Listing here is not an endorsement of the router over the vendor’s own API.

Hugging Faceread 2026-09-26

Catalog · 219 of 815 models matched · ids, context, prices, open weights · never scored

Parameter counts are read from the weights files of the model’s repository (safetensors metadata). Only open-weight models have one; a closed model’s size is undisclosed, not estimated.

Excluded — SWE-bench VerifiedRead on 2026-09-05: the newest of its 181 entries is dated 2026-02-17. None of the models above appear on it. Including a stale board would add a column of blanks, not a signal. source

Excluded — LiteLLM — Jev vs Haiku 4.5 on request classificationRead on 2026-09-22: a careful independent test — 80 authored cases, three runs, frozen results and reproduction scripts — but it compares exactly two models. This index drops any measure with fewer than three, because standardizing a pair places the loser at 35 and the winner at 65 whether they differ by two points or thirty. The figures are worth reading at the source; they are not worth a number here. source

§ 3

Method

Six rules, each so that a number here traces to a number there, and so that thin or self-reported evidence cannot buy a rank.

  1. 01Standardize, discount uncertainty

    Every figure is z-scored across the models its source lists and mapped to 0–100, so an Elo and a pass rate share one scale. No single measure may place a model more than two standard deviations from the field it was measured against: on a board whose models sit close together, a modest lead standardizes into an enormous figure, and one of those could outweigh four boards. A figure a source prints as a risk or an error rate is read the other way round: the raw stays as printed, and the model with the least of it is the one ahead. Where a publisher prints an error bar, the figure counts for less.

  2. 02Shrink, never impute

    A missing figure is left missing. Thin evidence is pulled toward 50, so one generous board cannot lift a barely-tested model. A source comparing fewer than three models is not standardized at all — two points have no spread to place anything against.

  3. 03Weight by a published scheme

    Each board’s share of the Overall index is fixed and printed beside the table, with its reason.

  4. 04Two publishers, and half the scheme

    Overall ranks a model once two independent publishers among the boards that build the index have scored it, and at least half the weighting printed beside the table has actually been measured. Below that the prior decides more than the boards do, so the model is listed with its coverage and no number — a floor is not a placing. A domain ranks on its own board evidence.

  5. 05A domain must earn its column

    Rank by offers a domain only when several boards measure it, or one board ranks at least twenty models there. A single board’s sub-score is not offered as a peer of Coding; it stays on each model’s page.

  6. 06Reports ✱ stay out of the headline

    A ✱ figure decides nothing a board has measured: in a column, the boards’ number stands, and a report speaks only where no board has. Those rows are listed with their ✱ score and take no number. One publication is one voice per column, however many rows it printed — and an independent write-up counts for more than a vendor’s own post.

  7. 07Estimate the rest, and say so

    A model with no Overall score is placed among models that have one, on capability measures only, and shown as ~55 — never ranked. A safety figure never stands in for a capability, nor does a figure for work nothing else here measures; those models are listed with a dash, not a guess.

§ 4

Limitations

An aggregate hides the disagreements that produced it. These are the ones worth knowing before you cite this page.

Effort tiers are matched by rule

The highest tier a source publishes is indexed; the exact label is kept beside the score.

Some boards score a scaffold

Terminal-Bench results are a model plus an agent framework; the framework is named beside the score.

Identity is inferred from names

Seven sources spell one model seven ways. Every merge is logged; a wrong one is possible.

Report figures are the publisher’s

Competitor numbers in a launch post may be copied, not re-run. They are marked, not laundered.

Sizes are what makers disclose

Counted from open weights or read from the name; a closed model is undisclosed, not small.

A snapshot, not a live feed

Read from each publisher on the date shown. Check the source before acting on a number.