Request a pilot

PublicAI Index

465 models, 28 domains, one scale.

A single benchmark is easy to target and easy to overfit. This index normalizes recognised public leaderboards onto one scale, discounts each score by the uncertainty its publisher reports, shrinks thin evidence, and weights the rest by a scheme published beside the table — overall, by category, and by domain. Figures from launch posts and blogs are indexed too, marked ✱ and kept out of the headline. Every model opens into a card with how to call it, and agents can ask the same questions over MCP.

Updated
Sep 6, 2026
Boards
5
Reports ✱
1
Measures
39
Categories
6
Domains
28
Models
465
With a callable id
146
Attribution
Scores belong to their publishers; PublicAI only normalizes and weights them.

CitePublicAI Foundation (2026). PublicAI Index: a weighted aggregate of public model leaderboards, snapshot 2026-09-06. https://publicai.io/model-index

§ 1 · Ranking

Rank by
Must include

63 ranked · 402 provisional · showing 100 of 443

#ModelIndexSources

Scores are 0–100 standardized: 50 is average across the models listed, not an absolute grade. Thin evidence is pulled toward 50. A model scored by fewer than 2 recognised boards is listed as provisional without a rank; figures from reports ✱ shape category and domain columns only. A filled badge means that source scored the model; hover for the figure. Open a row for the model card: scores by domain, how to call it, and every source figure. The first 100 of 443 are shown; narrow the filters to see the rest.

§ 2

Sources

Recognised leaderboards form the Overall index. Reports ✱ widen the domain columns. A catalog says where a model can be called. All first-party; every figure links to the page it was read from.

LMLMArena Textread 2026-09-06

Preference

7,999,020 votes across 400 models

Blind head-to-head votes. Measures what people prefer reading, which tracks style as well as correctness.

AAArtificial Analysisread 2026-09-06

General intelligence

A composite of ten evaluations Artificial Analysis runs itself, under one harness. Only fully evaluated models are indexed: rows the site stars as estimates (not every component run independently, or based on the lab’s own claims) are left out. Several components (Terminal-Bench among them) are also indexed on their own boards, so the two overlap.

TBTerminal-Benchread 2026-09-06

Agentic coding

Top entry dated Sep 3, 2026

Scores a model plus its agent scaffold (Codex, Claude Code, Grok Build). The scaffold is part of the result, not a constant.

ARCARC-AGI-2read 2026-09-06

Abstract reasoning

Reported per reasoning-effort tier. The same model spans a wide range across tiers; the highest published tier is indexed and recorded beside the score.

LBLiveBenchread 2026-09-06

General · Reasoning · Code generation · Agentic coding · Mathematics · Data analysis · Language · Instruction following

Questions are refreshed to resist contamination. The overall figure averages the publisher’s own categories, each of which is also indexed on its own.

IFM — K2 Horizon launchpublished 2026-09-03

Report by Institute of Foundation Models (MBZUAI) · 27 benchmarks across 5 categories · marked ✱ wherever it appears

Vendor launch post. Competitor figures are as IFM printed them; the post does not say whether they were re-run or taken from the competitors’ own reports. Effort tiers vary by column.

CATOpenRouterread 2026-09-06

Catalog · 146 of 465 models matched · ids, context, prices, open weights · never scored

Used as a catalog of where a model can be called and on what terms, not as a source of scores. Prices and availability are the router’s at the date read and move often. Listing here is not an endorsement of the router over the vendor’s own API.

Excluded — SWE-bench VerifiedRead on 2026-09-05: the newest of its 181 entries is dated 2026-02-17. None of the models above appear on it. Including a stale board would add a column of blanks, not a signal. source

§ 3

Method

Five steps, each chosen so that a number here can be traced to a number there, and so that thin or self-reported evidence cannot buy a rank.

  1. 01Standardize per measure

    Each figure a source publishes is z-scored across the models on it and mapped to 0–100 with 50 as that measure’s average. An Elo of 1504 and a 57.9% resolution rate become comparable, and a narrow-spread board is not drowned out by a wide one.

  2. 02Discount by published uncertainty

    Where a publisher reports an error bar, the score is discounted in proportion to how wide it is relative to the board’s spread. No error bar means face value — absence is not treated as evidence of a wide one.

  3. 03Never impute; shrink instead

    A model absent from a source is excluded from that term, not filled in. But thin evidence is pulled toward 50 by a prior worth 25% of the weight available, so one generous board cannot put a barely-tested model at the top. Scored by fewer than 2 recognised boards, a model is listed as provisional and not ranked.

  4. 04Weight by a published scheme

    Each recognised board’s headline figure carries a fixed share of the Overall index, stated beside the table with its reason. A board’s category figures shape only the domain they measure, so no board is counted twice.

  5. 05Mark reports ✱ and keep them out of the headline

    A launch post or a blog is evidence of a different grade: the publisher chose the benchmarks, the settings and the comparison set. Its figures are indexed for domain and category columns at 10% of a board’s share, never enter the Overall index, and never make a model rankable.

§ 4

Limitations

An aggregate hides the disagreements that produced it. These are the ones worth knowing before you cite this page.

Reasoning-effort tiers are matched by rule

Sources report several effort tiers per model. The highest published tier is indexed, and the exact label is kept beside each score. Where a source only published a lower tier, that row is weaker evidence than it looks.

Some sources score a scaffold, not a model

Terminal-Bench results are a model plus an agent framework — Codex, Claude Code, Grok Build. The framework is part of the number and is recorded beside each score.

Model identity is inferred from names

Five sources spell the same model five ways. They are matched by name after tier and vendor words are removed. The rule is tested against hand-checked cases and every merge is logged, but a wrong merge is possible; the model card shows the exact source labels.

Report figures are the publisher’s

A launch post does not always say whether competitor figures were re-run or copied, and its columns can mix effort tiers. They are shown with their mark and their publisher so the reader can weigh them, not laundered into board figures.

Access facts are a catalog’s, not a test

Ids, context windows and prices are read from a public catalog on the date shown and move often. Vendor sites are curated pointers. None of it is an endorsement, and none of it touches a score.

A snapshot, not a live feed

Figures were read from each publisher on the date shown. Leaderboards move; check the source before acting on a number.