{"name":"PublicAI Index","url":"https://publicai.io/model-index","mcp":"https://publicai.io/model-index/mcp","api":"https://publicai.io/model-index/api","feed":"https://publicai.io/model-index/feed.xml","generatedAt":"2026-09-26T06:03:03.526Z","counts":{"models":815,"ranked":57,"boards":19,"reports":5,"measures":119,"categories":9,"domains":26},"method":["Each measure is z-scored across the models its source lists and mapped to 0–100 (mean 50, sd 15), with z capped at ±2 (so 20–80): past two standard deviations a board has stopped discriminating and started extrapolating, and one such figure could outweigh four boards. The cap changes no board’s own order. A measure marked direction \"lower\" (a risk or error rate) is read the other way round; its raw stays as printed.","Where a source publishes an error bar, the figure’s weight is discounted by how wide it is relative to the board’s spread.","A model absent from a source is excluded from that term, never imputed; a prior worth 25% of the in-scope weight pulls thin evidence toward 50.","Overall ranks a model once 2 independent publishers among the boards that build the index have scored it — a board that shapes only a domain does not count toward it; otherwise the model is provisional. A category or domain ranks any model a recognised board has measured there.","In a category or domain, position counts only models a recognised board measured there; a model placed by report ✱ figures alone keeps its score and takes no number (position null). Filters hide rows without renumbering the rest.","Reports (launch posts, blogs) are marked ✱. They never enter the Overall index, and in a category or domain they decide nothing a recognised board measured: the boards' number stands, and a ✱ figure sets the score only for a model no board measured there (which takes no position). Pass reports=true to let ✱ figures into every score instead. An independent write-up carries 12 against a board measure's 25; a figure the model's own publisher printed carries 10. One publication is one voice in a scope: the measures it contributes there are averaged and enter once, so a post that printed five numbers does not outweigh a board that ran a fixed set on everyone.","scopes lists the domains offered as headline rankings: measured by at least 2 boards, or by one board that ranks at least 20 models there. Every other domain a board measured is still on the model pages and accepted as a scope by name.","A model with no Overall index gets an estimate ✱ (estimatedIndex): its figures on each shared capability measure are placed among models that have an index, and theirs is read at that position; outside their range the nearest anchor is a bound (ceiling or floor), not a point. Safety measures never anchor an estimate. Never a rank."],"overallWeighting":[{"source":"LMArena Text","weight":25,"rationale":"Largest evidence base of any board — millions of blind votes across hundreds of models — and the widest coverage. Discounted because preference tracks style as well as correctness.","url":"https://arena.ai/leaderboard/text"},{"source":"Artificial Analysis","weight":25,"rationale":"Ten evaluations across agents, coding, reasoning and knowledge, run by one third party under one harness on ~300 models, with a confidence interval under ±1%. Discounted because it is itself a composite whose components overlap other boards here.","url":"https://artificialanalysis.ai/leaderboards/models"},{"source":"Terminal-Bench","weight":15,"rationale":"Measures completed agentic work, not answers, and publishes error bars. Discounted because the agent scaffold is part of the score and the board is small.","url":"https://www.tbench.ai/leaderboard"},{"source":"ARC-AGI-2","weight":15,"rationale":"Hard and far from saturated, which keeps it discriminating at the top. Discounted for being one narrow skill and highly sensitive to reasoning-effort tier.","url":"https://arcprize.org/leaderboard"},{"source":"LiveBench","weight":20,"rationale":"Broad, refreshed to resist contamination, and covers reasoning, coding, maths and instruction following in one place.","url":"https://livebench.ai/"}],"sources":[{"id":"lmarena","name":"LMArena Text","kind":"board","publisher":"LMArena","url":"https://arena.ai/leaderboard/text","retrievedAt":"2026-09-26","measures":[{"name":"LMArena Text","category":"Human preference","domain":"Human preference"}],"caveat":"Blind head-to-head votes. Measures what people prefer reading, which tracks style as well as correctness."},{"id":"artificial-analysis","name":"Artificial Analysis","kind":"board","publisher":"Artificial Analysis","url":"https://artificialanalysis.ai/leaderboards/models","retrievedAt":"2026-09-26","measures":[{"name":"Artificial Analysis Intelligence Index","category":"Core abilities","domain":"General intelligence"}],"caveat":"A composite of ten evaluations Artificial Analysis runs itself, under one harness. Only fully evaluated models are indexed: rows the site stars as estimates (not every component run independently, or based on the lab’s own claims) are left out. Several components (Terminal-Bench among them) are also indexed on their own boards, so the two overlap."},{"id":"aa-gdpval","name":"GDPval-AA","kind":"board","publisher":"Artificial Analysis","url":"https://artificialanalysis.ai/evaluations/gdpval-aa","retrievedAt":"2026-09-26","measures":[{"name":"GDPval-AA","category":"Agents","domain":"Knowledge work"}],"caveat":"Pairwise-graded outputs on economically valuable knowledge-work tasks, run by Artificial Analysis and reported as an Elo with a confidence interval. Same publisher as the Intelligence Index and AA-Briefcase, so the three are one independent voice."},{"id":"aa-briefcase","name":"AA-Briefcase","kind":"board","publisher":"Artificial Analysis","url":"https://artificialanalysis.ai/evaluations/aa-briefcase","retrievedAt":"2026-09-26","measures":[{"name":"AA-Briefcase","category":"Agents","domain":"Knowledge work"}],"caveat":"Agentic knowledge-work tasks run by Artificial Analysis and reported as an Elo with a confidence interval. Same publisher as the Intelligence Index and GDPval-AA, so the three are one independent voice."},{"id":"terminal-bench","name":"Terminal-Bench","kind":"board","publisher":"Terminal-Bench","url":"https://www.tbench.ai/leaderboard","retrievedAt":"2026-09-26","measures":[{"name":"Terminal-Bench","category":"Coding","domain":"Agentic coding"}],"caveat":"Scores a model plus its agent scaffold (Codex, Claude Code, Grok Build). The scaffold is part of the result, not a constant."},{"id":"arc-agi-2","name":"ARC-AGI-2","kind":"board","publisher":"ARC Prize","url":"https://arcprize.org/leaderboard","retrievedAt":"2026-09-26","measures":[{"name":"ARC-AGI-2","category":"Reasoning","domain":"Reasoning"}],"caveat":"Reported per reasoning-effort tier. The same model spans a wide range across tiers; the highest published tier is indexed and recorded beside the score."},{"id":"livebench","name":"LiveBench","kind":"board","publisher":"LiveBench","url":"https://livebench.ai/","retrievedAt":"2026-09-26","measures":[{"name":"LiveBench","category":"Core abilities","domain":"General intelligence"},{"name":"LiveBench · Reasoning","category":"Reasoning","domain":"Reasoning"},{"name":"LiveBench · Coding","category":"Coding","domain":"Code generation"},{"name":"LiveBench · Agentic Coding","category":"Coding","domain":"Agentic coding"},{"name":"LiveBench · Mathematics","category":"Reasoning","domain":"Mathematics"},{"name":"LiveBench · Data Analysis","category":"Core abilities","domain":"Data analysis"},{"name":"LiveBench · Language","category":"Core abilities","domain":"Language"},{"name":"LiveBench · Instruction Following","category":"Core abilities","domain":"Instruction following"}],"caveat":"Questions are refreshed to resist contamination. The overall figure averages the publisher’s own categories, each of which is also indexed on its own."},{"id":"livecodebench","name":"LiveCodeBench","kind":"board","publisher":"LiveCodeBench","url":"https://livecodebench.github.io/leaderboard.html","retrievedAt":"2026-09-26","measures":[{"name":"LiveCodeBench","category":"Coding","domain":"Code generation"}],"caveat":"Pass@1 on LeetCode, AtCoder and Codeforces problems published after each model’s training cutoff, so it resists contamination. Reasoning-effort tier is part of the label; the highest published tier is indexed."},{"id":"aider","name":"Aider polyglot","kind":"board","publisher":"Aider","url":"https://aider.chat/docs/leaderboards/","retrievedAt":"2026-09-26","measures":[{"name":"Aider polyglot","category":"Coding","domain":"Agentic coding"}],"caveat":"Percent of 225 Exercism exercises solved by editing code through the aider tool, so the harness is part of the number. Effort tier is in the label; the highest published tier is indexed."},{"id":"bfcl","name":"BFCL","kind":"board","publisher":"UC Berkeley (Gorilla)","url":"https://gorilla.cs.berkeley.edu/leaderboard.html","retrievedAt":"2026-09-26","measures":[{"name":"BFCL v4","category":"Agents","domain":"Tool use"}],"caveat":"Overall accuracy across single-turn, multi-turn, agentic and hallucination-measurement function-calling tasks, run by the Gorilla team at UC Berkeley. Function-calling and prompting modes are one model; the better row is indexed."},{"id":"osworld","name":"OSWorld","kind":"board","publisher":"XLANG Lab","url":"https://osworld-v1.xlang.ai/","retrievedAt":"2026-09-26","measures":[{"name":"OSWorld","category":"Agents","domain":"Computer use"}],"caveat":"Success rate on 369 real desktop tasks. Entries are submitted runs, sometimes of a model inside an agent framework, which is recorded beside the score; step budgets vary by entry."},{"id":"mmmu","name":"MMMU-Pro","kind":"board","publisher":"MMMU","url":"https://mmmu-benchmark.github.io/","retrievedAt":"2026-09-26","measures":[{"name":"MMMU-Pro","category":"Knowledge","domain":"Multimodal understanding"}],"caveat":"MMMU-Pro overall accuracy. Only figures the organisers verified are indexed; starred, self-reported ones are left out. Human-expert baselines are not models and are skipped."},{"id":"kagi","name":"Kagi LLM Benchmark","kind":"board","publisher":"Kagi","url":"https://help.kagi.com/kagi/ai/llm-benchmark.html","retrievedAt":"2026-09-26","measures":[{"name":"Kagi LLM Benchmark","category":"Core abilities","domain":"General intelligence"}],"caveat":"Accuracy on Kagi’s unpublished set of reasoning, coding and instruction-following questions, run by Kagi. Small and private by design, which resists contamination but cannot be inspected."},{"id":"simplebench","name":"SimpleBench","kind":"board","publisher":"SimpleBench","url":"https://simple-bench.com/","retrievedAt":"2026-09-26","measures":[{"name":"SimpleBench","category":"Reasoning","domain":"Reasoning"}],"caveat":"Average of five runs on a private set of common-sense trick questions. Small and run by one team; human-baseline rows are not models and are skipped."},{"id":"vals","name":"Vals.ai","kind":"board","publisher":"Vals AI","url":"https://www.vals.ai/benchmarks","retrievedAt":"2026-09-26","measures":[{"name":"Vals · Legal Research Bench","category":"Professional","domain":"Legal"},{"name":"Vals · CaseLaw","category":"Professional","domain":"Legal"},{"name":"Vals · LegalBench","category":"Professional","domain":"Legal"},{"name":"Vals · Harvey Legal Agent Benchmark","category":"Professional","domain":"Legal"},{"name":"Vals · Finance Agent","category":"Professional","domain":"Finance"},{"name":"Vals · CorpFin","category":"Professional","domain":"Finance"},{"name":"Vals · TaxEval","category":"Professional","domain":"Finance"},{"name":"Vals · MortgageTax","category":"Professional","domain":"Finance"},{"name":"Vals · MedQA","category":"Professional","domain":"Medical"},{"name":"Vals · MedCode","category":"Professional","domain":"Medical"},{"name":"Vals · MedScribe","category":"Professional","domain":"Medical"},{"name":"Vals · BioMysteryBench","category":"Professional","domain":"Biology research"},{"name":"Vals · CyberBench","category":"Professional","domain":"Cybersecurity"},{"name":"Vals · SRE Bench","category":"Professional","domain":"IT operations"},{"name":"Vals · Web Search Index","category":"Agents","domain":"Web research"},{"name":"Vals · SWE-bench Verified","category":"Coding","domain":"Agentic coding"},{"name":"Vals · Vibe Code Bench","category":"Coding","domain":"Agentic coding"},{"name":"Vals · Code Migration","category":"Coding","domain":"Agentic coding"},{"name":"Vals · GPQA Diamond","category":"Reasoning","domain":"Science"},{"name":"Vals · MMLU Pro","category":"Knowledge","domain":"Academic knowledge"},{"name":"Vals · AIME","category":"Reasoning","domain":"Mathematics"},{"name":"Vals · ProofBench","category":"Reasoning","domain":"Mathematics"}],"caveat":"Twenty-two evaluations Vals AI runs itself, several built with domain experts and kept private, each reported as accuracy with a confidence interval. Agentic benchmarks score a model inside a scaffold, recorded beside the figure. One publisher, so all of them are one independent voice."},{"id":"jevbench","name":"JevBench","kind":"board","publisher":"Benchmark Heaven","url":"https://benchmarkheaven.com/jev-models","retrievedAt":"2026-09-26","measures":[{"name":"JevBench · Intelligence","category":"Decisions","domain":"Routing & classification"},{"name":"JevBench · Calibration","category":"Decisions","domain":"Calibration"}],"caveat":"Two of JevBench’s four axes: Intelligence and Calibration. Its headline score also weighs speed and price, which this index does not read as capability. Benchmark Heaven runs a router of its own and says so; that entry is excluded from its ranking. Half the decisions are sealed, and the public-minus-sealed gap it publishes is large for several entrants."},{"id":"helm-safety","name":"HELM Safety","kind":"board","publisher":"Stanford CRFM","url":"https://crfm.stanford.edu/helm/safety/latest/","retrievedAt":"2026-09-26","measures":[{"name":"HELM Safety · HarmBench","category":"Safety","domain":"Harm refusal"},{"name":"HELM Safety · SimpleSafetyTests","category":"Safety","domain":"Harm refusal"},{"name":"HELM Safety · Anthropic Red Team","category":"Safety","domain":"Harm refusal"},{"name":"HELM Safety · BBQ","category":"Safety","domain":"Fairness"},{"name":"HELM Safety · XSTest","category":"Safety","domain":"Safe-prompt compliance"}],"caveat":"Five safety scenarios scored by an LM judge, on a fixed model set Stanford chooses; releases are periodic, so the newest models can lag by months. Higher is safer on every column, XSTest included."},{"id":"enkrypt","name":"Enkrypt AI Safety Leaderboard","kind":"board","publisher":"Enkrypt AI","url":"https://leaderboard.enkryptai.com/","retrievedAt":"2026-09-26","measures":[{"name":"Enkrypt · Jailbreak risk","category":"Safety","domain":"Jailbreak resistance"},{"name":"Enkrypt · Harmful content risk","category":"Safety","domain":"Harm refusal"},{"name":"Enkrypt · CBRN risk","category":"Safety","domain":"Harm refusal"},{"name":"Enkrypt · Toxicity risk","category":"Safety","domain":"Toxicity avoidance"},{"name":"Enkrypt · Bias risk","category":"Safety","domain":"Fairness"},{"name":"Enkrypt · Insecure code risk","category":"Safety","domain":"Secure code"}],"caveat":"Automated red-teaming by one vendor of safety tooling; each column is the share of attacks that succeeded, so lower is safer. Method and prompt sets are Enkrypt’s own and not fully public."},{"id":"vectara","name":"Vectara Hallucination Leaderboard","kind":"board","publisher":"Vectara","url":"https://github.com/vectara/hallucination-leaderboard","retrievedAt":"2026-09-26","measures":[{"name":"Vectara · Factual consistency","category":"Safety","domain":"Factual grounding"}],"caveat":"One task — summarising short documents — scored by Vectara’s own hallucination detector. A model that declines to summarise is scored on what it did answer; the answer rate is published beside the figure."},{"id":"ifm-k2-horizon","name":"IFM — K2 Horizon launch","kind":"report","publisher":"Institute of Foundation Models (MBZUAI)","url":"https://ifm.ai/blog/k2/","retrievedAt":"2026-09-26","publishedAt":"2026-09-03","measures":[{"name":"GDPVal-AA","category":"Agents","domain":"Knowledge work"},{"name":"tau3-Banking","category":"Agents","domain":"Tool use"},{"name":"Toolathlon Verified","category":"Agents","domain":"Tool use"},{"name":"Automation Bench Public","category":"Agents","domain":"Workflow automation"},{"name":"Apex-Agents (pass@1)","category":"Agents","domain":"Workflow automation"},{"name":"MCPMark","category":"Agents","domain":"Tool use"},{"name":"BrowseComp","category":"Agents","domain":"Web research"},{"name":"WildClawBench","category":"Agents","domain":"Workflow automation"},{"name":"Terminal-Bench 2.1","category":"Coding","domain":"Agentic coding"},{"name":"SciCode","category":"Coding","domain":"Code generation"},{"name":"SWE-Atlas-QnA","category":"Coding","domain":"Repository Q&A"},{"name":"SWE Bench Pro","category":"Coding","domain":"Agentic coding"},{"name":"Humanity's Last Exam (without tools)","category":"Reasoning","domain":"Expert reasoning"},{"name":"GPQA Diamond","category":"Reasoning","domain":"Science"},{"name":"CritPt","category":"Reasoning","domain":"Science"},{"name":"AA-LCR","category":"Core abilities","domain":"Long context"},{"name":"AA-Omniscience Accuracy","category":"Knowledge","domain":"Factuality"},{"name":"AA-Omniscience Non-Hallucination","category":"Knowledge","domain":"Factuality"},{"name":"SWE-bench Verified","category":"Coding","domain":"Agentic coding"},{"name":"HMMT Feb 2026","category":"Reasoning","domain":"Mathematics"},{"name":"HLE","category":"Reasoning","domain":"Expert reasoning"},{"name":"BFCL v4","category":"Agents","domain":"Tool use"},{"name":"HumanEval+","category":"Coding","domain":"Code generation"},{"name":"MBPP+","category":"Coding","domain":"Code generation"},{"name":"AIME 2025","category":"Reasoning","domain":"Mathematics"},{"name":"AIME 2026","category":"Reasoning","domain":"Mathematics"},{"name":"LiveCodeBench v6","category":"Coding","domain":"Code generation"}],"caveat":"Vendor launch post. Competitor figures are as IFM printed them; the post does not say whether they were re-run or taken from the competitors’ own reports. Effort tiers vary by column."},{"id":"datacamp-astra-vs-fable","name":"DataCamp — GPT-6 Astra vs Claude Fable 5.1","kind":"report","publisher":"DataCamp","url":"https://www.datacamp.com/blog/gpt-6-astra-vs-claude-fable-5-1","retrievedAt":"2026-09-26","publishedAt":"2026-09-05","measures":[{"name":"FrontierMath Tier 4 (v2)","category":"Reasoning","domain":"Mathematics"},{"name":"Humanity's Last Exam, with tools","category":"Reasoning","domain":"Expert reasoning"},{"name":"AutomationBench","category":"Agents","domain":"Workflow automation"},{"name":"ExploitBench","category":"Professional","domain":"Cybersecurity"},{"name":"AA Intelligence Index (max effort)","category":"Core abilities","domain":"General intelligence"},{"name":"Terminal-Bench 4.0","category":"Coding","domain":"Agentic coding"},{"name":"DeepSWE v1.1","category":"Coding","domain":"Agentic coding"},{"name":"FrontierCode 1.1 Main","category":"Coding","domain":"Agentic coding"},{"name":"Internal database migration","category":"Coding","domain":"Agentic coding"}],"caveat":"Third-party write-up. Figures are the vendors’ own as compiled by the author; some rows are one vendor’s internal evaluations with no external replication. Two models only."},{"id":"cellcog-k2-horizon","name":"CellCog — K2 Horizon analysis","kind":"report","publisher":"CellCog","url":"https://cellcog.ai/blog/k2-horizon/","retrievedAt":"2026-09-26","publishedAt":"2026-09-04","measures":[{"name":"GDPVal-AA (Elo)","category":"Agents","domain":"Knowledge work"},{"name":"Toolathlon Verified","category":"Agents","domain":"Tool use"},{"name":"Terminal-Bench 2.1","category":"Coding","domain":"Agentic coding"},{"name":"SWE-bench Pro (strict)","category":"Coding","domain":"Agentic coding"},{"name":"MCPMark","category":"Agents","domain":"Tool use"},{"name":"SWE-Atlas-QnA (strict)","category":"Coding","domain":"Repository Q&A"}],"caveat":"Third-party write-up. Figures are from the vendors’ own reports as the author compiled them, not re-run. Effort tiers are not stated."},{"id":"deepseek-v4-1-flash","name":"DeepSeek — V4.1 Flash model card","kind":"report","publisher":"DeepSeek","url":"https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash","retrievedAt":"2026-09-26","publishedAt":"2026-09-10","measures":[{"name":"GPQA Diamond (Pass@1)","category":"Reasoning","domain":"Science"},{"name":"HLE (Pass@1)","category":"Reasoning","domain":"Expert reasoning"},{"name":"Codeforces (Rating)","category":"Coding","domain":"Code generation"},{"name":"MathArena Apex (Pass@1)","category":"Reasoning","domain":"Mathematics"},{"name":"Terminal-Bench 2.1 (Pass@1)","category":"Coding","domain":"Agentic coding"},{"name":"Terminal-Bench 3.0 (Pass@1)","category":"Coding","domain":"Agentic coding"},{"name":"Terminal-Bench 4.0 (Pass@1)","category":"Coding","domain":"Agentic coding"},{"name":"DeepSWE v1.1 (Resolved)","category":"Coding","domain":"Agentic coding"},{"name":"ProgramBench (Almost@1)","category":"Coding","domain":"Agentic coding"},{"name":"NL2Repo-Bench (Score)","category":"Coding","domain":"Repository Q&A"},{"name":"CyberGym (Pass@1)","category":"Professional","domain":"Cybersecurity"},{"name":"SEC-Bench Pro (Pass@1)","category":"Professional","domain":"Cybersecurity"},{"name":"ExploitGym (Pass@1)","category":"Professional","domain":"Cybersecurity"},{"name":"HLE w/ tools (Pass@1)","category":"Reasoning","domain":"Expert reasoning"},{"name":"AutomationBench (Pass@1)","category":"Agents","domain":"Workflow automation"},{"name":"Agent's Last Exam (Pass@1)","category":"Reasoning","domain":"Expert reasoning"},{"name":"Chartography w/ tools (Pass@1)","category":"Knowledge","domain":"Multimodal understanding"},{"name":"BabyVision w/ tools (Pass@1)","category":"Knowledge","domain":"Multimodal understanding"},{"name":"ZeroBench-main w/ tools (Pass@5)","category":"Knowledge","domain":"Multimodal understanding"}],"caveat":"Vendor model card. Competitor figures are as DeepSeek printed them; the card does not say whether they were re-run or taken from the competitors’ own reports."},{"id":"largitdata-jev","name":"LargitData — Jev on multi-turn RAG routing","kind":"report","publisher":"LargitData","url":"https://www.largitdata.com/en/blog/jev-system-one-model-open-source-benchmark/","retrievedAt":"2026-09-26","publishedAt":"2026-09-19","measures":[{"name":"Decision correct","category":"Decisions","domain":"Routing & classification"}],"caveat":"One hundred routing decisions over twenty conversations the authors wrote. A model is scored on agreeing with their intended route, so the figure is \"fits this routing scheme\", not a general accuracy."}],"excluded":[{"name":"SWE-bench Verified","url":"https://www.swebench.com/","reason":"Read on 2026-09-05: the newest of its 181 entries is dated 2026-02-17. None of the models above appear on it. Including a stale board would add a column of blanks, not a signal."},{"name":"LiteLLM — Jev vs Haiku 4.5 on request classification","url":"https://docs.litellm.ai/blog/jev-auto-router-benchmark","reason":"Read on 2026-09-22: a careful independent test — 80 authored cases, three runs, frozen results and reproduction scripts — but it compares exactly two models. This index drops any measure with fewer than three, because standardizing a pair places the loser at 35 and the winner at 65 whether they differ by two points or thirty. The figures are worth reading at the source; they are not worth a number here."}],"catalogs":[{"id":"openrouter","name":"OpenRouter","publisher":"OpenRouter","url":"https://openrouter.ai/models","retrievedAt":"2026-09-26","caveat":"Used as a catalog of where a model can be called and on what terms, not as a source of scores. Prices and availability are the router’s at the date read and move often. Listing here is not an endorsement of the router over the vendor’s own API."},{"id":"huggingface","name":"Hugging Face","publisher":"Hugging Face","url":"https://huggingface.co","retrievedAt":"2026-09-26","caveat":"Parameter counts are read from the weights files of the model’s repository (safetensors metadata). Only open-weight models have one; a closed model’s size is undisclosed, not estimated."}],"scopes":[{"category":"Human preference","domains":["Human preference"]},{"category":"Agents","domains":["Knowledge work","Tool use","Computer use"]},{"category":"Professional","domains":["Legal","Finance","Medical"]},{"category":"Coding","domains":["Agentic coding","Code generation"]},{"category":"Reasoning","domains":["Reasoning","Mathematics","Science"]},{"category":"Knowledge","domains":["Academic knowledge"]},{"category":"Decisions","domains":["Routing & classification","Calibration"]},{"category":"Core abilities","domains":["General intelligence","Data analysis","Language","Instruction following"]},{"category":"Safety","domains":["Harm refusal","Fairness","Safe-prompt compliance","Jailbreak resistance","Toxicity avoidance","Secure code","Factual grounding"]}],"access":"Recommended channel: the OpenRouter id through its OpenAI-compatible endpoint — one key, every listed model, ids stable across vendors. Use the vendor’s own API for first-party features, and the weights for self-hosting.","limits":["Reasoning-effort tiers are matched by rule; the highest published tier is indexed and the exact label kept.","Some sources score a scaffold (model + agent framework), recorded beside the figure.","Model identity is inferred from names; every merge is logged, but a wrong merge is possible.","Report figures are the publisher’s own and may mix effort tiers or copied competitor numbers.","A snapshot, not a live feed: check the source before acting on a number."]}