Sources and licences
Every number on the atlas comes from one of the sources below and is shown under its own name. Nothing is averaged or rescaled, and the coverage figures say how many of the 166 configurations each source actually reaches. Built 2026-09-23; what changed lists every change this site has been through. The list of model releases is not a source of numbers; each of its rows links to where it came from.
Artificial Analysis
- Licence
- Proprietary · no tier permits republishing these values · site terms · data platform terms
- Source
- https://artificialanalysis.ai/
- As of
- 2026-09-23
- Metrics and coverage
- tb4 166/166, speed 165/166, iq 166/166, knowAll 166/166, knowSwe 153/166, accuracy 166/166, price 166/166
- What it measures
- One independent harness runs every model, so all 166 configurations are directly comparable on its evaluations. It had not yet timed Claude Opus 5.5 at max effort on the snapshot’s day; its speed comes from EvalMap. Terminal-Bench here is AA’s own 198-trial run, not the official board. Artificial Analysis publishes two sets of terms. The site Terms of Use bind every visitor and permit personal, non-commercial access only. The Data Platform Terms bind paying subscribers and do permit public charts with attribution, but never a structured, tabular or machine-readable reproduction of individual values. Neither covers what this page does, so these numbers carry attribution as a matter of practice rather than of permission.
Terminal-Bench 4.0 official board
- Licence
- Apache-2.0 · terms
- Source
- https://www.tbench.ai/leaderboard
- As of
- 2026-09-02
- Pinned commit
b90196f7105235f869729f14aa90ed7a02bb865b- Metrics and coverage
- tb4 13/166
- What it measures
- Vendor-submitted runs on each model’s native agent (Claude Code, Codex, Grok Build, mini-SWE-agent) over 330 trials, with a published 95% interval. Only the 13 submitted configurations are covered.
DeepSWE v1.1
- Licence
- Apache-2.0 · terms
- Source
- https://deepswe.datacurve.ai/
- As of
- 2026-09-22
- Metrics and coverage
- deepswe 54/166
- What it measures
- 113 original long-horizon engineering tasks, pass@1 over four whole-benchmark passes on mini-SWE-agent, with a run-to-run 95% interval. Timeouts and context-window failures count as failures.
Epoch AI
- Licence
- CC BY 4.0 · terms
- Source
- https://epoch.ai/benchmarks
- As of
- 2026-09-22
- Metrics and coverage
- eci 101/166, gpqa 43/166, aime 45/166, fmath 35/166, simpleqa 28/166
- What it measures
- Five feeds from one publisher: the Capabilities Index plus four benchmarks Epoch runs itself — GPQA Diamond, OTIS Mock AIME, FrontierMath tiers 1-3 and SimpleQA Verified — each with the standard error of that run and the effort level it was run at. The index itself: a fitted latent-capability scale, not a percentage: anchored so Claude 3.5 Sonnet = 130 and GPT-5 = 150, unbounded above. Intervals are Epoch’s 90% bootstrap over 500 resamples, narrower in kind than the 95% intervals the benchmark boards publish. Epoch scores a model once, not once per reasoning effort.
LMArena leaderboard
- Licence
- CC BY 4.0 · terms
- Source
- https://huggingface.co/datasets/lmarena-ai/leaderboard-dataset
- As of
- 2026-09-11 to 2026-09-15
- Metrics and coverage
- arena 59/166, arenaCode 59/166, arenaMath 52/166, arenaExpert 59/166, arenaWeb 51/166, arenaFact 56/166, arenaAgent 39/166, arenaTools 39/166, arenaBash 39/166, arenaSteer 39/166
- What it measures
- Ten leaderboards of two kinds. Six are Bradley-Terry ratings from blind paired votes, and measure what people prefer rather than whether a task passed: the style-controlled text leaderboard, coding prompts only, maths prompts, expert prompts, votes scored for factuality, and the WebDev arena where two apps are built from one prompt. Four come from the agent arena, its headline score and three of its five signals (invented tools, bash recovery, steerability): each is an IPS estimate of how much a session’s outcome changes when the model runs the agent, from causal tracing of live sessions. Every value carries its 95% bounds and the number of votes, sessions or signal observations behind it; observation counts are not comparable across the signal boards, which record from under one to about fifty-five per session. EvalMap rounds the values, gives each interval as a half-width and matches the entries to configurations; nothing else is changed.
models.dev catalogue
- Licence
- MIT · terms
- Source
- https://models.dev/
- As of
- 2026-09-22
- Metrics and coverage
- ctx 144/166, price 139/166
- What it measures
- An open catalogue of what providers publish: list price, context window and the capability flags, across 222 providers. The price shown is the model owner’s own listing where there is one, because the same model is also resold and given away; the full spread travels with the value. It agrees with Artificial Analysis on the output price of most configurations, which is the point of carrying both.
LiteLLM model prices
- Licence
- MIT · terms
- Source
- https://github.com/BerriAI/litellm/blob/main/model_prices_and_context_window.json
- As of
- 2026-09-22
- Metrics and coverage
- ctx 115/166, price 112/166
- What it measures
- Cheapest listed output price across every provider LiteLLM knows serves the model. A provider count and the full spread travel with each value, because one model can be resold at very different rates.
Vals AI
- Licence
- No licence and no terms published · republished with attribution · their methodology
- Source
- https://www.vals.ai/benchmarks
- As of
- 2026-09-01 to 2026-09-21
- Metrics and coverage
- lawResearch 49/166, lawAgent 49/166, lawBench 46/166, tax 46/166, medCode 44/166, finAgent 49/166
- What it measures
- Six boards in domains nothing else on this page reaches: legal research, legal agent work, open legal reasoning, tax, medical billing codes and the work of a financial analyst. Vals runs every evaluation itself, on its own harness, and publishes the standard error and the reasoning effort of each run. Four of the six are graded against held-out sets it does not release, so those values cannot be reproduced from outside; the board says which. The benchmarks are not all theirs: the legal agent board is Harvey’s, with Harvey’s own scoring, LegalBench is Stanford CodeX’s dataset, the research rubrics were written by practising firms, and the medical coding set is Protege’s. Their cost and latency are per task rather than per million tokens or per second, so neither is carried here. Vals publishes no terms of use, no licence and no robots exclusion — checked path by path on the fetch date recorded in each extract — which is silence rather than permission, so these values carry attribution as a matter of practice.
EvalMap measurement
- Licence
- Measured here; every run is kept with the site’s sources · this page
- Source
- https://evalmap.ai/licenses/
- As of
- 2026-09-22
- Metrics and coverage
- speed 1/166
- What it measures
- Output speed EvalMap measured itself where Artificial Analysis has not yet timed a configuration. On 22 September, the day they reached AA, that was all seventeen new rows, the five Claude Opus 5.5 and the twelve GPT-6 Sol and GPT-6 Luna configurations; AA timed sixteen of them that evening and its values replaced EvalMap’s there, and on the snapshot’s day only Opus 5.5 at max effort still shows this one. Five public explainer prompts per reasoning effort, sent through each maker’s own CLI with no tools — Claude Code for Opus, Codex for Sol and Luna; the speed is the output tokens the API reports, thinking included, over the time from the first streamed chunk to the last, which is how Artificial Analysis defines it, and the value is the median of the five. It is one evening on one client, not AA’s 72-hour median, so read it as close, not as equal.
Why this exists, and what it is not
EvalMap is a reading tool: it gathers published measurements so they can be compared in one place, and recommends nothing. Nothing on the atlas is measured here; the one benchmark of EvalMap’s own is the jeep bench, whose scores stay in the jeep gallery and off the chart. Nothing here earns from the data — no advertising, no affiliate links, no sponsorship, no paid placement, nothing for sale. No source pays to appear and none can pay to move.
Nothing that identifies a visitor is collected: no cookies, no accounts, no third-party scripts or fonts, and choices are kept in the address bar so a link can be shared. Cloudflare hosts the files, and its cookieless page counter is deliberately left on so the readership is known; it reports totals only — pages, referrers, country, browser — and identifies nobody. A source owner who wants values removed or corrected only has to ask.
Contact
Questions about a value, corrections and requests from a source owner: support@evalmap.ai.
Non-endorsement
Artificial Analysis, Terminal-Bench 4.0 official board, DeepSWE v1.1, Epoch AI, LMArena leaderboard, models.dev catalogue, LiteLLM model prices, Vals AI, EvalMap measurement do not endorse EvalMap, are not affiliated with it and have not reviewed it. Their names and marks belong to them and are used here only to say where a value came from. Benchmark questions and answers remain the property of their creators, and externally sourced results keep their original terms.