Where the names come from. Each of the older builds lives in its own project named after the model that wrote it, and that name is the only record of who wrote what. Since 21 September a build comes as a project folder instead, with the word of the person who ran it for the model and the effort that wrote it. A build from EvalMap’s own bench names the model its session log records. Where the maker’s own name for the model differs, the card uses the maker’s: the stealth model Union Alpha was revealed as Unbiased’s Pareto 26.9, a blend of several models rather than one, and the build whose project says MiMo X Pro was written by what Xiaomi calls MiMo-X-Pro-Preview, a preview offered only inside its MiMo Desktop app. Omen Alpha is still a codename that no maker has claimed; its card repeats only the guess in its project’s name. The 9 builds older than the series say nothing in their project names about who wrote them, so their cards say the model was not recorded rather than guessing. Where a build states a reasoning effort or an unusual way of serving the model — a local GPU, a fast host — the card states it too.
Running here. The builds are shown here and nowhere else, served from this domain. The older ones are copies of pages their builders published themselves, and three kinds of thing were changed in them. Addresses: files that used to be fetched from a CDN or from Google Fonts are saved beside the build and pointed at locally, so nothing a build asks for is loaded from anywhere else. On-screen Russian, put into English as “In English” below describes. And a strip along the top, added to every build alike: it names the model that wrote the build, the day it went up and the way back here, because a build opened on its own says none of that. It is one block of plain markup and one stylesheet from this domain, with no script behind it and nothing kept in the browser. It moves the build down by its own height rather than covering it, so a display pinned to the top of the window sits just below the strip and controls pinned to the bottom stay on the bottom edge; on a touch screen one build, whose controls would otherwise lose their lower edge, keeps its place and the strip lies over its top. The word “hide” folds the strip away and gives the build the whole window until the page is loaded again. Since 21 September a build is published nowhere but here: it is built from its own source with its own Vite config, relative file paths the one setting added, and nothing in it is changed but that strip.
What the dates mean. The date on a card is the day that build was published, not how long the model took and not the day its maker released it. The frames were all captured later, on 23 September 2026, in one pass.
What the pictures are. Every frame was made the same way: open the copy published here in a headless browser, start the run, hold the accelerator for a few seconds, save the picture. The frames come from these copies rather than from the originals, so a tile shows what this page actually serves, English text and all. These builds ask for WebGPU, so the pictures are taken on a graphics card, an RTX 5090, in Chrome with WebGPU on and frames paced like a 60 Hz screen: each build is shown on the path it takes by itself, the one a visitor with a card gets. A prettier still is not a better model.
In English. 14 of the builds wrote their menus, timers and loading lines in Russian, and 5 named themselves in it. The copies here were put into English from a table of 406 phrases, by hand: the tool replaces text a browser shows — text in the HTML, the title, alt, placeholder and aria-label attributes, and the string literals the scripts draw their words from — and touches nothing else, so the comments, the names and the logic are still exactly what the model wrote. Every replacement is listed in sources/jeep-english.json, line by line, and the build each one landed in is counted in sources/jeep-english-log.json.
The bench. Since 22 September a build can also come from EvalMap’s own bench. The brief goes to the model it names exactly as the owner writes it — English, with one Russian phrase asking for a detailed landscape — and is never reworded. It runs unattended in OpenCode, or in Claude Code or Codex when the owner names one, in a new folder and a clean harness: none of the owner’s plugins, settings, extra tools or skills, and the tool that asks a person a question switched off, because an unattended run cannot answer it. Whatever the model hands back is published as it is; nobody fixes or finishes it. The card then says what the run took — time on the clock, tokens and dollars, from the harness’s own session records with any subagents it started counted in — and carries a short note by Claude, the agent that runs the bench, on how the run went and what came of it. The dollars are the harness’s pricing of the tokens, or the maker’s list price where the harness prices nothing: on a flat-rate plan or a subscription that is what they would cost, not what was charged. 4 of the builds came this way.
The score. Each card carries a number from 0 to 100, and so does the board at the top of this page. It is scored against one written rubric where 100 is a finished open-world driving game on the level of GTA VI — an absolute scale, not a curve over these builds, so a game written from one line of brief belongs in the thirties and anything over 60 would be a game a studio could sell. Six parts carry it: runs (10), world (25), because the brief asks for the landscape to be detailed, driving (20), race (15), look (15) and presentation (15). Each part is rated on the same ladder — rudimentary, a solid hobby browser game, a good commercial indie game, an AA production, AAA — and every card shows all six, so the total can be argued with part by part. Every build is looked at the same way and none is changed: it is started in a browser here, driven for half a minute or longer, and read where a picture cannot show whether a thing is there. It is one reader’s judgement, written down: Claude’s, the same agent that runs the bench. The rubric is tools/jeepbench/RUBRIC.md; when it changes its version changes, and scores from two versions are not compared. 41 of the 41 builds carry a score.
What it is not. No build was retried, and nothing in how one plays was fixed, or timed. The score is a judgement, not a measurement: one reader against a written ladder, with the six parts shown so the number can be argued with. It is given to the copy published here, which is all these builds have in common — so it ranks builds, not models under equal conditions. Only a bench build has a record of the harness it was written in, the effort it ran at and the brief it was given word for word; for an older build none of that is known, and two builds by the same model are only loosely comparable. The bench’s time, tokens and dollars say what a run cost, not how good its build is, and no number on the atlas comes from them.
The atlas links. 23 of the 41 cards name a model that is also among the configurations on the atlas, and link to the row there. The rest do not: a codename, preview releases, an agent, a blend of several models, a model none of the atlas’s sources covers, and the older builds whose author is unknown. A card links to its model, which the published benchmarks may have measured at a different reasoning effort than the build used.
Nothing that identifies you. This page keeps nothing in your browser and loads no third-party script of its own and no web font. Every file a build asks for comes from this domain: the CDN copies of three.js and the web fonts two builds used are saved here beside them. The exception is the one the atlas discloses too — Cloudflare’s zone adds its cookieless page counter (cloudflareinsights.com) to every HTML page on this site, these builds included; it sets no cookie and reports totals only. The builds are other people’s code rather than this site’s: they draw to a canvas and three of them remember something in the browser’s own storage — settings in one, a best run in the other two, and none of them sends anything anywhere.