EvalMapBench: solved against price.

4 models on 150 private tasks from one developer’s own repositories, against what each costs per million output tokens. Up is better, left is cheaper; the dashed line joins the models nothing else beats on both.

← back to the atlas · sources and licences

$3$10$300%20%40%60%80%100%Claude Opus 5.5 (max): 97.3% (146 of 150 tasks, 95% 93–99%), $20 per 1M output tokens (AA), run through Claude CodeGPT-6 Sol (high): 94.7% (142 of 150 tasks, 95% 90–97%), $10 per 1M output tokens (AA), run through Codex CLIGemini 3.8 Flash (high): 92.0% (138 of 150 tasks, 95% 87–95%), $3.75 per 1M output tokens (AA), run through Antigravity CLIGemini 3.1 Pro Preview: 81.5% (88 of 108 tasks, 95% 73–88%), $12 per 1M output tokens (AA), run through Antigravity CLIClaude Opus 5.5 (max) 97%GPT-6 Sol (high) 95%Gemini 3.8 Flash (high) 92%Gemini 3.1 Pro Preview 81%USD per 1M output tokens, log scaletasks solved, %

One attempt per task, graded only by hidden tests; the bar through a point is the 95% Wilson interval over tasks and leaves out run-to-run spread. The price is the list price per million output tokens as the atlas shows it, not what the run cost: several rows ran on a subscription. Each row names the CLI that drove it on hover, and part of the spread between rows is that CLI rather than the model. A hollow point has a verdict on only part of the set (Gemini 3.1 Pro Preview). Measured 2026-09-26.