EvalMapBench: solved against price.
4 models on 150 private tasks from one developer’s own repositories, against what each costs per million output tokens. Up is better, left is cheaper; the dashed line joins the models nothing else beats on both.
← back to the atlas · sources and licences
One attempt per task, graded only by hidden tests; the bar through a point is the 95% Wilson interval over tasks and leaves out run-to-run spread. The price is the list price per million output tokens as the atlas shows it, not what the run cost: several rows ran on a subscription. Each row names the CLI that drove it on hover, and part of the spread between rows is that CLI rather than the model. A hollow point has a verdict on only part of the set (Gemini 3.1 Pro Preview). Measured 2026-09-26.