Small tests, run here.
When a model arrives that no public board measures, or makes a claim nobody has checked, EvalMap sometimes asks the question itself. These are those answers: small, dated, each against the model it is sold as better than or cheaper than, and each saying how large its sample is.
Nothing on this page reaches the atlas. The prompts, scenarios and answers are not published, so that the tests can be run again on the next model without it having seen them; what leaves is these numbers and a description of the method. The best value in each column is shaded; where every row ties, none is.
Aion 3.5 against the GLM it is built onRecraft V4.1 Flash against Recraft V4.1
Aion 3.5 against the GLM it is built on: a long role-play in Russian
Aion Labs sells Aion 3.5 as a role-play system built on GLM, so the question is whether it keeps a long game straighter than GLM-5.3 does by itself.
| Modelroute | Memory in play−1 to +1, 170 questions a model | Hard questions at the end−1 to +1, 73 questions a model | Turns breaking no rule216 turns a model | Repeated four-word runs | Distinct word pairs | Seconds a turnmedian | Reasoning tokens a turnmean | $ a 36-turn game |
|---|---|---|---|---|---|---|---|---|
| Aion 3.5openrouter/aion-labs/aion-3.5 | +0.99 | +0.86 | 91% | 0.023 | 0.956 | 16.0 | 272 | $0.82 |
| Aion 3.5 Miniopenrouter/aion-labs/aion-3.5-mini | +0.99 | +0.97 | 96% | 0.028 | 0.951 | 31.1 | 1,374 | $0.24 |
| GLM-5.3openrouter/z-ai/glm-5.3 | +0.95 | +0.92 | 93% | 0.013 | 0.971 | 9.6 | 703 | $0.34 |
| GLM-5.3-Flashopenrouter/z-ai/glm-5.3-flash | +0.98 | +0.97 | 89% | 0.010 | 0.968 | 15.6 | 658 | $0.040 |
What it says
- Keeping track of a game is solved at this length (about 25,000 tokens by the last turn): every model answers nearly every memory question right. Aion's wrapper around several models does not cost it consistency, and does not buy any either.
- The hard questions are where they part, and Aion 3.5 comes last there: its misses are totals and look-ups, and in the one case read closely it had kept the purse right in the story and got the sum wrong when asked. Every miss was read by hand; none is the grader's.
- Aion 3.5 Mini keeps the rules best and is the slowest. Both Aion models repeat themselves a little more than GLM does.
- GLM-5.3-Flash is as good as any of them on memory at a nineteenth of Aion 3.5's price.
- Six games are a small sample: differences of 0.02–0.03 on the indices are noise.
How it was run
- Six generated scenarios of 36 turns each, in Russian. The player's turns are written in advance, so every model plays the same game, and every change to the world — who holds what, who is hurt, what was paid — is made by the player, so the true state of the world is known at every turn before any model plays it.
- Memory: after turns 12, 24 and 36 an out-of-game question is asked on a copy of the conversation, never in the game itself, about the state of the world and the town the system prompt describes. +1 for a right answer, −1 for a wrong one, 0 for “I don’t know”, as in AA-Omniscience. The first value an answer names is the one scored.
- Hard questions: once, after the last turn, questions about the whole game rather than its present — who held an item before its last hand-over, how many times it changed hands, totals paid and earned, who was hurt first — and look-ups across the town description. Scored the same way.
- Rules of the table: eight rules the system prompt sets — format, length, language and the speech habits of named characters — each checked by a script on every turn. A clean turn breaks none of them.
- Prose: distinct word pairs within windows of the narration, and how often a four-word sequence comes back from an earlier turn.
- Every model through OpenRouter with prompt logging refused (data_collection: deny); Aion offers no zero-retention route. Up to 8,000 output tokens a turn. Fourteen questions a model were voided because the scenario generator, not the model, made their answer ambiguous; the same fourteen for every model.
- The whole test cost $9.00 in API calls.
Recraft V4.1 Flash against Recraft V4.1: exact text and composition
Flash is sold as the fast, cheap V4.1, so the question is what the fifth of the price gives up.
| Modelroute | Exact text, Latinfloor / ceiling / estimate, 64 images | Exact text, Cyrillicfloor / ceiling / estimate, 64 images | Text of 1–2 wordsceiling | Text of 3–6 wordsceiling | Text of 8–11 wordsceiling | Composition, all128 images | Position | Two colours bound | Counting, measured / audited | Seconds an image | $ an image | $ a passing image |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Recraft V4.1vercel/recraft/recraft-v4.1 | 73 / 86 / 83% | 23 / 81 / 68% | 100% | 90% | 67% | 0.91 | 0.96 | 0.88 | 0.75 / 0.94 | 6.8 | $0.035 | $0.040 |
| Recraft V4.1 Flashvercel/recraft/recraft-v4.1-flash | 64 / 83 / 79% | 34 / 67 / 60% | 100% | 79% | 54% | 0.89 | 0.88 | 0.88 | 0.75 / 0.97 | 2.2 | $0.007 | $0.0085 |
What it says
- V4.1 writes long strings and Cyrillic better and places objects more reliably. On one or two words, on a single colour and on two objects side by side the two do not differ.
- Flash gives up little for a fifth of the price and a third of the time: a passing image costs $0.0085 against $0.040.
- Counting is not a difference between them: both fail it as measured for the detector's sake, and on audit both count nearly every image right.
- 32 prompts per part: a difference of one or two prompts is inside the noise.
How it was run
- 64 generated prompts, four images each at 1024×1024 through the Vercel AI Gateway, 256 scored images a model.
- Text: 32 prompts ask for a sign, label, cover or poster carrying an exact string of one to eleven words, half in Latin and half in Cyrillic script, some with a number. An image passes when the string is on it exactly, ignoring case. Read by two OCR engines (EasyOCR, PaddleOCR) and a vision model asked to transcribe without seeing the target. OCR misses decorative lettering and the vision model corrects misspellings toward real words — of nine passes only it saw, seven were right on inspection — so the share is given as a range: passed by a classical OCR (floor), by any reader (ceiling), and the estimate that counts the vision model's passes at seven in nine. The same engines read the same strings drawn in a plain font 32 times out of 32.
- Composition: 32 prompts in the manner of GenEval — counts, two objects, a colour, two colours bound to two objects, and position — checked with a Mask2Former detector; a colour passes when either the pixel colour of the object's mask or CLIP on the masked object names it.
- The detector undercounts: looked at one by one, about 13 of the 16 failed counting images show the right number of objects — cats, stone benches and vases the detector did not see. Counting is given both as measured and as audited.
- The whole test cost $10.92 in API calls.