We Raced AI Models to Find One Small Thing in a Huge Picture. Cutting It Up More Than Doubled What They Found.

Give an AI model a 4,000-pixel picture and ask it to find one small thing, and it mostly fails: it shrinks the picture before looking, and the thing disappears. Cut the picture into tiles, ask about every tile at once, and the same model finds more than twice as many. We raced LLMs and decision models through 79 full-size scenes to see who finds the target, how fast, and at what cost.
Watch them race
Each lane is a real run, replayed on its real clock. The box above the lanes says what the models are looking for, and the line under each name says what it is doing at that moment: checking tiles, comparing its best claims, or zooming in. Blue cells are questions still waiting for an answer, dark cells came back "not here", and amber cells claimed "found it". At the end, a green ring means the answer was right and a red cross means it was wrong; the white box is where the target really is. The fastest run on each scene is first. Pick another scene, hide models below, or save the race as a GIF.
Where's Waldo? (Where's Wally? in the UK) pages © Martin Handford, published by Walker Books and Candlewick Press. Shown here, from public research datasets, for research and commentary on how AI models search pictures; not endorsed by the rights holders. The letter and shape puzzles are ours.
The short version
- Gemini 3.5 Flash Lite found the most: 66 of 79 (84%) with 512-pixel tiles, at $0.0155 a scene. Smaller tiles found one more, at nearly four times the cost and almost twice the time.
- Cutting the picture up is what works. Shown each whole scene at once (in a pilot run earlier the same day), the same model found 26 of 79.
- The decision models answered fastest but found fewer: GPT-6 Luna Decisions and pplx-decider v1.1 each found 31 of 79, in a median 3.4 and 3.9 seconds a scene.
- On the same search, the decision models are fastest and cheapest. Run through the same yes-or-no zoom, they found 39% in under 4 seconds a scene. Only Claude Haiku 5.5 found more (48%), taking 16 seconds; Gemini 3.5 Flash Lite found 33% and GPT-6 Luna 22%. The LLMs' real edge is that they can point.
- pplx-decider v1.1 is by far the cheapest per find: $0.0004 a scene, about a tenth of a cent per target found, less than a third of the cheapest LLM run.
- Waldo is hard for everyone. The best run found him on 19 of 29 pages; Cloudflare's two decision models found one target between them in 158 tries.
How the models searched, and why LLMs and decision models search differently
Every model got the same scenes and the same target, and had to give one point; it counts if the point lands on the target (its box, plus 10% of its size). How a model can get to that point depends on what kind of answer it can give.
LLMs can say where
An LLM writes text, so it can answer "where?" directly: "yes, about a third of the way across and near the bottom". That lets it find the target in one step per piece of the picture.
- The whole picture at once. One question about the whole scene. Cheap, but the model shrinks a big picture before looking, and a small target turns into a few pixels.
- Tiles. The scene cut into overlapping squares (256, 512 or 1,024 pixels) and every square asked at the same moment: "is the target here, and if so where?" Because the questions run in parallel, 70 tiles take not much longer than the slowest one. The answer is the most confident "it's here".
- Tiles, then a double-check. Several tiles often say "it's here, 100% sure", and only one can be right. So one more question shows close-ups of the top four claims side by side and asks which is the target. This is what the LLM lanes in the race do.
Decision models can only say yes or no
A decision model doesn't write anything. It answers a typed question (yes or no, pick one, score) with a probability, which makes it quick and cheap, but it has no way to say "bottom left". So it has to narrow down instead:
- Yes or no on each tile, then zoom in. Every 512-pixel tile is asked "does this show the target?" at the same moment. The likeliest tile is cut into nine overlapping pieces and those are asked again, then the likeliest of those, three rounds in all, until a piece is about the size of the target. The answer is the middle of the last round, weighted by each piece's probability.
The probabilities are a real advantage: where LLMs claim "100% sure" for several wrong tiles, a decision model's numbers can be compared, so there is no tie to break. The cost is the zooming: three more rounds of questions one after another, each on a smaller and smaller crop, and the last crops can show only part of the target.
Comparing like with like
Each kind first got the strategy that suits what it can answer. To compare them on exactly the same search, we also ran the LLMs through the decision models' zoom: the same crops and the same yes-or-no question, with the LLM giving its probability of yes.
On that common ground the decision models are the quick, cheap searchers, but not the most accurate. GPT-6 Luna Decisions and pplx-decider v1.1 each found the target in 31 of 79 scenes (39%), in a median 3.4 and 3.9 seconds, at $0.0024 and $0.0004 a scene. Claude Haiku 5.5, zooming, found 38 (48%) but took 16.2 seconds a scene at $0.0061; Gemini 3.5 Flash Lite found 26 (33%) in 10.3 seconds at $0.0236, and GPT-6 Luna 17 (22%) in 13.5 seconds. The same LLMs did far better when they could point: Flash Lite found 66 (84%) with tiles and a double-check. So asking yes or no suits a decision model, and an LLM should be asked where.
Results: found, time and cost
How to read this table: "Found" is the share of all 79 scenes where the model pointed at the target; the next three columns split it by set (29 Waldo pages, 30 V* photos, 20 of our own puzzles). Time is the median a scene, from first question to final answer. Cost per find is the run's total cost divided by the targets it found.
| # | Model | Type | How it searched | Found | Waldo | V* | Ours | Median time | Cost per scene | Cost per find |
|---|---|---|---|---|---|---|---|---|---|---|
| 1 | gemini-3.5-flash-lite | LLM | Tiles of 256 px, then a double-check | 85% 67 of 79 | 62% 18 of 29 | 97% 29 of 30 | 100% 20 of 20 | 7.1 s | $0.0569 | $0.0671 |
| 2 | gemini-3.5-flash-lite | LLM | Tiles of 512 px, then a double-check | 84% 66 of 79 | 66% 19 of 29 | 90% 27 of 30 | 100% 20 of 20 | 4.1 s | $0.0155 | $0.0186 |
| 3 | claude-haiku-5.5 | LLM | Tiles of 256 px, then a double-check | 71% 56 of 79 | 38% 11 of 29 | 87% 26 of 30 | 95% 19 of 20 | 11.3 s | $0.0152 | $0.0215 |
| 4 | claude-haiku-5.5 | LLM | Tiles of 512 px, then a double-check | 70% 55 of 79 | 41% 12 of 29 | 77% 23 of 30 | 100% 20 of 20 | 10.9 s | $0.0051 | $0.0073 |
| 5 | gemini-3.5-flash-lite | LLM | Tiles of 1,024 px, then a double-check | 68% 54 of 79 | 41% 12 of 29 | 73% 22 of 30 | 100% 20 of 20 | 4.8 s | $0.0048 | $0.0070 |
| 6 | claude-haiku-5.5 | LLM | Tiles of 1,024 px, then a double-check | 68% 54 of 79 | 41% 12 of 29 | 73% 22 of 30 | 100% 20 of 20 | 11.0 s | $0.0030 | $0.0044 |
| 7 | gpt-6-lunano reasoning | LLM | Tiles of 1,024 px, then a double-check | 53% 42 of 79 | 28% 8 of 29 | 47% 14 of 30 | 100% 20 of 20 | 5.6 s | $0.0019 | $0.0036 |
| 8 | claude-haiku-5.5 | LLM | Yes/no on each tile, then zooms in | 48% 38 of 79 | 48% 14 of 29 | 37% 11 of 30 | 65% 13 of 20 | 16.2 s | $0.0061 | $0.0127 |
| 9 | gpt-6-lunano reasoning | LLM | Tiles of 512 px, then a double-check | 42% 33 of 79 | 21% 6 of 29 | 43% 13 of 30 | 70% 14 of 20 | 6.0 s | $0.0021 | $0.0049 |
| 10 | gpt-6-luna-decisions | Decision model | Yes/no on each tile, then zooms in | 39% 31 of 79 | 28% 8 of 29 | 40% 12 of 30 | 55% 11 of 20 | 3.4 s | $0.0024 | $0.0060 |
| 11 | pplx-decider-v1.1-27b | Decision model | Yes/no on each tile, then zooms in | 39% 31 of 79 | 31% 9 of 29 | 47% 14 of 30 | 40% 8 of 20 | 3.9 s | $0.0004 | $0.0010 |
| 12 | gpt-6-lunano reasoning | LLM | Tiles of 256 px, then a double-check | 34% 27 of 79 | 21% 6 of 29 | 60% 18 of 30 | 15% 3 of 20 | 7.7 s | $0.0046 | $0.0135 |
| 13 | gemini-3.5-flash-lite | LLM | Yes/no on each tile, then zooms in | 33% 26 of 79 | 28% 8 of 29 | 40% 12 of 30 | 30% 6 of 20 | 10.3 s | $0.0236 | $0.0718 |
| 14 | gpt-6-lunano reasoning | LLM | Yes/no on each tile, then zooms in | 22% 17 of 79 | 21% 6 of 29 | 27% 8 of 30 | 15% 3 of 20 | 13.5 s | $0.0020 | $0.0092 |
| 15 | clef | Decision model | Yes/no on each tile, then zooms in | 1% 1 of 79 | 3% 1 of 29 | 0% 0 of 30 | 0% 0 of 20 | 9.5 s | $0.0303 | $2.3945 |
| 16 | clef-flash | Decision model | Yes/no on each tile, then zooms in | 0% 0 of 79 | 0% 0 of 29 | 0% 0 of 30 | 0% 0 of 20 | 5.5 s | $0.0153 | – |
The LLM rows come from a tile-size test in which the three models ran at the same time, which slows each of them a little; the decision models ran one at a time. The second round, now running, reruns every LLM alone, adds 12 more current vision models, and runs the LLMs through the decision models' zoom as well, so the two kinds can be compared on exactly the same search.
LLMs against decision models
The decision models are quick and, in pplx-decider's case, remarkably cheap. Their problem isn't finding the right area: pplx-decider picked the tile containing the target in 60 of 79 scenes, and GPT-6 Luna Decisions in 50. They lose it in the last zoom rounds, where a 48-pixel piece shows only part of a letter or a face, and the final point misses by a median of 9 pixels (pplx-decider) to 18 (GPT-6 Luna Decisions). An LLM can simply say where in the tile the target is.
pplx-decider also has a long tail: on three of our scenes a question took close to a minute, so its slowest scenes ran past 50 seconds. GPT-6 Luna Decisions never did.
Cloudflare's Clef and Clef Flash found one target between them, out of 79 scenes each. Shown a crop with the target and one without, they gave the same probability either way, so the zoom never got started in the right place. They were also the most expensive decision models per scene.
We also tried TypeSafe's Jev, a decision model that reads text only, as the judge in the verify step, choosing between Flash Lite's own few-word descriptions of each claim. It made things worse (55 found against 58 with no verify step): the descriptions were often identical.
What cutting the picture up buys
For Gemini 3.5 Flash Lite, asking about 512-pixel tiles instead of the whole scene took finds from 26 to 66 out of 79, for a little more waiting (the tiles go at once) and about 35 times the cost: $0.0155 a scene instead of $0.0004. (The whole-scene run was a pilot earlier the same day, with a prompt that differed only in not asking for a short description.) Smaller is not always better: GPT-6 Luna found most with the largest tiles (42 with 1,024 pixels, 27 with 256), and Claude Haiku 5.5 found about the same at every size. The verify step helped Flash Lite (66 against 60 without it) and hurt Haiku (55 against 62).
Waldo himself
The 29 Where's Waldo pages were the hardest set by far. Waldo is a few dozen pixels on pages up to 4,000 pixels wide, often next to people in red and white stripes, and the scans are uneven. The best run, Gemini 3.5 Flash Lite on 512-pixel tiles, found him 19 times (18 on 256-pixel tiles); Claude Haiku 5.5 found him 12 times, and the best decision model, pplx-decider, 9. Pick a Waldo scene in the race above to watch each model hunt for him. The datasets mark only his head, so a point on his head or upper body counts as finding him (we first scored the head alone, and corrected that after seeing models point at his jumper).
How we tested
- Scenes. 29 Where's Waldo pages (1,280 to 4,000 pixels) with Waldo's head marked, from the public HereIsWally, Hey-Waldo and Kaggle waldo-yolo research sets; every box checked by eye. 30 V* Bench photos (2,000 to 3,300 pixels, targets 8 to 50 pixels). 20 scenes we generated at 4,000 by 3,000 pixels.
- Models and settings. Gemini models through Google's API, the rest through OpenRouter, each lab's current models, at their fastest reasoning setting where there is one (named in the table). Decision models through OpenRouter's Decisions API.
- Rules. One point per scene; a hit if it falls in the target's box grown by 10% of its larger side. For Waldo the box is his head and upper body: the datasets' head box, widened by half and extended down by one head height, checked by eye against every near miss. A question unanswered after 60 seconds counts as a miss. Tiles overlap by 15%.
- Time runs from the first question to the final answer and includes sending the tiles, a few megabytes a scene at 512 pixels, from one office connection; faster or slower connections will shift it.
- What this doesn't show. One run per configuration, on 79 scenes: differences of a few finds are within chance. It measures finding one known target, not describing a scene.
Quick answers
Can AI find Waldo?
Sometimes. The best configuration we tested, Gemini 3.5 Flash Lite reading the page in 512-pixel tiles, found Waldo on 19 of 29 Where's Waldo pages (measured 8 October 2026). Shown the whole page at once, it found him on 9.
Which AI model is fastest at finding something in a large image?
The decision models answered fastest: GPT-6 Luna Decisions took a median 3.4 seconds a scene and pplx-decider v1.1 3.9 seconds, against 4.1 seconds for Gemini 3.5 Flash Lite. But they found fewer targets: 31 of 79 each, against 66.
What is a decision model?
A model that answers a typed question about some input (yes or no, pick one, score) with a probability instead of writing text. Four take images through OpenRouter's Decisions API: GPT-6 Luna Decisions, pplx-decider v1.1, Clef and Clef Flash. They can't point, so we had them zoom in by asking yes or no about smaller and smaller pieces.
Why cut the picture into tiles?
Models shrink a large image before they look at it, so a small target turns into a few pixels. Asking about every 512-pixel tile at once lifted Gemini 3.5 Flash Lite from 26 to 66 finds out of 79 for only a little more waiting, because the tiles are asked in parallel.
How much does it cost to search a picture with AI?
From about a tenth of a cent per target found (pplx-decider v1.1, $0.0004 a scene) to about 7 cents (Gemini 3.5 Flash Lite on small tiles). Flash Lite on 512-pixel tiles, within one find of the most accurate run, cost $0.0155 a scene, about 2 cents per find.
Model pages: Gemini 3.5 Flash-Lite · Claude Haiku 5.5 · GPT-6 Luna