AI Checkpoint: Sol’s new fast lane

There’s a new Mistral, a cheaper Haiku, and a faster way to run Sol. We’ve also been testing surveys and hunting for tiny things in large pictures. Here’s what we found.
5–9 October 2026 · News cutoff: Friday, 15:20 Bangkok time (UTC+7)
The Friday test
Sol loses fewer games on the clock

OpenAI switched on Ultrafast for Sol on Thursday. It’s the same model on a faster serving tier, available through the Responses API.
We tried it Friday morning. Standard Sol lost nine of twelve one-minute chess games on time. Ultrafast lost three. Its ladder rating rose from 323 to 928; median moves took 1.4 seconds instead of 2.2.
The bill was less encouraging: $0.34 per game against $0.0243. Ultrafast tokens cost six times as much, and the runs used different amounts. We’d try it where slow answers lose the job. Twelve chess games won’t tell you what it saves in a support queue.
New this week
A cheaper Haiku, a new Mistral
Haiku 5.5 arrived on Wednesday. Base token prices are a tenth of Haiku 4.5’s: $0.10 input and $0.50 output per million. There’s a catch at 100,000 input tokens: go over and the rates become $0.50 and $2.50. Our launch-day runs used default medium effort. It improved a lot on Haiku 4.5, but Luna still beat it on listings, decks and ad planning. Our review →

Sonnet 5.5’s cache reads also got cheaper. Anthropic halved them to $0.10 per million tokens. How much that saves depends on how much of your input is cached.
Mistral Large 4 is a preview. It launched Tuesday, with European training and serving. Weights are promised for month-end. In our 6 October tests, high reasoning reduced invented product claims but hurt ad planning. We wouldn’t leave its customer-facing copy unchecked. Our review →
OpenAI’s Decisions API is in beta. The Tuesday release uses GPT-6 Luna for typed answers and choices with probabilities. Worth looking at for routing and classification. In our visual-search tests reported Thursday, its workflow found 31 of 79 targets; Flash Lite’s tiled workflow found 66. Fast answers still need checking.
ChatGPT gets interactive answers. Wednesday’s Intelligent UI rollout adds charts, diagrams and small tools in ChatGPT, starting with paid plans and expanding to Free and Go on Thursday. This is the ChatGPT interface, not a standard API response.
Google announced the Gemini agent on Thursday. The pitch: work across documents, inboxes and business systems, with shared context and model routing. We haven’t tested it; check which features your account can use.
From our bench
The sums were right. Seven analyses still made things up.
We launched SurveyBench on Thursday. Models write a survey, then analyse simulated responses containing traps: tiny groups, filtered bases and skewed samples.
Haiku got every set numeric question right. Seven of its nine analyses still drew findings the data didn’t support.
Sol and Opus 5.5 led the analysis ratings, with overlapping uncertainty ranges. We can’t call a winner. The suite is small and synthetic, the judges are models, and human calibration is still to come.
One thing to try
Give the model a closer look


In our 8 October tests, Flash Lite found 26 of 79 targets with a whole-image pilot, and 66 with a tiled workflow, at provider default settings. Splitting the image cost about 35 times more. The pilot prompt also differed slightly, so cropping alone doesn’t get all the credit.
Try it: send overlapping crops in parallel. Ask each for the target’s location and a short description, then compare close-ups of the best candidates. Test that last check: it helped Flash Lite but hurt Haiku in our 512px runs.
That’s this week. If you try the image trick, keep count of the misses as well as the finds.