Models / Kev 4B
Kev 4B
Early evidence, 5 results from 1 source, on 1 use case. Kev 4B trails most of the field on what it has been measured on; it is not on sale through OpenRouter, so it is here for reference.
Compared with the models available today (it is not on sale itself), in the bottom quarter for speed under pressure.
Results as of 8 October 2026, from 5 results on 1 source; prices checked 8 Oct 2026.
- Price per million tokens
- Not available through OpenRouter
- Speed
- No speed yet
Not timed: Artificial Analysis, our source for speed, has not tested it - Intelligence Index
- Not measured
Artificial Analysis has not scored it - Spring Prompt overall
- Not ranked yet
needs our benchmarks and two groups of results - Developer
- Jared Palmer
Weights not published - Released
- Not recorded
no launch date in our tracker or our sources
How good is it, and for what?
Each use case ranks the available models its benchmarks measured. The bar shows its percentile in that field, best to the right. Open a row for the results behind it.
Product listingsTurning a sparse product feed and photos into listings that can go live
- CatalogBench: has not measured this model.
Decks from an analysisTurning a finished analysis into a deck you could present as it is
- DeckBench: has not measured this model.
User surveysPlanning a user survey and reading its results without being misled
- SurveyBench: has not measured this model.
Marketing planningPlanning a year of ad spend without overspending
- ROASBench: has not measured this model.
Agents and tool useMulti-step tasks with tools: support desks, coding agents, function calls
- tau2-bench: has not measured this model.
- Berkeley Function Calling Leaderboard (BFCL) V4: has not measured this model.
- OpenHands Index: has not measured this model.
- Microsoft STATE-Bench: has not measured this model.
- Artificial Analysis: has not measured this model.
- Vending-Bench 2: has not measured this model.
Professional workReal tasks from banking, consulting and law, business documents and freelance projects
- APEX-Agents: has not measured this model.
- GDP.pdf: has not measured this model.
- Remote Labor Index: has not measured this model.
Reasoning and knowledgeHard questions across science, maths and general knowledge
- Artificial Analysis: has not measured this model.
WritingWhat people prefer in blind comparisons, and judged writing quality
- Arena (formerly LMArena): has not measured this model.
- UGI Leaderboard: has not measured this model.
Sticking to the factsSummarising without inventing things, and factual answers
- Vectara Hallucination Leaderboard: has not measured this model.
- Arena (formerly LMArena): has not measured this model.
- SimpleQA Verified (Epoch AI): has not measured this model.
Speed under pressureGood decisions against a real clock (fast chess)
| Benchmark | Kev 4B | Rank among available | Best available |
|---|---|---|---|
| BulletBenchBulletBench Lightning 10+1: ladder Elo | 33 | 12th of 15 | Gemini 3.5 Flash-Lite 892 |
Nearest alternatives
- The most capable open-weights modelMiMo-V2.6-Pro$0.54 blended · Intelligence Index 46.3
Against Kev 4B. "Similar intelligence" means within 4 points or better.
Against the alternatives
The models you would most likely weigh it against: the leaders of the groups above.
Scroll sideways to see every alternative.
| Kev 4B | Claude Opus 5.5 leads overall | GPT-6.1 Sol near the top overall | GPT-6 Astra near the top overall | GPT-6 Sol near the top overall | |
|---|---|---|---|---|---|
| Price per million tokens | – | $8.00 | $4.00 | $20.00 | $4.00 |
| Tokens a second | – | 97 | 55 | 52 | – |
| Intelligence Index | – | 57.6 | 51.8 | 52.7 | 47.6 |
| Spring Prompt overall | – | 85 | 81 | 79 | 69 |
| Rank among available models, by use case | |||||
| Product listings | – | 6th | 2nd | 1st | 3rd |
| Decks from an analysis | – | 5th | 3rd | 1st | 2nd |
| User surveys | – | 2nd | 1st | – | – |
| Marketing planning | – | 5th | – | 1st | 2nd |
| Agents and tool use | – | 2nd | 3rd | 4th | 5th |
| Professional work | – | 3rd | 9th | 5th | 13th |
| Reasoning and knowledge | – | 1st | 5th | 3rd | 11th |
| Writing | – | 4th | 15th | 19th | 51st |
| Sticking to the facts | – | 2nd | 18th | 24th | 49th |
| Speed under pressure | 19th of 23 | – | – | 14th | 10th |
Shaded figures are better than Kev 4B's. A dash means no result.
Where it has been measured
- BulletBench5 results
- CatalogBenchNot measured
- DeckBenchNot measured
- ROASBenchNot measured
- SurveyBenchNot measured
- APEX-AgentsNot measured
- Arena (formerly LMArena)Not measured
- Artificial AnalysisNot measured
- Berkeley Function Calling Leaderboard (BFCL) V4Not measured
- GDP.pdfNot measured
- Microsoft STATE-BenchNot measured
- OpenHands IndexNot measured
- Remote Labor IndexNot measured
- SimpleQA Verified (Epoch AI)Not measured
- tau2-benchNot measured
- UGI LeaderboardNot measured
- Vectara Hallucination LeaderboardNot measured
- Vending-Bench 2Not measured
Every result
5 published results from 1 source, each in the source's own units, with the configuration that produced it.
BulletBench · measured by Spring Prompt · 5 results
| Measure | Value | Rank | Configuration | Dated |
|---|---|---|---|---|
| BulletBench Lightning 10+1: cost per gameUS dollars, lower is better | $0.0006 | 3 of 23 | Kev 4B provider default reasoning |
7 Oct 2026 |
| BulletBench Lightning 10+1: games lost on time% of games, lower is better | 66.7% | 18 of 23 | Kev 4B provider default reasoning |
7 Oct 2026 |
| BulletBench Lightning 10+1: invalid moves% of moves, lower is better | 0.0% | 1 of 23 | Kev 4B provider default reasoning |
7 Oct 2026 |
| BulletBench Lightning 10+1: ladder Eloladder Elo, higher is better | 33 0–270 |
20 of 23 | Kev 4B provider default reasoning |
7 Oct 2026 |
| BulletBench Lightning 10+1: median move timemilliseconds, lower is better | 1.0 s | 17 of 23 | Kev 4B provider default reasoning |
7 Oct 2026 |
Not shown: Not your cost: prices are those charged on the run date.