Models / Step 5 Preview
Step 5 Preview
Early evidence, 11 results from 1 source, on 2 use cases. Step 5 Preview is strongest at writing; it is not on sale through OpenRouter, so it is here for reference.
Compared with the models available today (it is not on sale itself), mid-field on writing and sticking to the facts, the use cases measured so far.
Results as of 8 October 2026, from 11 results on 1 source; prices checked 8 Oct 2026.
- Price per million tokens
- Not available through OpenRouter
- Speed
- No speed yet
Not timed: Artificial Analysis, our source for speed, has not tested it - Intelligence Index
- Not measured
Artificial Analysis has not scored it - Spring Prompt overall
- Not ranked yet
needs our benchmarks and two groups of results - Developer
- StepFun, based in China
Weights not published - Released
- Not recorded
no launch date in our tracker or our sources
How good is it, and for what?
Each use case ranks the available models its benchmarks measured. The bar shows its percentile in that field, best to the right. Open a row for the results behind it.
Product listingsTurning a sparse product feed and photos into listings that can go live
- CatalogBench: has not measured this model.
Decks from an analysisTurning a finished analysis into a deck you could present as it is
- DeckBench: has not measured this model.
User surveysPlanning a user survey and reading its results without being misled
- SurveyBench: has not measured this model.
Marketing planningPlanning a year of ad spend without overspending
- ROASBench: has not measured this model.
Agents and tool useMulti-step tasks with tools: support desks, coding agents, function calls
- tau2-bench: has not measured this model.
- Berkeley Function Calling Leaderboard (BFCL) V4: has not measured this model.
- OpenHands Index: has not measured this model.
- Microsoft STATE-Bench: has not measured this model.
- Artificial Analysis: has not measured this model.
- Vending-Bench 2: has not measured this model.
Professional workReal tasks from banking, consulting and law, business documents and freelance projects
- APEX-Agents: has not measured this model.
- GDP.pdf: has not measured this model.
- Remote Labor Index: has not measured this model.
Reasoning and knowledgeHard questions across science, maths and general knowledge
- Artificial Analysis: has not measured this model.
WritingWhat people prefer in blind comparisons, and judged writing quality
| Benchmark | Step 5 Preview | Rank among available | Best available |
|---|---|---|---|
| Arena (formerly LMArena)Arena Text: overall · reported | 1,454 | 49th of 167 | Claude Opus 4.6 1,505 |
- UGI Leaderboard: has not measured this model.
Sticking to the factsSummarising without inventing things, and factual answers
| Benchmark | Step 5 Preview | Rank among available | Best available |
|---|---|---|---|
| Arena (formerly LMArena)Arena Text factuality: overall · reported | 1,452 | 45th of 114 | Claude Opus 5.5 1,507 |
- Vectara Hallucination Leaderboard: has not measured this model.
- SimpleQA Verified (Epoch AI): has not measured this model.
Speed under pressureGood decisions against a real clock (fast chess)
- BulletBench: has not measured this model.
Nearest alternatives
- The most capable open-weights modelMiMo-V2.6-Pro$0.54 blended · Intelligence Index 46.3
Against Step 5 Preview. "Similar intelligence" means within 4 points or better.
Against the alternatives
The models you would most likely weigh it against: the leaders of the groups above.
Scroll sideways to see every alternative.
| Step 5 Preview | Claude Opus 5.5 leads overall | GPT-6.1 Sol near the top overall | GPT-6 Astra near the top overall | GPT-6 Sol near the top overall | |
|---|---|---|---|---|---|
| Price per million tokens | – | $8.00 | $4.00 | $20.00 | $4.00 |
| Tokens a second | – | 97 | 55 | 52 | – |
| Intelligence Index | – | 57.6 | 51.8 | 52.7 | 47.6 |
| Spring Prompt overall | – | 85 | 81 | 79 | 69 |
| Rank among available models, by use case | |||||
| Product listings | – | 6th | 2nd | 1st | 3rd |
| Decks from an analysis | – | 5th | 3rd | 1st | 2nd |
| User surveys | – | 2nd | 1st | – | – |
| Marketing planning | – | 5th | – | 1st | 2nd |
| Agents and tool use | – | 2nd | 3rd | 4th | 5th |
| Professional work | – | 3rd | 9th | 5th | 13th |
| Reasoning and knowledge | – | 1st | 5th | 3rd | 11th |
| Writing | 53rd of 189 | 4th | 15th | 19th | 50th |
| Sticking to the facts | 55th of 153 | 2nd | 18th | 24th | 49th |
| Speed under pressure | – | – | – | 14th | 10th |
Shaded figures are better than Step 5 Preview's. A dash means no result.
Where it has been measured
- Arena (formerly LMArena)11 results
- BulletBenchNot measured
- CatalogBenchNot measured
- DeckBenchNot measured
- ROASBenchNot measured
- SurveyBenchNot measured
- APEX-AgentsNot measured
- Artificial AnalysisNot measured
- Berkeley Function Calling Leaderboard (BFCL) V4Not measured
- GDP.pdfNot measured
- Microsoft STATE-BenchNot measured
- OpenHands IndexNot measured
- Remote Labor IndexNot measured
- SimpleQA Verified (Epoch AI)Not measured
- tau2-benchNot measured
- UGI LeaderboardNot measured
- Vectara Hallucination LeaderboardNot measured
- Vending-Bench 2Not measured
Every result
11 published results from 1 source, each in the source's own units, with the configuration that produced it.
Arena (formerly LMArena) · reported by Arena (formerly LMArena) · 11 results
| Measure | Value | Rank | Configuration | Dated |
|---|---|---|---|---|
| Arena Agent: confirmed task successIPS effect estimate, higher is better | 0.05 0.01–0.08 |
22 of 50 | Step 5 Preview (arena agent) | 2 Oct 2026 |
| Arena Agent: praise over complaintIPS effect estimate, higher is better | -0.05 -0.10–0.00 |
38 of 50 | Step 5 Preview (arena agent) | 2 Oct 2026 |
| Arena Agent: steerabilityIPS effect estimate, higher is better | 0.02 -0.00–0.04 |
21 of 50 | Step 5 Preview (arena agent) | 2 Oct 2026 |
| Arena Agent: tool groundingIPS effect estimate, higher is better | 0.00 0.00–0.01 |
8 of 50 | Step 5 Preview (arena agent) | 2 Oct 2026 |
| Arena Text factuality: overallArena rating, higher is better | 1,452 1,442–1,461 |
62 of 182 | Step 5 Preview | 2 Oct 2026 |
| Arena Text: business, management and financeArena rating, higher is better | 1,468 1,442–1,494 |
42 of 406 | Step 5 Preview | 2 Oct 2026 |
| Arena Text: creative writingArena rating, higher is better | 1,421 1,396–1,445 |
84 of 411 | Step 5 Preview | 2 Oct 2026 |
| Arena Text: expert promptsArena rating, higher is better | 1,472 1,438–1,505 |
86 of 363 | Step 5 Preview | 2 Oct 2026 |
| Arena Text: instruction followingArena rating, higher is better | 1,445 1,427–1,464 |
73 of 413 | Step 5 Preview | 2 Oct 2026 |
| Arena Text: overallArena rating, higher is better | 1,454 1,443–1,465 |
74 of 413 | Step 5 Preview | 2 Oct 2026 |
| Arena Text: writing, literature and languageArena rating, higher is better | 1,436 1,415–1,457 |
74 of 412 | Step 5 Preview | 2 Oct 2026 |
Not shown: Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.