Models / Step 5 Preview

StepFun

Step 5 Preview

Early evidence, 11 results from 1 source, on 2 use cases. Step 5 Preview is strongest at writing; it is not on sale through OpenRouter, so it is here for reference.

Compared with the models available today (it is not on sale itself), mid-field on writing and sticking to the facts, the use cases measured so far.

Results as of 8 October 2026, from 11 results on 1 source; prices checked 8 Oct 2026.

Price per million tokens
Not available through OpenRouter
Speed
No speed yet
Not timed: Artificial Analysis, our source for speed, has not tested it
Intelligence Index
Not measured
Artificial Analysis has not scored it
Spring Prompt overall
Not ranked yet
needs our benchmarks and two groups of results
Developer
StepFun, based in China
Weights not published
Released
Not recorded
no launch date in our tracker or our sources

How good is it, and for what?

Each use case ranks the available models its benchmarks measured. The bar shows its percentile in that field, best to the right. Open a row for the results behind it.

Product listingsTurning a sparse product feed and photos into listings that can go live Not measured
  • CatalogBench: has not measured this model.
Decks from an analysisTurning a finished analysis into a deck you could present as it is Not measured
  • DeckBench: has not measured this model.
User surveysPlanning a user survey and reading its results without being misled Not measured
  • SurveyBench: has not measured this model.
Marketing planningPlanning a year of ad spend without overspending Not measured
  • ROASBench: has not measured this model.
Agents and tool useMulti-step tasks with tools: support desks, coding agents, function calls Not measured
  • tau2-bench: has not measured this model.
  • Berkeley Function Calling Leaderboard (BFCL) V4: has not measured this model.
  • OpenHands Index: has not measured this model.
  • Microsoft STATE-Bench: has not measured this model.
  • Artificial Analysis: has not measured this model.
  • Vending-Bench 2: has not measured this model.
Professional workReal tasks from banking, consulting and law, business documents and freelance projects Not measured
  • APEX-Agents: has not measured this model.
  • GDP.pdf: has not measured this model.
  • Remote Labor Index: has not measured this model.
Reasoning and knowledgeHard questions across science, maths and general knowledge Not measured
  • Artificial Analysis: has not measured this model.
WritingWhat people prefer in blind comparisons, and judged writing quality 53rd of 189 available
BenchmarkStep 5 PreviewRank among availableBest available
Arena (formerly LMArena)Arena Text: overall · reported 1,454 49th of 167 Claude Opus 4.6 1,505
  • UGI Leaderboard: has not measured this model.
Sticking to the factsSummarising without inventing things, and factual answers 55th of 153 available
BenchmarkStep 5 PreviewRank among availableBest available
Arena (formerly LMArena)Arena Text factuality: overall · reported 1,452 45th of 114 Claude Opus 5.5 1,507
  • Vectara Hallucination Leaderboard: has not measured this model.
  • SimpleQA Verified (Epoch AI): has not measured this model.
Speed under pressureGood decisions against a real clock (fast chess) Not measured
  • BulletBench: has not measured this model.

Nearest alternatives

  • The most capable open-weights modelMiMo-V2.6-Pro$0.54 blended · Intelligence Index 46.3

Against Step 5 Preview. "Similar intelligence" means within 4 points or better.

Against the alternatives

The models you would most likely weigh it against: the leaders of the groups above.

Scroll sideways to see every alternative.

Step 5 PreviewClaude Opus 5.5
leads overall
GPT-6.1 Sol
near the top overall
GPT-6 Astra
near the top overall
GPT-6 Sol
near the top overall
Price per million tokens – $8.00$4.00$20.00$4.00
Tokens a second – 975552–
Intelligence Index – 57.651.852.747.6
Spring Prompt overall – 85817969
Rank among available models, by use case
Product listings – 6th2nd1st3rd
Decks from an analysis – 5th3rd1st2nd
User surveys – 2nd1st––
Marketing planning – 5th–1st2nd
Agents and tool use – 2nd3rd4th5th
Professional work – 3rd9th5th13th
Reasoning and knowledge – 1st5th3rd11th
Writing 53rd of 189 4th15th19th50th
Sticking to the facts 55th of 153 2nd18th24th49th
Speed under pressure – ––14th10th

Shaded figures are better than Step 5 Preview's. A dash means no result.

Where it has been measured

Every result

11 published results from 1 source, each in the source's own units, with the configuration that produced it.

Arena (formerly LMArena) · reported by Arena (formerly LMArena) · 11 results
MeasureValueRankConfigurationDated
Arena Agent: confirmed task successIPS effect estimate, higher is better 0.05
0.01–0.08
22 of 50 Step 5 Preview (arena agent) 2 Oct 2026
Arena Agent: praise over complaintIPS effect estimate, higher is better -0.05
-0.10–0.00
38 of 50 Step 5 Preview (arena agent) 2 Oct 2026
Arena Agent: steerabilityIPS effect estimate, higher is better 0.02
-0.00–0.04
21 of 50 Step 5 Preview (arena agent) 2 Oct 2026
Arena Agent: tool groundingIPS effect estimate, higher is better 0.00
0.00–0.01
8 of 50 Step 5 Preview (arena agent) 2 Oct 2026
Arena Text factuality: overallArena rating, higher is better 1,452
1,442–1,461
62 of 182 Step 5 Preview 2 Oct 2026
Arena Text: business, management and financeArena rating, higher is better 1,468
1,442–1,494
42 of 406 Step 5 Preview 2 Oct 2026
Arena Text: creative writingArena rating, higher is better 1,421
1,396–1,445
84 of 411 Step 5 Preview 2 Oct 2026
Arena Text: expert promptsArena rating, higher is better 1,472
1,438–1,505
86 of 363 Step 5 Preview 2 Oct 2026
Arena Text: instruction followingArena rating, higher is better 1,445
1,427–1,464
73 of 413 Step 5 Preview 2 Oct 2026
Arena Text: overallArena rating, higher is better 1,454
1,443–1,465
74 of 413 Step 5 Preview 2 Oct 2026
Arena Text: writing, literature and languageArena rating, higher is better 1,436
1,415–1,457
74 of 412 Step 5 Preview 2 Oct 2026

Not shown: Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.