cd /news/large-language-models/step-5-preview-pricing-api-costs-and… · home › topics › large-language-models › article
[ARTICLE · art-147471] src=mindstudio.ai ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Step 5 Preview Pricing: API Costs and October 15 Weights Release

Stepfun has set Step 5 Preview API pricing at $1 per million uncached input tokens, $5 per million cached input tokens, and $2.70 per million output tokens, with reasoning tokens billed at the output rate, and committed to an open-weight release on October 15. The mixture-of-experts model carries 600 billion total parameters with 27 billion active per token, supports a one-million-token context and up to 64,000 output tokens, and accepts text, images, and video. In an independent Kingbench 3 test across eight coding and reasoning tasks, Step 5 Preview scored 83.75%, landing between GLM 5.3 Flash and GLM 5.3.

by read7 min views1 publishedOct 8, 2026
Step 5 Preview Pricing: API Costs and October 15 Weights Release
Image: Mindstudio (auto-discovered)

Stepfun's Step 5 Preview API pricing per million tokens, cached input discounts, and the planned October 15 open-weight release explained.

What does Step 5 Preview cost to use? #

Stepfun prices Step 5 Preview at $1 per million uncached input tokens, $5 per million cached input tokens, and $2.70 per million output tokens, with reasoning tokens billed as part of output. That’s the published international API rate for the preview version currently accessible through Stepfun’s API. The model supports a million tokens of context and up to 64,000 output tokens per generation, and it accepts text, images, and video as input, though most early testing has focused on text prompts run through coding agents.

TL;DR #

  • Step 5 Preview costs $1 per million tokens for standard input,$5 per million for cached input, and**$2.70 per million** for output, with reasoning tokens folded into the output rate.

  • The model is a mixture of experts architecture with 600 billion total parameters and 27 billion active per token, meaning only a fraction of the network fires for any given token.

  • Stepfun has committed to an open-weight release on October 15 , after which developers can presumably run the model outside the hosted API.

  • Context length runs to one million tokens with output capped at 64,000 tokens, a combination that matters for long coding sessions or large document workloads.

  • In an independent Kingbench 3 test across eight coding and reasoning tasks, Step 5 Preview scored 83.75% , landing between GLM 5.3 Flash and GLM 5.3 on the comparison chart.

  • The cached input discount only pays off if your workflow repeatedly sends shared context, such as a long system prompt or a stable codebase snippet, across many calls.

  • Standout results in testing included a playable archery game with physics and alocal LoRA fine-tuning workflow that produced a working web app querying a locally trained model.

  • ✕a coding agent

  • ✕no-code

  • ✕vibe coding

  • ✕a faster Cursor

The one that tells the coding agents what to build.

How does the pricing structure actually work? #

Step 5 Preview uses a three-tier token pricing model that splits input into cached and uncached categories, then bills output separately. Uncached input, the default rate for any new or unique prompt content, runs $1 per million tokens. Cached input, which applies when the API recognizes previously processed context being reused, costs $5 per million tokens. Output tokens, including any tokens the model spends on internal reasoning before producing a final answer, are billed at $2.70 per million.

That cached rate looks counterintuitive at first since it’s higher than the uncached rate, not lower. The practical implication is that caching discounts depend on your specific setup and provider terms rather than being a universal cost saver. For developers running coding agents that repeatedly send the same system instructions or shared file context across many turns, the actual savings or added cost will depend on call volume and how much of each prompt is genuinely reusable versus freshly generated. The bottom line: total spend is still driven mostly by how many calls an agent makes and how much it generates per call, not by the sticker price alone.

What’s inside Step 5 Preview’s architecture? #

Stepfun describes Step 5 as a mixture of experts (MoE) model with 600 billion parameters in total, of which 27 billion are active for any given token. This is the same general design philosophy behind several recent large language models: rather than running every parameter for every token, the model routes each token through a subset of specialized expert networks. The result is a model that behaves, in terms of inference cost and speed, closer to a much smaller dense model, while retaining the capacity benefits of a much larger overall parameter count.

This matters for pricing because MoE architectures are part of why providers can offer large models at relatively accessible per-token rates. The 27 billion active parameters is the number that most directly correlates with inference compute cost per token, even though the full 600 billion parameters need to be stored and made available for routing.

When are the open weights coming out? #

Stepfun has stated that Step 5’s weights are planned for release on October 15. Until that date, the only way to use Step 5 is through Stepfun’s hosted API, which is what’s reflected in the preview pricing above. Once the weights are available, developers will presumably be able to self-host the model, which opens up different cost tradeoffs entirely: no per-token API fees, but the hardware burden of running a 600-billion-parameter MoE model locally or on rented infrastructure.

Given the parameter count, self-hosting Step 5 will not be a casual undertaking. A model of this scale typically requires multi-GPU setups or quantized versions to run outside of a well-resourced data center, though specific VRAM or hardware requirements for the open release have not yet been published. Anyone planning around the October 15 date should treat it as a weights release, not necessarily a plug-and-play local deployment on consumer hardware.

Other agents start typing. Remy starts asking. #

Scoping, trade-offs, edge cases — the real work. Before a line of code.

Is Step 5 Preview worth the cost for coding work? #

Based on an independent benchmark run using Kingbench 3, a suite of eight coding and reasoning tasks executed inside the Open Code agent framework, Step 5 Preview scored 67 out of 80 points, or 83.75%. Tasks included building a multi-elevator simulation, a 3D interactive object (a contact lens case), a folding table animation, an SVG illustration, a playable archery game, a counting/reasoning problem, a full local LoRA fine-tuning pipeline with a working web interface, and a 3D wristwatch.

Seven of the eight tasks scored 8 out of 10 or higher, including a 10/10 on the counting problem and two 9.5/10 scores: one for the archery game (which included wind effects, moving targets, and a working leaderboard) and one for the fine-tuning task (which produced a trained Gemma 2B model via LoRA, fused weights, and a functioning local web app, despite some factual errors in the generated training data). The weakest result was the 3D wristwatch, which the reviewer called unremarkable at 6/10.

For context, GLM 5.3 scored 91.25% and GLM 5.3 Flash scored 78.75% on the same chart using their own historical results, while MiMo V2.6 Pro and MiMo V2.6 Flash scored 69.38% and 72.50% respectively. Step 5 Preview’s 83.75% places it above both Flash-tier models and the Pro-tier MiMo model, though behind full GLM 5.3. Whether that performance justifies the API cost depends on the kind of work: for interactive prototyping and local experimentation tasks, the combination of a working playable game and a complete fine-tuning pipeline suggests the model handles multi-step, verifiable coding tasks well.

Who should consider using Step 5 Preview right now? #

Developers building interactive prototypes, such as browser-based games or 3D visualizations, and those experimenting with local model fine-tuning workflows, are the clearest fit based on the demonstrated results. The model’s ability to chain together data generation, a training job, weight saving, and a working inference interface in a single agent session is a nontrivial capability that not every coding model handles cleanly.

Teams with high-volume production workloads should weigh the $2.70 per million output token rate against alternatives, since output costs (including reasoning tokens) will dominate the bill for any agent that does extensive multi-step work before arriving at a final answer. The preview nature of the API also means pricing or capabilities could shift once the full release and open weights land on October 15.

Frequently Asked Questions #

How much does Step 5 Preview cost per million tokens?

Step 5 Preview costs $1 per million uncached input tokens, $5 per million cached input tokens, and $2.70 per million output tokens. Reasoning tokens are billed under the output rate.

When will Step 5’s weights be released?

Stepfun has stated the weights are planned for release on October 15. Until then, the model is only accessible through Stepfun’s hosted preview API.

What is Step 5’s context window and output limit?

Step 5 Preview supports up to one million tokens of context and allows up to 64,000 output tokens per generation.

Other agents ship a demo. Remy ships an app. #

Real backend. Real database. Real auth. Real plumbing. Remy has it all.

What architecture does Step 5 use?

Step 5 is a mixture of experts model with 600 billion total parameters and 27 billion active parameters per token, meaning only a portion of the model’s parameters are used for any given token.

How does Step 5 Preview compare to other models on coding benchmarks?

In an independent Kingbench 3 test across eight coding and reasoning tasks, Step 5 Preview scored 83.75%, ahead of GLM 5.3 Flash (78.75%) and both MiMo V2.6 variants, but behind GLM 5.3 (91.25%).

── more in #large-language-models 4 stories · sorted by recency
── more on @stepfun 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/step-5-preview-prici…] indexed:0 read:7min 2026-10-08 · —