Agent backends
Power tool-using loops that need reasoning, code execution plans, and a long working context at a low token price.
Independent tool overview
Step-3.5-Flash is StepFun's Apache-2.0 open-weight reasoning model for coding and agent workflows. Its sparse mixture-of-experts architecture has about 196.8B total parameters but activates roughly 11B per generated token, pairing a 256K context window with relatively low hosted-API pricing.
Visit the official Step-3.5-Flash site ↗
Overview
Step-3.5-Flash is a text-generation foundation model optimized for tool use, coding, long-context work, and agent loops. It can be called through StepFun's OpenAI-compatible international API, used through supported third-party hosts, or self-hosted from published weights.
The model's efficiency comes from sparse activation, sliding-window attention, and multi-token prediction. Only a subset of experts runs for each token, but the deployment still has to store and serve a much larger 196B-class model. The 11B active-parameter figure should not be read as an 11B-model memory requirement.
StepFun has since published the newer Step-3.7-Flash, which adds multimodal capabilities and broader updates. Step-3.5-Flash remains available and can still make sense when an existing text, coding, or agent workload is already validated on it or when its current price and deployment support are preferable.
Use cases
The strongest fit depends on the job you need the product to complete, not the size of its feature list.
Power tool-using loops that need reasoning, code execution plans, and a long working context at a low token price.
Support repository exploration, implementation, debugging, and terminal-oriented engineering workflows.
Process large text or code inputs within a published 256K-token context window, subject to quality testing at the intended length.
Run or adapt published weights when a team has the substantial hardware and operations capacity required.
Use the official hosted endpoint when its quality is sufficient and per-token economics matter more than access to the newest flagship model.
Capabilities
Routes each token through eight of 288 routed experts plus a shared expert, activating about 11B parameters from the full model.
Uses a hybrid of sliding-window and full attention to support large codebases, documents, and accumulated agent context.
Includes an MTP head intended to improve decoding throughput by predicting multiple tokens in a forward pass.
Training and published evaluations emphasize coding, terminal work, browsing, research, and multi-step task execution.
The international StepFun endpoint works with common Chat Completions clients using the model ID step-3.5-flash.
Weights, technical material, deployment recipes, quantized variants, and integration examples are publicly available under Apache 2.0.
Official guides cover agent and coding tools including OpenClaw, Claude Code, Codex CLI, Roo Code, Kilo Code, and others.
Process
Step 1
Start with the official API for fast evaluation; consider local or dedicated deployment only when privacy, customization, or scale justifies the infrastructure.
Step 2
Include actual codebases, tools, languages, context lengths, failure cases, and output constraints from the intended workflow.
Step 3
Use StepFun's OpenAI-compatible endpoint or a documented integration, store the API key server-side, and pin the intended model ID.
Step 4
Validate arguments, sandbox code, restrict credentials and filesystem access, and require approval for consequential actions.
Step 5
Track not only token price but reasoning length, tool calls, latency, retries, completion rate, and human correction time.
Step 6
Re-run the same evaluations against Step-3.7-Flash and other current candidates before standardizing on an older model generation.
Cost
The model weights are available under Apache 2.0, but self-hosting still incurs hardware and operations costs. StepFun's international pay-as-you-go API lists Step-3.5-Flash at $0.10 per 1M uncached input tokens, $0.02 per 1M cached input tokens, and $0.30 per 1M output tokens. Step Plan subscriptions offer quota-based access for high-frequency agent and coding use.
No model license fee
Download and deploy under Apache 2.0 while paying the resulting infrastructure and operations costs.
$0.10 input / $0.30 output per 1M tokens
Official hosted inference through StepFun's international platform.
$6.99
Entry subscription for supported agent and coding integrations.
$9.99
Higher quota for daily productivity.
$29
Higher-volume plan for heavy agent use.
$99
Largest published Step Plan quota for professional workloads.
Pricing checked . Check current pricing at the source ↗
Assessment
Compare
The right alternative depends on the specific output, workflow, controls and budget your project requires.
Consumer
A newer high-speed agentic model positioned specifically for OpenClaw and similar long-running tool workflows.
Explore GLM-5-Turbo →Coding
A more compact open-weight option designed around agentic coding when local deployment size matters.
Explore Qwen3-Coder-Next →Business Operations
A broader open-model ecosystem with hosted chat and API options for reasoning, coding, and general-purpose workloads.
Explore DeepSeek →Questions
StepFun publishes the weights and repository under the Apache 2.0 license. Teams should still review the license and any third-party dependencies or datasets relevant to their deployment.
The model contains about 196.8B parameters, but its router selects a subset of experts so roughly 11B participate in generating each token. That reduces computation; it does not shrink the full weight set to an 11B model.
Yes, official full and quantized releases support local deployment, and StepFun gives examples using high-end consumer or compact AI hardware. Memory, storage, speed, and context capacity still depend on the quantization and machine.
The international platform currently lists $0.10 per 1M uncached input tokens, $0.02 per 1M cached input tokens, and $0.30 per 1M output tokens for Step-3.5-Flash. Reasoning length can materially affect the effective cost of a task.
Yes, the official specification lists 256K. Long-context support is a capacity claim, not proof that accuracy stays constant across every position or task, so evaluate it at the lengths you actually need.
Coding and agent workflows are its main positioning, and StepFun publishes integrations and benchmark results for them. Production selection should still use repository-specific tests, tool-call completion rates, security checks, and total task cost.
Test both. Step-3.7-Flash is newer and multimodal, while Step-3.5-Flash remains a low-cost, text-focused option with established deployment recipes. The better choice depends on accuracy, modality, latency, compatibility, and cost in the actual workflow.
Bottom line
Step-3.5-Flash remains an attractive open-weight engine for cost-sensitive coding and agent experiments, especially through the inexpensive official API. Its sparse design is efficient but does not make self-hosting trivial, and StepFun's own known-issues section is a useful warning against trusting long, specialized agent runs without supervision. Evaluate it beside Step-3.7-Flash and current peers before committing a new production stack.
Visit Step-3.5-Flash website ↗.png)
Qwen3-Max-Thinking - Alibaba's new flagship reasoning model competitive with models like Claude 4.5 Opus, GPT 5.2-Thinking, and Gemini 3 Pro across benchmarks

Claude Opus 4.6 - Anthropic's upgrade to its most powerful model line, featuring multi-agent collaboration, 1M context window, and new Office integrations

ERNIE 5.0 - Baidu's omni-modal, top-ranked Chinese model

Tiny Aya - Cohere Labs' new open-source multilingual small model covering 70+ languages in just 3.35B parameters

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.
Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.
To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.
Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.
Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.