# Step 3.5 Flash: a 196B model that behaves like an 11B one

> Sparse MoE with 11B active parameters, three-way multi-token prediction, a 3:1 sliding window attention ratio for 256K context. The benchmark table is impressive, and the footnotes on that table matter as much as the numbers.

**stepfun-ai/Step-3.5-Flash** — Fast, Sharp & Reliable Agentic Intelligence

- Repository: https://github.com/stepfun-ai/Step-3.5-Flash
- Website: https://static.stepfun.com/blog/step-3.5-flash/
- Stars: 2,083 · Forks: 87
- Language: C++
- License: Apache-2.0
- Published: 2026-10-07 · Updated: 2026-10-07 · Language: en
- Canonical page: https://hysenlabs.com/projects/stepfun-ai-step-3-5-flash

## The architecture: sparse activation plus multi-token prediction

The headline numbers are 196B total parameters with 11B activated per token. That is a sparse Mixture of Experts, and the README's phrase for it is intelligence density: only a fraction of the weights participate in any given token, so the model can be large in knowledge while behaving like a small model in cost.

Two specific mechanisms do the work on top of that. The first is three-way Multi-Token Prediction, labelled MTP-3. Rather than emitting one token per forward pass, the model predicts three, which is where the throughput figure comes from. The README states 100 to 300 tokens per second in typical use, peaking at 350 for single-stream coding tasks, and makes the reasoning explicit: chatbots are built for reading while agents must reason fast, so multi-step reasoning chains need immediate responses rather than a comfortable reading pace.

The second mechanism is context handling. The model supports a 256K context window using a 3:1 sliding window attention ratio, integrating three sliding-window layers for every full-attention layer. This is the standard trade in long context: full attention over 256K tokens is quadratic and expensive, so most layers attend to a local window while a minority retain global attention to keep the model coherent across a long document or codebase. Three-to-one means global attention is the minority, which is the aggressive end of that trade and explains how the cost advantage holds up at 128K context.

The training side is described more briefly: a scalable reinforcement learning framework that drives consistent self-improvement. That phrase is doing a lot of work without specifics, and the technical report PDF in the repository is where the detail would live if you want it.

One number worth isolating, because it is the most concrete capability claim and the one most easily compared across model launches: 74.4% on SWE-bench Verified and 51.0% on Terminal-Bench 2.0. The second figure is the more interesting of the two, since terminal work is closer to what an agent actually does than a patch benchmark.

## Reading the benchmark table honestly

The comparison table is wide and the footnotes at the bottom change how you should read it. Both need attention.

First, the footnotes. An em dash means the score is not publicly available or not tested. An asterisk means the original score was inaccessible or lower at the time StepFun gathered the data. That second convention is unusual and it is favorable to the competition in a specific way: where a rival's number could not be found, Step did not quietly drop the row or assume parity, and where the available number was worse they still used it. The table has asterisks against DeepSeek V3.2, GLM-4.7 and MiniMax M2.1 on most reasoning rows. Kimi's figures appear as two values separated by a slash, corresponding to K2 Thinking and K2.5.

Second, the rows where Step 3.5 Flash leads and the rows where it does not, because a table where one column wins everything is usually a table built for one conclusion.

On agentic work it is at or near the top throughout. Tau-squared-Bench is 88.2 against 87.4 for GLM-4.7 and 86.6 for MiniMax M2.1. xbench-DeepSearch for May 2025 is 83.7 against 78.0 and 76.0. GAIA with no file access is 84.5 against 75.1, which is the widest gap in the table. ResearchRubrics is 65.3 against 62.0.

On BrowseComp the story is closer. Plain BrowseComp is 51.6, and GLM-4.7 posts 52.0, so Step does not lead that row. BrowseComp-ZH is 66.9 against 66.6 for GLM-4.7, effectively level.

On reasoning it leads nearly every row: AIME 2025 at 97.3 against 96.1 for Kimi K2.5, HMMT February 2025 at 98.4 against 95.4, IMOAnswerBench at 85.4 against 81.8. MiniMax M2.1 is notably behind here, at 83.0 on AIME and 60.4 on IMOAnswerBench.

On coding, LiveCodeBench-V6 is 86.4 against 85.0, SWE-bench Verified is 74.4 against 74.0 for MiniMax, and Terminal-Bench 2.0 is 51.0 against 50.8 for Kimi K2.5. Those last two are within half a point of the field, which is a more useful fact than the SWE-bench headline suggests.

There is also a second pair of xbench-DeepSearch rows, for May 2025 and October 2025, and the score drops from 83.7 to 56.3. A twenty-seven point move on the same benchmark name is a change in the benchmark rather than a change in the model, and it is a useful reminder that agentic search benchmarks were rewritten between those two dates. Comparing a current model to a figure produced before that rewrite is not a like-for-like comparison.

The chart above the table includes shadowed bars for enhanced performance using Parallel Thinking, linked to a separate arXiv paper. So the bar chart shows two numbers per model while the table shows one. Decide which you are reading before you quote a figure.

## Scaffolding moves the agentic numbers more than the model does

Three rows in the agentic section describe the same benchmark measured with and without a component called Context Manager, and the gaps are larger than most of the gaps between models.

BrowseComp goes from 51.6 to 69.0. BrowseComp-ZH goes from 66.9 to 73.7. Those are increases of 17.4 points and 6.8 points respectively, from adding scaffolding around the model rather than changing the model.

For comparison, the entire spread on SWE-bench Verified between Step 3.5 Flash at 74.4 and the lowest listed competitor at 73.1 is 1.3 points. So on agentic tasks, the choice of harness moved the number several times further than the choice of model.

This is not an argument against the model. It is an argument about what the benchmark measures. BrowseComp asks a question that cannot be answered from a model's weights alone, since the agent has to search, read several sources, reconcile them and answer. A context manager that handles what the agent has gathered is doing real work that a raw completion pass cannot. The honest reading is that the 69.0 figure describes Step 3.5 Flash plus a good harness, and the 51.6 figure describes Step 3.5 Flash without one.

The comparison rows use the same convention where they exist at all. DeepSeek V3.2 goes from 51.4 to 67.6 with a Context Manager, GLM-4.7 from 52.0 to 67.5, Kimi K2.5 from 60.6 to 74.9. So the scaffolding helps every model by a similar order of magnitude, and Kimi K2.5 leads both the plain and the augmented BrowseComp rows. That pattern is the more informative one: the harness is a shared lever, and the model differences underneath it are a couple of points.

Only BrowseComp-ZH with a Context Manager has a Step entry, 73.7, with dashes for every other model. There is no comparison available there, which means that particular 6.8-point improvement cannot be attributed to the model rather than the harness.

## Deployment: local hardware, vLLM, llama.cpp, and cookbooks

The README claims accessible local deployment on high-end consumer hardware, naming a Mac Studio M4 Max and an NVIDIA DGX Spark as examples, with the stated benefit being data privacy without giving up performance. Treat that as a claim about hardware that is expensive in a specific way rather than a general statement about consumer machines, since a 196B model at any usable quantization needs memory that those two products happen to have.

The repository contents tell you which runtimes are supported. There is a `llama.cpp/` directory, which suggests CPU and Metal paths through that runtime. There is `step3.5_vllm_v0.15.1.Dockerfile` together with `step3.5_vllm_v0.15.1.patch`, which is a vLLM container pinned to a specific version plus a patch applied on top of it. Pinning a version and shipping a patch is how you deploy a model whose architecture support has not yet landed upstream, and the version number in both filenames is the one to match if you build it yourself. There is also `step_3p5_flash_tech_report.pdf` at the root, which is where the architecture details behind the claims above would be documented.

The README's own table of contents covers architecture details, quick start, local deployment, use on agent platforms, cookbooks, known issues and future directions, a co-development section and licensing, so the sections beyond what the page shows are documented rather than absent.

Agent platform coverage is where the cookbooks come in. Four are linked at the top of the README: an OpenClaw guide, a Claude Code best practices guide, a Roo Code integration guide, and a local agent guide for macOS. Four cookbooks for four specific agent runtimes is a deliberate choice about who the model is for, and the answer is coding agents rather than chat interfaces.

Access points are broad. Weights are on Hugging Face under stepfun-ai, with ModelScope as a second host. There is a Discord, an arXiv paper at 2602.10604, and a blog post on StepFun's own site which is also listed as the repository homepage. A badge links to a chat on OpenRouter and the model slug there carries a `:free` suffix, which indicates an OpenRouter-hosted free tier. Another badge links to a Hugging Face Space for trying the model without local hardware.

Licensing is Apache-2.0, stated in the repository metadata, in the YAML front matter of the README, and in a badge. For a model of this size that is the most consequential single fact in the repository, because Apache-2.0 places no restriction on commercial use of the weights.

## Conclusion

Step 3.5 Flash makes a coherent bet rather than a collection of features. Put 196B parameters behind an 11B active path, predict three tokens per forward pass, and spend one full-attention layer for every three sliding-window layers, and you get a model whose serving cost is closer to a mid-sized dense model than to a frontier MoE. The table backs that up on cost more convincingly than on capability: at 128K context on a Hopper GPU the decoding cost is given as 1.0x against 6.0x for DeepSeek V3.2 and 18.9x for both Kimi K2 Thinking and GLM-4.7. Capability is where the footnotes do the work. Asterisks across the comparison table mark scores that were inaccessible or lower when Step's team gathered them, which means the comparisons are honest where they are unstarred and conservative where they are not, but it also means the ranking depends on whose number was available. Two practical notes before you plan around it. The Context Manager rows move BrowseComp from 51.6 to 69.0 and BrowseComp-ZH from 66.9 to 73.7, so agentic scores here describe a scaffold as much as a model, and the gap between the with and without columns is a better measure of scaffolding value than of raw capability. And the repository carries a llama.cpp directory, a vLLM 0.15.1 Dockerfile and a matching patch, so local deployment is a documented path rather than an aspiration. Apache-2.0 with weights published on Hugging Face and ModelScope means the commercial question has a clear answer.

## FAQ

### How fast is 3.5 flash?

The README states 100 to 300 tokens per second in typical use, peaking at 350 tokens per second for single-stream coding tasks. The speed comes from three-way Multi-Token Prediction, which emits three tokens per forward pass instead of one. The comparison table gives 100 tok/s with MTP-3 and EP8 at 128K context on a Hopper GPU, at a decoding cost of 1.0x against 6.0x for DeepSeek V3.2 and 18.9x for Kimi K2 Thinking and GLM-4.7.

### How many parameters does Step 3.5 Flash activate per token?

11B of 196B. It is a sparse Mixture of Experts model that selectively activates only 11B of its 196B parameters for each token, which the README calls intelligence density. The comparison table lists MiniMax M2.1 at 10B activated out of 230B and MiMo-V2 Flash at 15B out of 309B, so Step sits in a similar active-parameter range with more total capacity.

### How do I run Step 3.5 Flash locally?

The repository ships a llama.cpp directory and a vLLM container pinned to version 0.15.1, step3.5_vllm_v0.15.1.Dockerfile, together with a matching step3.5_vllm_v0.15.1.patch to apply on top. The README names a Mac Studio M4 Max and an NVIDIA DGX Spark as target hardware for local deployment. Weights are published on Hugging Face and ModelScope, and a Hugging Face Space is available for trying the model without running it locally.

### What benchmark scores did Step 3.5 Flash reach?

74.4% on SWE-bench Verified and 51.0% on Terminal-Bench 2.0, which is the coding pair. On reasoning it reports 97.3 on AIME 2025 and 98.4 on HMMT February 2025. On agentic tasks it reports 88.2 on Tau-squared-Bench, 51.6 on BrowseComp rising to 69.0 with a Context Manager, 66.9 on BrowseComp-ZH rising to 73.7, and 83.7 on the May 2025 version of xbench-DeepSearch. The README footnotes that an asterisk on a competitor score means it was inaccessible or lower when gathered.

## Sources

- [Issues](https://github.com/stepfun-ai/Step-3.5-Flash/issues)
- [License: Apache-2.0](https://github.com/stepfun-ai/Step-3.5-Flash/blob/main/LICENSE)
- [Project website](https://static.stepfun.com/blog/step-3.5-flash/)
- [README](https://github.com/stepfun-ai/Step-3.5-Flash/blob/main/README.md)
- [stepfun-ai/Step-3.5-Flash on GitHub](https://github.com/stepfun-ai/Step-3.5-Flash)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/stepfun-ai-step-3-5-flash
