Model or dataset
QwenLM/Qwen3.8-Flash-Next avatar
QwenLM/Qwen3.8-Flash-Next

Qwen3.8-Flash-Next: a 125B open-weight preview of the Qwen4 architecture

Qwen3.8-Flash-Next is the foundation model developed by Qwen Team, Alibaba Group.

403 stars21 forksUnknownLicense varies

At a glance

What is it?
Qwen3.8-Flash-Next is a multimodal MoE model whose weights are open, positioned as an early look at the attention, residual and embedding design the Qwen4 family will use. It is a serving and research artefact, not a drop-in chat model.
Who is it for?
Adopt Qwen3.8-Flash-Next if you have multi-GPU serving capacity and want to evaluate the GDN plus QSA architecture before the Qwen4 family lands, or if you need a coding and office-task model whose training cost is roughly one ninth of Qwen3.7-Plus by the project's own account. Do not adopt it if you are looking for a single-GPU local chat model, a fine-tuning base with documented recipes, or anything with a published licence and release notes.
Can I use it commercially?
Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
Is it still maintained?
Yes. The repository last received commits 33 days ago.
What is it written in?
GitHub does not report a main language for this repository.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

What Qwen3.8-Flash-Next is, and the job it is meant to do

The README describes Qwen3.8-Flash-Next as a multimodal MoE model that doubles as an early preview of the architecture used in Qwen4. The stated parallel is explicit: it plays the same role Qwen3-Next played for Qwen3.5, where the hybrid Gated DeltaNet plus Gated Attention design was released ahead of the series that would adopt it. The pattern here repeats. The team is publishing architectural changes before the full Qwen4 model family is built on top of them, so the community can examine them first.

The audience follows from that framing. This is not aimed at someone who wants a chat endpoint. It is aimed at engineers who need to read a new attention and residual design against real weights, and at teams that serve large models and care about the cost curve. The headline numbers are a 125B-parameter main model, an additional 51B of N-gram embeddings, and 6B parameters activated per token. Sparse activation is the reason a model of that total size is discussed in serving terms at all. The README also claims that compared with Qwen3.7-Plus, training takes about one ninth as much while delivering superior coding and office-task capability. Treat that as a vendor comparison, not an independent measurement, because the evaluation results live in the blog post rather than in the repository.

Gated DeltaNet, QSA, gated residual and N-gram embedding

Four changes are described, and they are worth separating because they fail in different ways.

Attention is a GDN plus QSA hybrid. Gated DeltaNet compresses history, and Qwen Sparse Attention uses a compressed lightweight indexer to select important context at micro-block granularity. The stated goal is reducing the cost of attention on long sequences. The trade-off is inherent: a lightweight indexer that decides what matters is a learned approximation, and when it misjudges, the tokens it dropped are not recoverable at inference time. Long-context quality on retrieval-style tasks is the place to look for that failure.

Residual is Gated Residual. It widens the residual stream into four branches and controls reads and writes with a dynamic gate. The README frames this as strengthening cross-layer information flow and training stability. This is the least externally visible change of the four, and the hardest to evaluate from a serving endpoint.

Embedding is N-gram Embedding. A table is looked up using local context to scale capacity with very little extra computation, and the README states the table can be offloaded to host memory and overlapped with model computation through asynchronous prefetching. That is a concrete systems claim, and it is also a dependency: if your serving stack does not implement the offload and prefetch path, you pay for 51B of embeddings somewhere in your memory budget.

Optimization is the Muon optimizer, refined around orthogonalization accuracy, the division of labour between Muon and AdamW, and the splitting of fused parameters, with the scaling law refitted. This matters to anyone reproducing training, and to nobody running inference.

Downloading the weights and serving a first request

The README points to two distribution points: the Hugging Face Hub under the Qwen organisation, and ModelScope for users unable to reach Hugging Face. Most frameworks can pull the files automatically from the model ID. For a manual download, the README names `huggingface download` or `git clone`, and on the ModelScope side `modelscope download` or `git clone`, with instructions on the model page.

The ModelScope route is selected through environment variables rather than a different client, which is a small but useful detail when you are containerising a deployment.

bash
SGLANG_USE_MODELSCOPE=true
VLLM_USE_MODELSCOPE=true

The simplest local path documented is Hugging Face Transformers, which the README treats as the model-definition framework and which now includes serving. Launching a server is one command.

bash
# from a Python environment with transformers installed
transformers serve Qwen/Qwen3.8-Flash-Next --port 8000 --continuous-batching

According to the README, an OpenAI-compatible API then becomes available at `http://localhost:8000/v1`. That is the point at which you can point an existing OpenAI client at the endpoint and send a request. The README does not state the memory footprint of this configuration, so what you actually see depends on the hardware you run it on and on whether the framework you use implements the embedding offload.

For production serving, the README demonstrates SGLang, vLLM and TokenSpeed. The SGLang example given is:

bash
sglang serve --model-path Qwen/Qwen3.8-Flash-Next --port 8000

Note that the SGLang snippet in the README is truncated at `--port 8000 --`, so the full flag set for that command is not documented in the repository. Check the SGLang project for the rest. On Apple Silicon, `mlx-vlm` is listed as supporting vision and text, with original checkpoints convertible and quantized MLX versions searchable on Hugging Face. `llama.cpp` is listed as supporting text and vision, with GGUF files on the Hub. Unsloth provides a local UI for running and training, with a linked guide for quants.

Where this model is the wrong tool

The repository is thin in ways that matter for adoption decisions. There are no release notes beyond a single news entry dated 2026-08-26 pointing at the blog, so the repository itself does not document rollback, deprecation or version-to-version behaviour. The README does not state a licence. It does not document a minimum GPU count, a quantisation recipe, or a fine-tuning procedure. The benchmark section is a pointer to the blog rather than a table.

That combination rules out several use cases. If you need to fine-tune with a documented recipe, this repository does not give you one. If you need a licence you can clear with legal before shipping, the README is silent and you have to read the model page. If you are running on a single consumer GPU, nothing here suggests that is a supported configuration; the 125B figure is the total parameter count, and sparse activation does not remove the need to hold weights. If you need long-context behaviour you can reason about formally, the QSA indexer is a learned selection step, and the repository does not publish the conditions under which it drops context that matters.

The last push to the repository was on 2026-08-27, which is recent, but the repository contents are two files: README.md and tech_report.pdf. There is no code, no configuration, no test suite. Anyone expecting a reference implementation will be disappointed; the implementations live in Transformers, SGLang, vLLM, llama.cpp and mlx-vlm.

How it differs from a dense open-weight release

The most useful comparison is not with another named model but with the shape of release this is. A conventional open-weight drop ships a checkpoint, a licence, an evaluation table and a set of framework integrations, and the architectural discussion is background. Qwen3.8-Flash-Next inverts that: the architecture is the product, and the weights exist so the architecture can be examined.

Against a dense model of comparable activated size, the difference is where the parameters sit. A dense model spends its capacity uniformly on every token. Here, 6B parameters activate per token out of a 125B main model plus 51B of N-gram embeddings, and the README's efficiency argument rests on that gap. The cost of the gap is that the effective capacity depends on routing and on the embedding lookup, both of which are framework-visible behaviour rather than fixed weights. Two serving stacks that implement the N-gram embedding offload differently can produce different memory profiles from the same checkpoint.

Against Qwen3.7-Plus, the README makes a direct claim: about one ninth the training cost with superior coding and office-task results. It is worth reading the blog for the evaluation setup, because the repository does not reproduce it. The honest position is that the efficiency direction is plausible from the architecture, and the magnitude is a vendor number.

Maintenance, upgrade surface and licence exposure

The repository's last push was on 2026-08-27. It is not archived. There are no retrieved releases, so there is no version history to track and no changelog to diff against. In practice this means the upgrade surface is not the repository at all: it is the inference frameworks. When Transformers, SGLang or vLLM change how they handle the GDN plus QSA hybrid or the N-gram embedding table, that is where behaviour shifts, and the repository will not tell you. Pin your framework versions the way you would pin any dependency, because the checkpoint alone does not determine runtime behaviour.

The licence is the larger open question. The README does not state one, and the repository metadata does not either. The weights are hosted on Hugging Face and ModelScope, so the model page is where terms would appear, and that is the document to read before any commercial deployment. Nothing in the repository describes redistribution terms, attribution requirements, or acceptable-use restrictions. This is not a legal opinion, and it is not a reason to avoid the model; it is a reason to resolve the question from the model page rather than from the GitHub README, which is silent.

Editorial conclusion

Adopt Qwen3.8-Flash-Next if you have multi-GPU serving capacity and want to evaluate the GDN plus QSA architecture before the Qwen4 family lands, or if you need a coding and office-task model whose training cost is roughly one ninth of Qwen3.7-Plus by the project's own account. Do not adopt it if you are looking for a single-GPU local chat model, a fine-tuning base with documented recipes, or anything with a published licence and release notes. Before committing, verify three things: the exact licence terms on the Hugging Face or ModelScope model page, which inference framework in your stack already lists Qwen3.8-Flash-Next support, and whether the N-gram embedding offload path is exposed by that framework, since that is the mechanism the README credits for the capacity gain.

Frequently asked questions

What is Qwen3.8-Flash-Next?

It is a multimodal MoE foundation model from the Qwen Team, released as an early preview of the architecture used in Qwen4. The README describes it as a 125B-parameter main model with an additional 51B of N-gram embeddings and 6B parameters activated per token.

How do I download Qwen3.8-Flash-Next?

The weights are on the Hugging Face Hub and on ModelScope. Most frameworks download them automatically from the model ID Qwen/Qwen3.8-Flash-Next, or you can fetch them manually with huggingface download, modelscope download or git clone.

Which frameworks support serving Qwen3.8-Flash-Next?

The README demonstrates SGLang, vLLM and TokenSpeed for deployment, and lists Hugging Face Transformers, llama.cpp, mlx-vlm for Apple Silicon, and Unsloth for local use. The README's SGLang command is truncated, so check that project for the full flag set.

What licence is Qwen3.8-Flash-Next released under?

The README does not state a licence, and the repository metadata does not either. The terms would appear on the Hugging Face or ModelScope model page, so that is the document to check before deploying.

Official sources

  1. Issues
  2. Project website
  3. QwenLM/Qwen3.8-Flash-Next on GitHub
  4. README
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/qwenlm-qwen3-8-flash-next.svg)](https://hysenlabs.com/projects/qwenlm-qwen3-8-flash-next)
Community notes

Community notes