Model or dataset
QwenLM/Qwen3.8 avatar
QwenLM/Qwen3.8

Qwen3.8: Alibaba's Open-Weight LLM Series with Flexible Reasoning Control

Qwen3.8 is the large language model series developed by Qwen team, Alibaba Group.

4,205 stars317 forksUnknownApache-2.0

At a glance

What is it?
Qwen3.8 is the latest series from Alibaba's Qwen team, releasing open-weight language models under Apache-2.0. The repository covers three successive generations: Qwen3.5 (multimodal, 201 languages, sparse MoE), Qwen3.6 (agentic coding focus, thinking preservation), and Qwen3.8 (announced as bringing a Qwen-Max-class model to open release, with tunable reasoning depth).
Who is it for?
Developers and researchers who want to run a high-capability, open-weight language model locally or integrate one into an existing pipeline will find Qwen3.8 directly deployable through Hugging Face Hub or ModelScope. The Apache-2.0 licence allows commercial use.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 45 days ago.
What is it written in?
GitHub does not report a main language for this repository.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What This Repository Covers

The QwenLM/Qwen3.8 repository is the official information hub for three successive model generations: Qwen3.5, Qwen3.6, and Qwen3.8. It contains the README, a licence file, and GitHub workflows. The model weights themselves are not stored in this repository; they are hosted on Hugging Face Hub at huggingface.co/Qwen and on ModelScope at modelscope.cn.

Developers use this repository to understand the capabilities of each generation, find the correct model IDs for loading weights, and track news about new releases. Issues and discussions are active; the README directs users to post questions in the GitHub Issues tab and share ideas in Discussions.

The repository is not a training codebase and does not contain inference code. It is a release announcement and documentation hub. Inference integrations are handled by third-party frameworks such as vLLM, SGLang, and the standard Hugging Face transformers library.

The Qwen3.8 Release: Open Weights at Qwen-Max Level

The README describes Qwen3.8 as bringing a Qwen-Max-class model to open release for the first time. Qwen-Max is Alibaba's closed proprietary model; releasing weights at that capability level as Apache-2.0 is the headline claim for this generation.

Two Qwen3.8 models were released: Qwen3.8-27B on 2026-08-14 and Qwen3.8-2.4T-A95B on 2026-08-12. The 2.4T-A95B designation indicates 2.4 trillion total parameters with 95 billion active parameters per forward pass, a pattern consistent with a sparse mixture-of-experts architecture.

The README lists four improvements in Qwen3.8. Core capabilities are described as comprehensive improvements across coding, professional work, research, and long-horizon agentic tasks. Agent execution is described as stronger autonomous planning and better handling of environment feedback, leading to more reliable end-to-end task completion. Downstream compatibility covers broader support for popular development tools. Flexible thinking control introduces reasoning depth as a configurable parameter.

Downloading and Running Qwen3.8 Models

Model weights are available on Hugging Face Hub and on ModelScope. Most LLM frameworks support loading from Hugging Face by specifying a model ID:

bash
# Example model IDs for Qwen3.8
# Qwen/Qwen3.8-27B
# Qwen/Qwen3.8-2.4T-A95B

The README states that models can also be downloaded manually using `huggingface download` or `git clone` from the model page. For users who cannot access Hugging Face Hub, ModelScope is the recommended alternative. Supported frameworks can switch to ModelScope by setting an environment variable:

bash
SGLANG_USE_MODELSCOPE=true
VLLM_USE_MODELSCOPE=true

The README directs readers to check the individual model cards on Hugging Face Hub for detailed benchmark results, hardware requirements, and usage examples. Benchmark data is not reproduced in the repository README itself.

Qwen3.5 Architecture: Vision-Language and 201 Languages

Qwen3.5 is the foundation generation that introduced several architectural changes. The README describes four enhancements. A unified vision-language foundation uses early fusion training on multimodal tokens to achieve what the README calls cross-generational parity with Qwen3 across reasoning, coding, agents, and visual understanding.

The efficient hybrid architecture combines Gated Delta Networks with sparse Mixture-of-Experts, described as delivering high-throughput inference with minimal latency overhead. The README also describes scalable reinforcement learning generalization across environments with progressively complex task distributions.

Global linguistic coverage is listed as expanding to 201 languages and dialects. This is a concrete distinction from many other open-weight models, which focus on a handful of high-resource languages. Qwen3.5 models released between 2026-02-16 and 2026-03-02 include sizes from 0.8B to 397B, with an initial flagship MoE model of 397B-A17B.

Qwen3.6: Stability and Thinking Preservation

Qwen3.6 was released in April 2026, building on the Qwen3.5 foundation. The README describes it as prioritising stability and real-world utility, shaped by direct community feedback.

Two specific improvements are listed. Agentic coding is described as the model handling front-end workflows and repository-level reasoning with greater fluency. Thinking preservation is a new feature that retains thinking context across conversation history, described as streamlining iterative development and reducing overhead.

Two Qwen3.6 models were released: Qwen3.6-35B-A3B on 2026-04-16 (35 billion total parameters, 3 billion active) and Qwen3.6-27B on 2026-04-22 (27 billion parameters). The model IDs are Qwen/Qwen3.6-35B-A3B and Qwen/Qwen3.6-27B on Hugging Face Hub.

Reasoning Control: reasoning_effort and preserve_thinking

Qwen3.8 introduces two parameters for controlling reasoning behaviour. The reasoning_effort parameter tunes the depth of reasoning applied to a query. The README describes this as allowing the user to trade off compute against reasoning quality depending on the task.

The preserve_thinking parameter retains reasoning context from historical messages in a conversation. This extends the thinking preservation concept introduced in Qwen3.6, making it available to the full Qwen3.8 generation. For iterative tasks, such as debugging a file across multiple turns, retaining thinking context across turns avoids the model re-deriving conclusions it has already reached.

The README does not document the API syntax for these parameters. Developers would need to consult the model cards on Hugging Face Hub or the inference framework documentation for the exact call signatures.

How Qwen3.8 Compares to Meta Llama

Meta Llama is another well-known family of open-weight language models released by Meta AI, available for download from Hugging Face Hub. Both Qwen3.8 and Llama are open-weight models suitable for local deployment and research.

The qualitative difference in approach is the training focus. The Qwen3.8 README emphasises support for 201 languages and dialects, multimodal vision-language capabilities introduced in Qwen3.5, and an architecture that combines Gated Delta Networks with sparse MoE. The Qwen series originates from Alibaba, which has strong research output in multilingual and East Asian language processing.

Developers whose applications require strong multilingual coverage, particularly for East Asian languages, will find Qwen3.8 a natural first choice to evaluate. For English-primary use cases, both families are worth comparing on the specific tasks that matter to the application.

Maintenance Status and Licence

The last push to the repository was on 2026-08-17. The repository is not archived. There are no GitHub releases tagged in this repository; model weight releases are announced via the README News section and hosted on Hugging Face Hub and ModelScope. The licence is Apache-2.0, which permits commercial use and modification.

The repository structure is minimal: README.md, LICENSE, and a .github/ directory. All substantive technical documentation lives in the model cards on Hugging Face Hub. Issues and Discussions on the GitHub repository serve as the community support channel.

The Qwen team is part of Alibaba Group. The release cadence visible in the News section shows model releases roughly every two to three months across the Qwen3.x series.

Editorial conclusion

Developers and researchers who want to run a high-capability, open-weight language model locally or integrate one into an existing pipeline will find Qwen3.8 directly deployable through Hugging Face Hub or ModelScope. The Apache-2.0 licence allows commercial use. The 27B model requires GPU memory in the range typical for that model size; the 2.4T-A95B variant is a very large mixture-of-experts model suited to multi-GPU deployments. Engineers looking for a specific benchmark comparison between Qwen3.8 and competing models should check the Qwen3.8-27B and Qwen3.8-2.4T-A95B model cards on Hugging Face Hub, as the repository README itself defers benchmark data to those pages.

Frequently asked questions

Can I run Qwen3.8 locally?

Yes. The model weights for Qwen3.8-27B and Qwen3.8-2.4T-A95B are available for download from Hugging Face Hub. Most LLM inference frameworks load them by specifying the model ID such as Qwen/Qwen3.8-27B.

How do I install the Qwen3.8-27B model?

Download the weights from Hugging Face Hub using the model ID Qwen/Qwen3.8-27B, either through your inference framework's automatic download or manually via huggingface download or git clone as the README describes. For users who cannot access Hugging Face Hub, ModelScope is the alternative.

What is Qwen3.8?

Qwen3.8 is the latest generation in Alibaba's open-weight LLM series. The README describes it as bringing a Qwen-Max-class model to open release for the first time, with improvements in coding, research, and long-horizon agentic tasks, and a configurable reasoning_effort parameter.

Is Qwen3.8 open source?

The model weights are released under the Apache-2.0 licence, which permits use, modification, and distribution including for commercial purposes. The training code and data are not released through this repository.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. Project website
  4. QwenLM/Qwen3.8 on GitHub
  5. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/qwenlm-qwen3-8.svg)](https://hysenlabs.com/projects/qwenlm-qwen3-8)