Model or dataset
QwenLM/Qwen avatar
QwenLM/Qwen

QwenLM/Qwen: the original Qwen repo, what still works and where to get its successor

The official repo of Qwen (通义千问) chat & pretrained large language model proposed by Alibaba Cloud.

21,881 stars1,977 forksPythonApache-2.0

At a glance

What is it?
The QwenLM/Qwen repository ships the 1.8B to 72B Qwen and Qwen-Chat weights, a CLI demo, a web demo and an OpenAI-compatible API server. The README now states the repo is no longer actively maintained, so the decision is whether you have a reason to stay on this generation.
Who is it for?
Adopt QwenLM/Qwen when you need to reproduce or fine-tune the original Qwen and Qwen-Chat checkpoints, or when you specifically want the 1.8B, 7B, 14B or 72B sizes with their GPTQ and KV cache quantization notes. Do not adopt it for new deployments: the README states the repository is no longer actively maintained and points readers to QwenLM/Qwen2, and the pinned transformers range in requirements.txt will fight any current environment.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Activity is slowing. The repository last received commits 6 months ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

What QwenLM/Qwen actually contains, and who it was built for

This is the release repository for the first Qwen generation from Alibaba Cloud: four base models (Qwen-1.8B, Qwen-7B, Qwen-14B, Qwen-72B) and four chat models (Qwen-1.8B-Chat, Qwen-7B-Chat, Qwen-14B-Chat, Qwen-72B-Chat), each also published in Int4 and Int8 quantized variants. The README's model table lists the release dates as 23.08.03 for Qwen-7B, 23.09.25 for Qwen-14B, and 23.11.30 for both Qwen-1.8B and Qwen-72B. Pretrained token counts range from 2.2T for the smallest model to 3.0T for the 14B and 72B models, and the context length is 32K for everything except Qwen-14B, which the table lists at 8K.

The intended audience is narrow and practical. Someone who wants a Chinese-and-English model they can run on their own hardware, fine-tune with Q-LoRA, or serve behind an OpenAI-compatible endpoint. The README states the models were pretrained on multilingual data with a focus on Chinese and English, and that the chat models were aligned with human preference using SFT and RLHF. The repository ships the surrounding tooling as well: cli_demo.py, web_demo.py, openai_api.py, finetune.py, run_gptq.py and an eval/ directory. If you are choosing a model today rather than reproducing a result from 2023, the README's own notice at the top sends you to QwenLM/Qwen2 instead.

How the pieces fit together: weights, tokenizer, and the serving path

The architecture is a standard decoder-only transformer stack, but the repository's real mechanism is the toolchain around it. Weights live on Hugging Face and ModelScope, not in the repository. The tokenizer is not a plain SentencePiece model: the repository includes examples/qwen_extra.tiktoken and examples/qwen_extra_vocab.txt, and tokenization_note.md is a dedicated document explaining the tokenization scheme, which is a detail most model repos do not bother to document.

The serving path has three distinct entry points. cli_demo.py is the minimal interactive loop, useful for confirming that a checkpoint loads at all. web_demo.py is a Gradio interface, and it needs requirements_web_demo.txt rather than the main requirements file, which is a split people miss. openai_api.py exposes the model behind an OpenAI-compatible HTTP API so existing client code can point at a local endpoint. For throughput, vllm_wrapper.py in examples/ wraps the model for vLLM. The README also points at GPTQ quantization through run_gptq.py and KV cache quantization, both aimed at reducing the memory footprint the model table quantifies: minimum GPU memory for generating 2048 tokens with Int4 ranges from 2.9GB on Qwen-1.8B to 48.9GB on Qwen-72B.

The agent and tool-use material lives in examples/ rather than in the core code. function_call_examples.py, function_call_finetune_examples.py, react_demo.py with react_prompt.md, and langchain_tooluse.ipynb cover function calling and ReAct-style loops. The model table marks tool usage as supported across all four sizes.

Installing QwenLM/Qwen and running a first prompt

The repository gives requirements.txt at the root. Its most consequential line is the transformers pin: `transformers>=4.32.0,<4.38.0`. That upper bound is the first thing to check against your environment, because a current transformers install will violate it. Install into a fresh virtual environment rather than an existing one.

bash
pip install -r requirements.txt

That file also pulls accelerate, tiktoken, einops, scipy and transformers_stream_generator==0.0.4. The stream generator is pinned to an exact version, so do not let a resolver upgrade it.

With dependencies in place, the CLI demo is the shortest path to a working response. The README describes cli_demo.py as the quickstart for simple inference, and it downloads the checkpoint on first run.

bash
python cli_demo.py

You should get an interactive prompt in the terminal. If you want the browser interface instead, it has its own dependency file, so install that first and then launch the script:

bash
pip install -r requirements_web_demo.txt
python web_demo.py

For programmatic use, openai_api.py starts an OpenAI-compatible server, which means an existing client can be repointed at it. The README does not state the port in the excerpt available here, so read the top of openai_api.py for the argument that sets it before wiring anything to it. For fine-tuning, finetune.py is the entry point, and the model table lists minimum Q-LoRA memory from 5.8GB on Qwen-1.8B to 61.4GB on Qwen-72B.

The maintenance notice is the main limitation, not the model quality

The README opens with an important notice stating that this repository is no longer actively maintained, citing substantial codebase differences, and directing readers to QwenLM/Qwen2. The last push recorded for the repository is 2026-03-05, which is more than six months before the current date, so nothing here should be described as under active development. Treat this as a frozen generation. Bug reports and pull requests against cli_demo.py or openai_api.py have no obvious path to a fix.

The practical consequences are concrete. The transformers upper bound of 4.38.0 means the code was written against an API surface that has since moved, and the README does not document a migration path or a tested newer combination. The RLHF stage of alignment is described in the README as not released, so the chat models cannot be reproduced end to end from this repository. There is no rollback or upgrade documentation in the README at all, which matters if you are pinning this in a production image.

The wrong-tool case is any new project. If you are starting a Chinese or English assistant today and have no dependency on these exact checkpoints, the maintenance notice in the README is the project telling you not to begin here. This repository is for reproduction, for fine-tuning runs against the original weights, and for reading how the first Qwen generation handled tokenization and quantization.

QwenLM/Qwen2 versus staying on the original checkpoints

The direct alternative is QwenLM/Qwen2, and the README itself names it. The difference is not a version bump in the same codebase: the README explains the split by citing substantial codebase differences, which is why the two live in separate repositories rather than as branches. That distinction matters for anyone planning to port code. A script written against this repository's openai_api.py or finetune.py is not a drop-in match for the successor repository, and the transformers constraint here will not carry over.

What you give up by moving is access to this specific generation: the 1.8B, 7B, 14B and 72B checkpoints with their published release dates, the tokenization notes, and the GPTQ and KV cache quantization documentation that the README lists as repository contents. What you gain by moving is a codebase that is not carrying a no-longer-maintained notice.

A second, quieter alternative is to skip the chat models and use the base Qwen checkpoints with your own alignment step, since the README states the RLHF stage was not released. That is a real option for teams with their own preference data, but it is a larger project than loading Qwen-7B-Chat and calling it done.

Licensing and the cost of staying on a frozen release

The repository is listed under Apache-2.0, and the LICENSE file is at the root. That is not the whole picture. The root also contains a NOTICE file, a file named Tongyi Qianwen LICENSE AGREEMENT, and a file named Tongyi Qianwen RESEARCH LICENSE AGREEMENT. Three separate agreements plus a NOTICE means the licence that applies to a given checkpoint is not something you can infer from the repository's top-level licence field alone. Check the model card for the exact weights you download, and read the NOTICE. This is a description of what is in the repository, not legal advice; if the distinction between the standard and research agreements affects your use, that is a question for your own counsel.

The upgrade cost is the part worth pricing before you start. Because the README states the repository is no longer maintained and points to QwenLM/Qwen2, there is no upgrade path within this repository. Moving to the successor means re-validating your inference wrapper, your fine-tuning script and your serving configuration against a different codebase, not bumping a version. If you pin this release, budget for that migration as a rewrite of the integration layer rather than a dependency update, and note that the README does not document rollback or an upgrade procedure to soften it.

Editorial conclusion

Adopt QwenLM/Qwen when you need to reproduce or fine-tune the original Qwen and Qwen-Chat checkpoints, or when you specifically want the 1.8B, 7B, 14B or 72B sizes with their GPTQ and KV cache quantization notes. Do not adopt it for new deployments: the README states the repository is no longer actively maintained and points readers to QwenLM/Qwen2, and the pinned transformers range in requirements.txt will fight any current environment. Before committing, verify that the model card for the exact checkpoint you want still resolves on Hugging Face or ModelScope, and confirm which licence file covers it, since the repository root holds three separate agreements.

Frequently asked questions

What is QwenLM/Qwen mainly used for?

It is the release repository for the original Qwen and Qwen-Chat models from 1.8B to 72B, used for local inference, fine-tuning and serving behind an OpenAI-compatible API. The README also documents GPTQ and KV cache quantization for reducing memory use.

Is QwenLM/Qwen totally free?

The repository is listed under Apache-2.0, but the root also contains a NOTICE file and two separate Tongyi Qianwen licence agreements, one of them a research licence. Which agreement covers a given checkpoint has to be checked on that model's card.

Is QwenLM/Qwen better than ChatGPT?

The README does not make that comparison. It reports benchmark results for the Qwen models and states that the chat models were aligned with human preference using SFT and RLHF, with the RLHF stage not released.

Is it safe to use QwenLM/Qwen?

The README does not contain a safety or risk assessment section, so it is silent on this. It does state the chat models are aligned with human preference through SFT and RLHF, and that the RLHF stage has not been released.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. QwenLM/Qwen on GitHub
  4. README
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/qwenlm-qwen.svg)](https://hysenlabs.com/projects/qwenlm-qwen)
Community notes

Community notes