Model or dataset
QwenLM/Qwen3-Coder avatar
QwenLM/Qwen3-Coder

Qwen3-Coder: Alibaba's Open-Weight Code Model for Agentic Tasks

Qwen3-Coder is the code version of Qwen3, the large language model series developed by Qwen team.

16,840 stars1,256 forksPythonLicense varies

At a glance

What is it?
Qwen3-Coder is a family of open-weight language models from the Qwen team, built for coding agents and local development. The largest variant offers 256K native context and is designed for repository-scale code understanding, but the GitHub repository last received a push on 2026-03-24.
Who is it for?
Qwen3-Coder suits developers who want an open-weight coding model they can run locally or serve with vLLM, particularly for agentic workflows that involve tool calls and long-context repository understanding. The 30B-A3B variant is the practical choice for local hardware, while the 480B-A35B version requires substantial infrastructure.
Can I use it commercially?
Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
Is it still maintained?
Activity is slowing. The repository last received commits 6 months ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Qwen3-Coder Is and Who It Targets

Qwen3-Coder is a family of large language models developed by the Qwen team, designed for coding agents and local development rather than general text generation. The README positions it as a model for agentic tasks: coding agents that call tools, navigate repositories, and interact with environments, rather than a model you primarily prompt in a chat interface.

Three main variants are available. Qwen3-Coder-480B-A35B-Instruct is the largest, using a mixture-of-experts architecture with 480 billion total parameters and 35 billion active. Qwen3-Coder-30B-A3B-Instruct is a smaller option with 30 billion total parameters and 3 billion active parameters, aimed at teams with limited inference budgets. Qwen3-Coder-Next is built on the Qwen3-Next-80B-A3B-Base architecture with hybrid attention and MoE, and the README describes it as achieving results comparable to Claude Sonnet on agentic coding and agentic browser-use tasks.

The README states the model was trained on large-scale executable task synthesis, environment interaction, and reinforcement learning to develop coding and agentic capabilities. The intended users are developers building agent frameworks, engineers who need a capable local coding assistant, and teams looking for an open-weight alternative to hosted coding models.

Architecture and Context Length

All Qwen3-Coder variants support a native context window of 256K tokens. The README notes this can be extended to up to 1 million tokens using Yarn, and describes the design as optimized for repository-scale understanding. A 256K token window allows the model to hold hundreds of source files in context simultaneously, which is relevant for tasks like cross-file refactoring, understanding large API surfaces, or following complex dependency chains.

The model supports 358 coding languages according to the README, ranging from mainstream languages like C, C++, Python, Java, and JavaScript to specialized languages including ABAP, Agda, Alloy, and Brainfuck. This breadth covers most professional and research programming contexts.

The hybrid attention and MoE architecture in the Qwen3-Coder-Next variant is described as enabling a favorable tradeoff: the README states it achieves results comparable to much larger dense models on agentic coding benchmarks while requiring significantly lower inference costs due to the sparse activation of the MoE layers. The total parameter count is larger than the active parameter count, which is the key property of MoE models.

Getting the Model and Its Dependencies

Model weights are distributed through Hugging Face and ModelScope, not as a pip package. The README directs users to search for checkpoints with names starting with Qwen3-Coder- in the Qwen organization on either platform. Several weight formats are available, including the base instruct versions, FP8 quantized variants for reduced memory use, and GGUF variants for local inference with tools like Ollama.

To run the model programmatically, the repository lists these core dependencies in requirements.txt:

code
torch
transformers
accelerate
safetensors
vllm

Install them with:

bash
pip install torch transformers accelerate safetensors vllm

The README notes that Qwen3-Coder function calling relies on an updated tool parser in both SGLang and vLLM. The relevant parser is linked from the Hugging Face model card. Teams using tool calls must ensure they are using the new version, as the standard tool parsers in these frameworks predate the Qwen3 tokenizer updates.

For local inference without a GPU server, the GGUF variants published on Hugging Face allow loading the model through llama.cpp-compatible tools including Ollama. The README lists Qwen3-Coder-Next-GGUF as an available checkpoint.

Agentic Coding: What Hands-Free Mode Means in Practice

The README describes several concrete use-case examples: building a working website from a description, organizing a cluttered desktop, building a game from scratch, and producing ASCII art. These examples reflect the model's intended deployment mode: an agent that takes a high-level task, uses tools to gather information, writes code, and iterates without manual step-by-step prompting.

The model is described as supporting platforms including Qwen Code, CLINE, and Claude Code, and uses a specially designed function call format. This means it can be dropped into existing agentic tool-call frameworks that support OpenAI-compatible function calling, as long as the updated tool parser is in place.

The README flags fill-in-the-middle capability, which is distinct from the standard generate-from-left mode. Fill-in-the-middle lets the model generate code that fits between an existing prefix and an existing suffix, which is useful for code editing tasks where you want to fill a gap in a known structure rather than generate from scratch. This is a standard feature for code-focused models.

Where Qwen3-Coder Falls Short

The GitHub repository last received a push on 2026-03-24. The repository contains no published GitHub releases. For a model repository rather than a framework, this does not necessarily mean the model weights are outdated, but it does mean the example code, fine-tuning utilities, and evaluation scripts in the repository have not been updated since that date.

The license for the model weights is listed as unknown in the repository metadata. Teams planning to deploy Qwen3-Coder in a product should verify the specific license terms on the Hugging Face model card before building on it, since open-weight models vary significantly in their commercial use terms.

Function calling requires a specific updated tool parser that is not bundled in the current stable releases of SGLang or vLLM. The README warns about this and links to the relevant update, but it adds a setup step that is easy to miss. Deploying with the standard tool parser produces incorrect function call behavior.

The 480B-A35B variant is impractical on a single consumer or workstation GPU. Even with FP8 quantization, this model requires multi-GPU server infrastructure. The 30B-A3B variant is the realistic option for most local deployments.

Qwen3-Coder vs Smaller or Earlier Code Models

The README positions Qwen3-Coder against Claude Sonnet on agentic benchmarks, stating that Qwen3-Coder-Next achieves comparable results. The agentic tasks tested include agentic coding and agentic browser-use. The README does not include raw benchmark numbers for the agentic task comparisons; it frames the comparison in terms of the efficiency-performance tradeoff.

Within the Qwen3-Coder family itself, the 30B-A3B-Instruct and the 480B-A35B-Instruct versions serve different hardware constraints. The 30B model requires far less GPU memory and can run on a single high-end GPU, while the 480B model needs a multi-GPU setup. The MoE architecture means that runtime cost scales with active parameters (3B or 35B) rather than total parameters, making both models cheaper to run than their total parameter counts suggest.

Developers already using Qwen2.5-Coder, which appears in related search queries, are the most direct upgrade path. The README does not document specific differences from that earlier model, beyond that Qwen3-Coder represents a newer generation with longer context support and the hybrid attention architecture.

Fine-tuning and Evaluation Infrastructure

The repository includes a finetuning directory and a qwencoder-eval directory alongside the main README and example code. This structure suggests the repository is intended as a resource for researchers and developers who want to adapt the model or reproduce evaluation results, not just as documentation for end users.

The finetuning scripts are present in the repository for teams that want to adapt Qwen3-Coder to a specific codebase, domain, or tool-calling schema. The README does not document the fine-tuning process, so teams pursuing this should consult the finetuning directory directly.

The model's open-weight status means weights can be downloaded and hosted on private infrastructure without routing requests through an external API. For teams with data privacy requirements or those who need predictable inference latency, self-hosted deployment via vLLM is the standard approach given the dependencies listed in requirements.txt.

Editorial conclusion

Qwen3-Coder suits developers who want an open-weight coding model they can run locally or serve with vLLM, particularly for agentic workflows that involve tool calls and long-context repository understanding. The 30B-A3B variant is the practical choice for local hardware, while the 480B-A35B version requires substantial infrastructure. Teams building on this model should note that the GitHub repository last received a push on 2026-03-24, which is more than six months before this writing, and that the function calling feature requires an updated tool parser in both SGLang and vLLM. Verify the tokenizer is current before deploying, as the README warns that the special tokens and their token IDs were updated to maintain consistency with Qwen3.

Frequently asked questions

Is Qwen3-Coder good at coding tasks?

According to the README, Qwen3-Coder-Next achieves results comparable to Claude Sonnet on agentic coding and agentic browser-use benchmarks. It supports 358 programming languages and a 256K token native context window, allowing it to handle repository-scale tasks.

Can I run Qwen3-Coder locally?

Yes. Model weights are available as GGUF variants on Hugging Face, which can be loaded by Ollama and other llama.cpp-compatible tools for local inference. The 30B-A3B variant is the practical choice for single-GPU setups; the 480B-A35B model requires multi-GPU infrastructure.

How do I install Qwen3-Coder in Ollama?

Qwen3-Coder-Next-GGUF weights are published on Hugging Face and ModelScope. Download the GGUF file and load it through Ollama or any compatible llama.cpp-based tool. The repository does not include Ollama-specific setup instructions, so follow Ollama's standard procedure for loading a custom GGUF model.

How do I use Qwen3-Coder?

Download a checkpoint from Hugging Face or ModelScope, install the dependencies listed in requirements.txt (torch, transformers, accelerate, safetensors, vllm), and serve the model with vLLM. For tool-calling or agent use, install the updated tool parser referenced in the README before running function call workflows.

How do I use Qwen3-Coder in VS Code?

The README lists VS Code as one of the supported platforms via integration with tools like CLINE and Qwen Code. Install the relevant VS Code extension and configure it to point at a locally running vLLM server or the hosted API endpoint. The repository does not provide step-by-step VS Code setup instructions.

Official sources

  1. Issues
  2. QwenLM/Qwen3-Coder on GitHub
  3. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/qwenlm-qwen3-coder.svg)](https://hysenlabs.com/projects/qwenlm-qwen3-coder)