Model or dataset
turboderp-org/exllamav2 avatar
turboderp-org/exllamav2

ExLlamaV2: running EXL2 and GPTQ models on consumer GPUs

A fast inference library for running LLMs locally on modern consumer-class GPUs

4,631 stars342 forksPythonMIT

At a glance

What is it?
ExLlamaV2 is a Python inference library for local LLMs on modern consumer GPUs. Its README now says the project is archived in favour of ExLlamaV3, which changes the calculation for anyone choosing it today.
Who is it for?
Adopt ExLlamaV2 if you already run EXL2 weights and want a working Python inference path with a dynamic generator and a mature server option in TabbyAPI. Do not adopt it for a new long-lived deployment, because the README states the project is archived and development continues on ExLlamaV3.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Activity is slowing. The repository last received commits 7 months ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What ExLlamaV2 is for, and who it is not for

ExLlamaV2 is an inference library for running local LLMs on modern consumer GPUs. That sentence sets the audience: people with a single gaming-class card, or a couple of them, who want to load a quantized model and generate text from Python. It is not a hosted service, not a training framework, and not a general model runner that abstracts away the hardware.

The library loads its own EXL2 quantization format as well as the 4-bit GPTQ models the earlier ExLlama supported. EXL2 is described as based on the same optimization method as GPTQ, with support for 2, 3, 4, 5, 6 and 8-bit quantization. That bit-width range is the practical reason people pick it: you choose a quantization that fits your VRAM rather than accepting one fixed size.

The README now opens with a note that the project is archived for now, and that development continues on ExLlamaV3. That is the first thing to weigh. Anyone starting a new project on ExLlamaV2 is adopting a library whose own documentation redirects new work elsewhere. The code still exists, the releases still exist, and TabbyAPI still exists, but the direction of travel is stated plainly.

Who it suits: someone with an existing EXL2 model collection, a Python pipeline that already imports the library, or a TabbyAPI deployment they do not want to rebuild. Who it does not suit: anyone who wants a project with a forward roadmap, and anyone without an NVIDIA GPU, since installation assumes the CUDA Toolkit.

The dynamic generator and how inference actually flows

The centrepiece of the current API is the dynamic generator, introduced in v0.1.0 and described as consolidating all inference, sampling and speculative decoding features of the previous two generators into one API. The one documented exception is FP8 cache; the README points to a Q4 cache mode as supported and performing better, with an evaluation document in doc/qcache_eval.md.

The generator handles single prompts, batches and streaming. In the batched form you pass a list of prompts and get a list of outputs back. Under the hood the README attributes dynamic batching, smart prompt caching, K/V cache deduplication and a simplified API to this generator. Those are the mechanisms that matter for throughput: instead of running prompts one at a time, the library schedules them together and reuses shared prefix state.

Streaming is exposed through an async job object, ExLlamaV2DynamicJobAsync, which you iterate over with async for. Each iteration yields a result dictionary, and the README's example reads the "text" key from it. That is the shape of the data flow: tokenizer encodes the input, the job runs against the generator, and partial results arrive as dictionaries you pull text out of.

Paged attention arrived via Flash Attention 2.5.7 and later. The README states this as a version floor, which is worth noting because it ties the library to a specific external dependency generation. The repository layout reflects the same architecture: a Python package alongside a compiled extension, exllamav2_ext, with separate C++ source files for GEMM, quantization, RoPE, sampling, cache and attention work. The Python side orchestrates; the extension does the arithmetic.

Installing ExLlamaV2 and running a first generation

The README gives three installation routes. Source and PyPI both produce the JIT version, meaning the C++ extension is built the first time the library is used and cached in ~/.cache/torch_extensions. Releases ship prebuilt wheels that already contain the extension binaries.

For the source route you need the CUDA Toolkit and either gcc on Linux or Build Tools for Visual Studio on Windows, plus a matching PyTorch. The README's commands are:

bash
git clone https://github.com/turboderp/exllamav2
cd exllamav2
pip install -r requirements.txt
pip install .

By default that compiles and installs the Torch C++ extension. If you want to skip the build, the README documents an environment variable that switches to the JIT path:

bash
EXLLAMA_NOCOMPILE= pip install .

With the library installed, the repository ships a test script. The README's example passes a model path and a prompt, and notes a flag for splitting across multiple GPUs:

bash
python test_inference.py -m <path_to_model> -p "Once upon a time,"
# Append the '--gpu_split auto' flag for multi-GPU inference

Expect the first run to spend time compiling if you took the JIT route; subsequent runs reuse the cached extension. There is also a console chatbot in the examples directory, invoked with a model path, a prompt format and automatic GPU splitting:

bash
python examples/chat.py -m <path_to_model> -mode llama -gs auto

The -mode argument selects the prompt format. The README says raw produces a chatlog-style chat that works with base models and various finetunes, that -modes lists all available formats, and that -sp supplies a custom system prompt. If you would rather install from a release, the README warns that the wheel must match your platform, Python version, CUDA version and PyTorch version, because the Torch C++ extension ABI breaks with every new PyTorch release.

Where the wheel matching and the archive note bite

The ABI warning is the sharpest practical limitation in the README. A prebuilt wheel is not a portable artifact. It is bound to a Python version, a CUDA version and a PyTorch version at once. Upgrade PyTorch and the wheel you installed may stop loading. The JIT path sidesteps that by building against whatever you have, at the cost of a compile on first use and a toolchain on the machine.

Windows users face the same toolchain requirement as Linux users, just with a different compiler: Build Tools for Visual Studio instead of gcc. The setup script confirms the platform split, choosing /Ox on Windows and -O3 elsewhere for the C++ flags, with a separate set of nvcc flags.

The archive note is the larger limitation. The README states the project is archived for now and that development continues on ExLlamaV3. That is not a deprecation schedule with dates, but it does mean bug reports and feature requests have a stated destination elsewhere. For a research machine or an existing deployment this may not matter. For a product you intend to maintain for years, it is a reason to look at the successor before writing integration code.

One more boundary: the README's performance table is explicitly described as quick tests comparing against ExLlama V1, with the caveat that speeds vary across GPUs and that slow CPUs can still be a bottleneck. It is a comparison against the project's own predecessor, not a neutral benchmark, and the numbers come from the README rather than from any independent measurement.

ExLlamaV2 versus llama.cpp, vLLM and Ollama

The nearest alternative in spirit is llama.cpp, which also targets local inference on modest hardware but takes a different route: it centres on the GGUF format and CPU or mixed CPU/GPU execution, with its own quantization scheme. ExLlamaV2 does not read GGUF. It reads EXL2 and GPTQ. If your model collection is GGUF, ExLlamaV2 is the wrong tool and no amount of configuration changes that. If your collection is EXL2 at 4.0 or 5.0 bits per weight, llama.cpp is the one that needs conversion.

vLLM sits at the other end. It is built for serving many concurrent requests on server-class GPUs, with paged attention and continuous batching as core features. ExLlamaV2 also has paged attention and dynamic batching, but its stated target is modern consumer GPUs, and the README's own numbers are single-card throughput figures. Choosing between them is mostly a question of whether you are serving a queue of users or running a model on your own desk.

Ollama is a different category again: a packaged runner with model management built in. ExLlamaV2 is a library. It gives you a Python API and expects you to bring your own server, which is why the README names TabbyAPI as the official and recommended backend. TabbyAPI provides an OpenAI-compatible API with Hugging Face model downloading, embedding model support and Jinja2 chat templates. There is also ExUI, a standalone single-user web UI, and loaders in text-generation-webui and lollms-webui.

None of these alternatives is a drop-in swap. The format each one reads is the dividing line.

Maintenance, licensing and the upgrade question

The repository is not archived on GitHub itself, and the last push was on 2026-03-04. The most recent release listed is v0.3.2 from 2025-07-13, preceded by v0.3.1 in May 2025 and v0.3.0 in May 2025. So the code has moved since the last tagged release, but the README's own note says the project is archived for now and points to ExLlamaV3.

That combination is the upgrade cost in one sentence: you are adopting a library whose documentation directs future work to a different repository. There is no migration guide in the README, and it does not document a rollback path or a support window. If you build on ExLlamaV2, the practical question is how much of your code touches the library's API surface. The dynamic generator is the current consolidated API, so code written against it is at least written against the newest interface rather than one of the two earlier generators it replaced.

The licence is MIT, which is permissive and places few obligations on how you redistribute or modify the code. That is a fact about the licence text, not advice about your situation; if you are embedding the library in a commercial product, have someone qualified read the licence and the licences of its dependencies. The requirements file pulls in torch, safetensors, numpy pinned to ~=1.26.4, tokenizers, and a handful of others, so your dependency review does not stop at the MIT header.

Editorial conclusion

Adopt ExLlamaV2 if you already run EXL2 weights and want a working Python inference path with a dynamic generator and a mature server option in TabbyAPI. Do not adopt it for a new long-lived deployment, because the README states the project is archived and development continues on ExLlamaV3. Before committing, check that a prebuilt wheel exists for your exact PyTorch, Python and CUDA combination, or plan on the JIT build path and its first-run compile.

Frequently asked questions

How do I install ExLlamaV2?

The README gives three routes: from source with git clone, pip install -r requirements.txt and pip install .; from a GitHub release wheel matched to your platform, Python, CUDA and PyTorch versions; or from PyPI with pip install exllamav2, which is the same as the JIT version. Source installation needs the CUDA Toolkit plus gcc on Linux or Build Tools for Visual Studio on Windows.

What is the difference between ExLlamaV2 and ExLlamaV3?

The README states that ExLlamaV2 is archived for now and that development continues on ExLlamaV3, which it links to. The README does not describe ExLlamaV3's features or provide a migration path, so the only documented difference is where development is happening.

Does ExLlamaV2 work with GGUF models?

The README describes support for GPTQ models and the project's own EXL2 format, with 2, 3, 4, 5, 6 and 8-bit quantization. GGUF is not mentioned in the installation or quantization sections, so the documented formats are EXL2 and GPTQ.

What is the difference between ExLlamaV2 and koboldcpp?

The README does not mention koboldcpp, so no comparison can be made from it. What the README does document is that ExLlamaV2 is a Python inference library for EXL2 and GPTQ weights on consumer GPUs, with TabbyAPI as the official backend server and ExUI as a standalone web UI.

Official sources

  1. Issues
  2. License: MIT
  3. README
  4. Releases
  5. turboderp-org/exllamav2 on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/turboderp-org-exllamav2.svg)](https://hysenlabs.com/projects/turboderp-org-exllamav2)