ExLlamaV3: Quantized Local LLM Inference with EXL3 and Parallel Execution
An optimized quantization and inference library for running LLMs locally on modern consumer-class GPUs
At a glance
- What is it?
- ExLlamaV3 is a Python inference library for running large language models on consumer NVIDIA GPUs using EXL3 quantization, a format based on QTIP that achieves high compression with low quality loss. It supports tensor-parallel and expert-parallel inference, CPU offloading with AVX2 and AVX512, and continuous batching, and integrates with TabbyAPI for an OpenAI-compatible server.
- Who is it for?
- ExLlamaV3 is the right choice for users who want to run large LLMs on one or a few consumer NVIDIA GPUs at high quality with the EXL3 quantization format, and who are comfortable with the CUDA build environment or with using prebuilt release wheels. It is not suitable for AMD GPU users, whose hardware is not targeted by the current CUDA-specific kernels, and it is not a drop-in for environments already using GGUF models and llama.cpp without additional conversion work.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What ExLlamaV3 Does and Who Needs It
Running a large language model locally on a consumer GPU involves a fundamental trade-off: model size determines quality, but model size also determines whether the model fits in VRAM. Quantization reduces the memory footprint of a model by representing weights in fewer bits, at a cost in output quality that depends heavily on the quantization method.
ExLlamaV3 addresses this trade-off with EXL3, a quantization format based on QTIP that the library's README describes as providing flexible quantization. The library targets modern consumer-class GPUs, which positions it for hobbyists and researchers who own a gaming or workstation GPU but not data center hardware.
The intended users are people who want to run models that are too large for their VRAM at full precision and want the best possible quality per gigabyte of GPU memory. A secondary audience is developers building inference pipelines who want a local backend with an OpenAI-compatible API surface through the recommended TabbyAPI server.
EXL3 Quantization and the QTIP Foundation
The README identifies EXL3 as based on QTIP, a quantization method. EXL3 supports sub-8-bit weight quantization and is the primary format for model weights in this library. The cache layer separately supports 2 to 8 bit quantization, which reduces the memory used by the KV cache during generation.
The library also includes a convert.py script for quantizing models to EXL3 format. The science/ directory contains the research-oriented tooling: sc_measure.py, sc_optimize.py, sc_rfn_probe.py, and sc_trace.py are utility scripts for analyzing and optimizing quantization.
On Windows, the Triton library is required because the attention, cache, and recurrent kernels are written in Triton. The pyproject.toml lists triton-windows as a Windows-only dependency. On Linux, Triton is already bundled with PyTorch in most CUDA builds. The library does not import without Triton on Windows, so installing without the Windows Triton package results in an import error.
The CUDA requirement is 12.4 or later. PyTorch 2.6.0 or later is required as a dependency. The pyproject.toml flavor extras (cu124, cu126, cu128, cu129, cu130, cu132) correspond to specific CUDA builds of PyTorch.
Installation and Getting Started
The recommended installation path is a prebuilt wheel from the releases page. The README shows the general form:
pip install https://github.com/turboderp-org/exllamav3/releases/download/v0.0.6/exllamav3-0.0.6+cu128.torch2.8.0-cp313-cp313-linux_x86_64.whlThe wheel filename encodes the CUDA version (cu128 in this example) and the PyTorch version (torch2.8.0). Picking the wheel that matches your installed CUDA and Python version avoids the build step entirely.
The PyPI installation also works:
pip install exllamav3However, the README notes that the PyPI package does not include a prebuilt CUDA extension, so it requires the CUDA toolkit and build prerequisites (VS Build Tools on Windows, gcc on Linux, python-dev headers) and compiles the extension on first import.
For source installation with uv, specifying a CUDA flavor extra selects the matching PyTorch build automatically:
git clone https://github.com/turboderp-org/exllamav3
cd exllamav3
uv venv
uv sync --extra cu130The examples/ directory contains scripts for common use cases: chat.py for interactive chat, chat_console.py for a console UI, generator.py for basic generation, async_generator.py for async workflows, and multimodal.py for image-capable models.
Parallel Inference and CPU Offloading
ExLlamaV3 supports two forms of parallel inference: tensor parallelism and expert parallelism. Tensor parallelism splits model layers across multiple GPUs, which allows running a model that is too large for a single GPU across two or more cards. Expert parallelism targets mixture-of-experts (MoE) models, distributing expert layers across GPUs.
For setups with insufficient total GPU memory, the CPU offloading feature moves layers to system RAM and computes them using AVX2 or AVX512 vector instructions. The README describes this as allowing large MoE models to run with limited GPU resources. The performance cost is significant compared to pure GPU execution, but it enables running very large models that would otherwise be impossible on the available hardware.
Continuous and dynamic batching are supported for serving multiple requests simultaneously. Speculative decoding is listed as a generation feature, which accelerates generation by using a smaller draft model to propose tokens that the main model then verifies. The examples/branch_decode/ directory covers speculative and branch decoding use cases.
TabbyAPI Integration and OpenAI-Compatible API
ExLlamaV3 is a library, not a server. To expose a REST API for local LLM inference, the README recommends TabbyAPI as the official and recommended backend server. The README describes TabbyAPI as providing an OpenAI-compatible API, HF model downloading, embedding model support, and HF Jinja2 chat templates, with a startup script that manages prerequisites.
This design separates concerns: ExLlamaV3 provides the inference engine and quantization, while TabbyAPI provides the server layer with API routing, chat template handling, and embedding endpoints. Developers who only need the inference primitives (loading a model, running generation) can use ExLlamaV3 directly through its Python API without TabbyAPI.
A Transformers plugin is also available at examples/transformers_integration.py, which connects ExLlamaV3 models to the Hugging Face Transformers generation API. This is relevant for researchers who have existing code using the Transformers generate() interface and want to swap in ExLlamaV3 as the backend without changing their generation code.
Limitations and Comparison with GGUF and llama.cpp
ExLlamaV3 targets NVIDIA GPUs through CUDA. AMD GPU support is not described in the README. Users with AMD hardware should look at alternatives such as llama.cpp with ROCm support or other ROCm-compatible inference tools.
The EXL3 model format is specific to this library. Models published on Hugging Face in GGUF format (used by llama.cpp and Ollama) require conversion to EXL3 before use with ExLlamaV3. The repository includes convert.py for quantizing from HuggingFace-format models, but there is no direct GGUF-to-EXL3 conversion documented in the README. GGUF models must go through the full HuggingFace format first.
llama.cpp with GGUF quantization is the most common alternative for local inference. It supports both NVIDIA and AMD GPUs, runs on CPU only if needed, has broader model format support out of the box, and has a larger community of prebuilt quantized models available. ExLlamaV3 trades that broad hardware compatibility for more aggressive NVIDIA-specific optimization and the EXL3 quantization format.
The Windows installation requires the triton-windows package in addition to standard dependencies, which adds a build step for Windows users who install from PyPI rather than prebuilt wheels.
Maintenance and License
The latest release is v1.5.3, published on 2026-09-27. Previous releases were v1.5.2 (2026-09-26) and v1.5.1 (2026-09-20), and the last push was on 2026-09-27. The release cadence shows several updates per week in the most recent period, reflecting active development. ExLlamaV3 is licensed under MIT.
Editorial conclusion
ExLlamaV3 is the right choice for users who want to run large LLMs on one or a few consumer NVIDIA GPUs at high quality with the EXL3 quantization format, and who are comfortable with the CUDA build environment or with using prebuilt release wheels. It is not suitable for AMD GPU users, whose hardware is not targeted by the current CUDA-specific kernels, and it is not a drop-in for environments already using GGUF models and llama.cpp without additional conversion work. Verify that CUDA 12.4 or later and PyTorch 2.6 or later are installed before attempting setup. The current release is v1.5.3.
Frequently asked questions
How does ExLlamaV3 compare to Ollama for local LLM inference?
ExLlamaV3 uses the EXL3 quantization format and provides CUDA-optimized kernels for NVIDIA GPUs, with tensor parallelism and speculative decoding for advanced setups. Ollama uses GGUF models and targets broader hardware including CPU-only and AMD GPU setups with a simpler one-command install. ExLlamaV3 requires more setup but targets higher throughput on NVIDIA hardware through its specific quantization and parallel inference design.
Does ExLlamaV3 work on AMD GPUs?
The README does not describe AMD GPU support. The CUDA kernel requirements and the CUDA version specification (12.4 or later) indicate the library targets NVIDIA hardware. AMD GPU users should look at alternatives with ROCm support.
What is the difference between EXL3 and GGUF quantization?
EXL3 is the quantization format used by ExLlamaV3, based on QTIP and optimized for NVIDIA GPU inference with this specific library. GGUF is the format used by llama.cpp and compatible tools, with broader hardware support including CPU-only and AMD GPU inference. The two formats are not directly interchangeable; conversion from HuggingFace model format to EXL3 uses the convert.py script in this repository.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/turboderp-org-exllamav3)