# GPTQModel ships AWQ, ParoQuant, EXL3 and GGUF next to GPTQ, and pins numpy below Python 3.14

> ModelCloud/GPTQModel is a quantization toolkit with native kernels for several accelerator families, whose own comparison table gives Transformers, vLLM, and SGLang identical support rows while the paragraph beneath it limits SGLang to five format combinations. Its dependency file has one exact pin, in one interpreter branch, and its two build files disagree about the copyright year.

**ModelCloud/GPTQModel** — LLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM, and SGLang.

- Repository: https://github.com/ModelCloud/GPTQModel
- Website: https://x.com/Qubitium
- Stars: 1,269 · Forks: 213
- Language: Python
- License: NOASSERTION
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/modelcloud-gptqmodel

## The package name is GPTQModel and four of its twelve methods are not GPTQ

The capability table lists twelve quantization methods. GPTQ is one of them.

The others are AWQ, ParoQuant, GGUF, FP8, Exllama V3, EoRA, Group Aware Act Reordering, QQQ, Rotation, GPTAQ, and FOEM. Four of those are called out in the text as currently native GPT-QModel paths rather than integrations: GGUF, FP8, EXL3, and ParoQuant.

AWQ is the notable one, because it is a different quantization method with its own paper and its own ecosystem, and here it is a first-class method with its own formats rather than a compatibility shim. The keywords in the packaging metadata list both autogptq and autoawq alongside the current names, which is the project telling you which two tools it grew out of.

The architecture claim that goes with this is modularity across four layers, method, format, backend, and kernel, with method-specific controls preserved where they matter and a shared model lifecycle underneath. New methods are meant to plug into the same calibrate, quantize, validate, save, load, and serve path.

So the name understates the scope. Anyone arriving looking for a GPTQ implementation will find a framework whose fastest-growing surface is the list of model families it recognises.

## The table gives SGLang the same row as vLLM, and the next paragraph limits it

The comparison table has five columns: GPT-QModel, Transformers, vLLM, SGLang, and a fifth labelled Lora Training.

Reading the rows, the Transformers, vLLM, and SGLang columns are identical in all twelve rows. Each supports GPTQ, AWQ, EoRA, Group Aware Act Reordering, GPTAQ, and FOEM, and each is marked as not supporting ParoQuant, GGUF, FP8, EXL3, QQQ, or Rotation.

Three different runtimes presented with an identical support profile is either true at method granularity or suspicious. The paragraph immediately below suggests the first, because it then restricts SGLang at format granularity: loading is limited to the GPTQ method with the GPTQ, GPTQ v2, and Marlin formats, and the AWQ method with the GEMM and Marlin formats.

So a cell reading yes in the table can mean three things depending on the row: the method is supported, the specific format you need is supported, or the combination is. The table cannot tell them apart, and the table is what a reader scans first.

The Lora Training column is a different axis again, since it is not a runtime. It marks ParoQuant as supported while marking EoRA as not, which is the reverse of what the runtime columns say for those two rows.

## Five engine argument names are translated on the way to SGLang

The SGLang integration is not a straight pass-through, and the translation table is the most concrete thing in the documentation.

Five engine arguments are renamed. The tensor parallelism size becomes the parallelism size, the GPU memory utilisation fraction becomes a static memory fraction, the maximum model length becomes the context length, the seed becomes a random seed, and the eager enforcement flag becomes a flag that disables CUDA graphs.

Then two dtype behaviours. An explicit legal dtype is preserved as given, and the deprecated torch_dtype spelling is normalised into SGLang's own string dtype names.

Finally a caveat about two AWQ formats. The GEMV fast and LLM AWQ formats still require torch.float16, and the sentence adds that those formats are not in the list SGLang supports anyway.

That last detail is the shape of the whole integration: a compatibility layer that normalises names and types so the same call site works across engines, plus a list of formats that remain outside the supported set. The method and format matrix that follows is the table that would tell you whether your specific combination works, and it is cut off in the visible document after the first row.

## numpy is the only exact pin in the file, and it applies to one interpreter branch

The dependency file has twenty-three lines and one of them uses an exact version. Everything else is a floor.

The exact pin is numpy at 2.2.6, and it carries an environment marker restricting it to Python versions below 3.14. Above that marker, the same package is asked for at 2.3.0 or newer, with no ceiling.

That is a deliberate split. Below the newest interpreter the resolver gets one known-good array version; above it, where the project cannot have tested as much, it gets a floor and whatever compatibility fixes have shipped.

The rest of the file follows the floor pattern, including the two that matter most for a quantization toolkit. torch is asked for at 2.8.0 or newer and transformers at 5.14.0 or newer. Neither is pinned, and neither has a ceiling.

For a project that compiles device kernels, an unpinned torch is the consequential one. A torch major release can change an API the extension layer calls, and nothing in the file prevents pip from taking it.

## transformers 5 and protobuf 7 are floors on a toolkit meant to be drop-in

Two dependency floors put this project at the leading edge of its ecosystem.

transformers is required at 5.14.0 or newer. protobuf is required at 7.34.0 or newer. Both are major versions well ahead of what most environments carry, and both are minimums rather than preferences, so an environment on an older major cannot install the toolkit without upgrading a large dependency of its own.

The rest of the file makes the same demand. datasets at 3.6.0 or newer, pyarrow at 21 or newer, safetensors at 0.7.0 or newer, accelerate at 1.13.0 or newer, pillow at 11.3.0 or newer, torchao at 0.18.0 or newer.

Then there is the toolchain side. ninja and maturin are both in the runtime dependency list, and maturin is the Rust build tool, which means the install can compile native code. That lines up with the packaging metadata declaring a C++ language classifier alongside the Python ones, and with a native extension tree in the repository that is collected as a namespace package rather than a regular one.

A handful of the remaining names are hard to place from the file alone: a device information library, a tokenizer helper, a logging library, a memory usage library, and a regular expression helper, plus a serialisation library. None of them is described in the visible documentation, so whether each is a thin wrapper or a load-bearing piece is something you would only learn by reading the imports.

## setup.py and pyproject.toml disagree about the copyright year

There are two build files and both carry SPDX headers. The headers do not match.

The packaging manifest's header covers 2024 to 2026 for both the company and the individual author. The legacy setup script's header covers 2024 to 2025 for the same two parties, with everything else about the attribution identical, including the licence identifier and the contact address.

A year of drift in a copyright header is not a bug in the sense that it breaks an install. It is a symptom of the two files having different maintenance histories, and this project has a reason for having both.

The setup script does three things and delegates the rest. It reads the version by executing gptqmodel/version.py and pulling the dunder version attribute out of the resulting namespace. It collects the regular packages while excluding tests. And it appends any namespace packages whose names start with the extension prefix, so the native extension tree is included even though it is not a regular package.

Everything else, the name, the metadata, the dependencies, and the extras, comes from the manifest. The script exists to compute a version and to make the extension packages visible to setuptools.

The version being read by executing a file is worth keeping in mind next time a version string looks wrong, and it connects to a fix listed in a recent release, where startup version reporting is described as more accurate.

## The GGUF path is an internal shim, and no external gguf package is needed

One section of the documentation is devoted to loading Prism and Bonsai GGUF checkpoints, and its whole point is what is not required.

The text says those checkpoints load through the native GGUF loading path and an internal GGUF runtime shim, and that no external gguf package from the package index is needed. A recent release also records updated native GGUF support.

The example is short and shows two things beyond the load call itself. There is a profile argument, with a low memory profile named alongside the default, which suggests memory behaviour is selectable at load time rather than automatic. And generation returns a sequence of tokens that you decode through the model's own tokenizer, so the object returned by load is a working model rather than a checkpoint handle.

```py
from gptqmodel import GPTQModel

model = GPTQModel.load("prism-ml/Bonsai-1.7B-gguf")
# or: model = GPTQModel.load("prism-ml/Bonsai-1.7B-gguf", profile="low_memory")

tokens = model.generate(
    "Who wrote Romeo and Juliet?",
    max_new_tokens=128,
)[0]

print(model.tokenizer.decode(tokens, skip_special_tokens=True))
```

Shipping your own reader for a third-party container format is a real maintenance cost, and it is also what removes a dependency. For an ecosystem where the reference GGUF library has moved slowly, that trade can be the sensible one.

## The build requirement caps setuptools, and the quality extra pins one tool exactly

The build system declaration is narrow on purpose. It requires setuptools at 77.0.1 or newer and below 83, with nothing else in the build requirements, and uses the standard setuptools backend.

An upper bound on a build tool is unusual. Most projects floor it. A ceiling says the project has decided something in a later setuptools does not work, which is worth knowing before you bump it in your own build pipeline and wonder why the install starts failing.

The extras tell a similar story. The test extra takes pytest at a floor of 8.3.5, pytest-timeout at 2.3.1, and the parameterised test library unpinned. The quality extra pins one tool to an exact version, and has a second tool commented out in place. So there is one linter, pinned exactly, and no import sorting configured.

That is a coherent minimal setup for a project whose heavy dependencies are already unpinned floors. The linter is pinned so that formatting and lint results are stable between contributors, while the runtime libraries are left to float.

One more thing the root of the repository carries: a licenses directory, alongside a CREDITS file and a checkpoint document. A licences tree in a repository like this is the third-party notice set for the kernels and code borrowed or vendored, which is a fair amount of attribution to be carrying by hand.

## Conclusion

GPTQModel fits a team with a model in a supported architecture and a target accelerator they cannot choose, who wants one tool covering several quantization methods and can accept a moving dependency floor. Four things to check first. Whether your method and format combination is actually supported, since the capability table is method-level while the runtime restrictions are format-level. Which accelerator backend you land on, because kernels are selected per device and the release notes describe that selection changing. Whether the transformers and torch floors suit your environment, since both are minimums rather than pins. And what the version string reports, because version reporting was listed as fixed in a recent release and the version comes from executing a Python file at build time.

## FAQ

### how to install gptqmodel

It is published to PyPI as GPTQModel, requires Python 3.10 or newer, and takes its dependency list from requirements.txt. Native extensions are built with ninja and maturin, both of which are runtime dependencies, and the setup script collects the extension tree as namespace packages.

### gptqmodel vs auto gptq

GPTQModel is the successor tool with a broader scope. The packaging keywords list both autogptq and autoawq, and the project supports twelve quantization methods rather than GPTQ alone, including AWQ, ParoQuant, GGUF, FP8, EXL3, QQQ, Rotation, EoRA, and FOEM.

### What quantization methods does GPTQModel support?

Twelve are listed: GPTQ, AWQ, ParoQuant, GGUF, FP8, Exllama V3, EoRA, Group Aware Act Reordering, QQQ, Rotation, GPTAQ, and FOEM. GGUF, FP8, EXL3, and ParoQuant are described as currently native quantization and runtime paths rather than integrations.

### Which hardware does GPTQModel accelerate on?

The description names NVIDIA, AMD, Intel GPUs and Intel, AMD, and Apple CPUs. The documentation banner lists NVIDIA CUDA, AMD ROCm, Huawei Ascend, Intel XPU, and CPU, with Transformers, vLLM, and SGLang as the runtime integrations. Native MPS quantization was added in the 7.3.4 release.

### Does GPTQModel work with vLLM and SGLang?

It integrates with both. For SGLang, loading is limited to the GPTQ method with the GPTQ, GPTQ v2, and Marlin formats, and the AWQ method with the GEMM and Marlin formats, and five vLLM-style engine argument names are translated on the way through.

### How is the GPTQModel version determined?

The setup script reads it by executing gptqmodel/version.py and pulling the dunder version attribute from the result, so the version lives in a Python file at build time rather than in the packaging manifest, which declares the version as dynamic.

## Sources

- [Issues](https://github.com/ModelCloud/GPTQModel/issues)
- [ModelCloud/GPTQModel on GitHub](https://github.com/ModelCloud/GPTQModel)
- [Project website](https://x.com/Qubitium)
- [README](https://github.com/ModelCloud/GPTQModel/blob/main/README.md)
- [Releases](https://github.com/ModelCloud/GPTQModel/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/modelcloud-gptqmodel
