AngelSlim: Tencent's Compression Toolkit for LLMs, VLMs and Diffusion Models
Model compression toolkit engineered for enhanced usability, comprehensiveness, and efficiency.
At a glance
- What is it?
- AngelSlim bundles post-training quantization, quantization-aware distillation, speculative decoding and sparse attention into one Python toolkit. It is aimed at engineers who already have a model checkpoint and need a smaller or faster one, not at people looking for a one-command converter.
- Who is it for?
- Adopt AngelSlim if you are compressing a model family it already lists in configs/ or scripts/ptq, such as Qwen3, DeepSeek, Hunyuan or GLM, and you have the GPU memory to run the calibration and training stages yourself.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 5 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What AngelSlim compresses, and who ends up using it
AngelSlim is a Python model compression toolkit from Tencent. The README describes it as "a more accessible, comprehensive, and efficient toolkit for large model compression", and the release notes back that up with breadth: quantization, distillation, speculative decoding, sparse attention and reasoning early exit all live in the same repository. The topics list names audio, diffusion, LLM and VLM targets, so the intended scope is not text-only language models.
The practical audience is an engineer holding a checkpoint that is too large or too slow for the hardware they have. The news entries give a concrete shape to that problem: a Hy4 preview build compressed from 1.5TB to 214 GiB with a reported 0.7% drop on SWE-bench Pro, later run on a laptop RTX 4090 plus a four-GPU A4000 server at 1.02 tokens/s. That is not a serving story for a production cluster. It is a story about making a model fit somewhere it previously could not run at all.
The second audience is training-side. AngelSpec, released in July 2026, is a torch-native disaggregated speculative-decoding training framework with draft methods led by DFly and MTP + TTT, and the README reports up to 2.40x average end-to-end speedup over autoregressive decoding for DFly on Hy3-A21B. If your job is to train a drafter rather than quantize a target model, that is the entry point.
Four mechanisms in one repository, and how they compose
Quantization is the largest surface. The README lists FP8-Static, W4A8-FP8, NVFP4, INT4, 2-bit and 1.25-bit paths, with named algorithms including SmoothQuant, DAQ, TEQUILA for ternary quantization and Sherry for 1.25-bit. Configs sit under configs/ (the README points to configs/flux and configs/seed_oss as examples) and post-training quantization scripts under scripts/ptq.
Distillation is the second. Since May 2026 the project supports distillation for full-precision HuggingFace models and for quantized QAT-style models. In July 2026 it added scale-only quantization-aware distillation on Megatron-Core for Qwen3-MoE and Hy3, with TP/EP/CP/SP distributed training. This is the part that most distinguishes AngelSlim from a plain PTQ script collection: the compression pipeline can continue training the model after quantization rather than freezing it.
Speculative decoding is the third. Eagle3 training and deployment covers LLMs, VLMs and audio models as of v0.3.0. DFlare is described as a block-diffusion speculative decoding framework with layer-wise fusion, and D-Cut is an adaptive verification depth pruning technique. SpecExit takes a different angle entirely: it is a reasoning early-exit algorithm, with a vLLM pull request referenced in the release notes.
The fourth is sparse attention. Stem dynamically selects top-k key blocks for block-sparse attention, targeting the prefill stage of long-context models.
The composition matters more than any single piece. A realistic pipeline is quantize, then distill to recover accuracy, then attach a drafter, then serve. Each stage has its own documentation page, and the README does not describe a single orchestrated command that runs all of them.
Installing AngelSlim and running a first quantization
The repository ships setup.py and a requirements/ directory, and the README links to the documentation site and a technical report rather than giving a pip line. That is worth flagging: there is no install command in the README itself, so the setup steps below are the ones the repository layout implies, not a quoted procedure. Confirm them against angelslim.readthedocs.io before relying on them.
A source install from a git checkout is the path setup.py is written for. The version logic reads git and appends a local segment describing the CUDA and torch build, and it falls back to the version recorded in PKG-INFO when building from an sdist:
git clone https://github.com/Tencent/AngelSlim.git
cd AngelSlim
pip install -e .Because the version string depends on the torch and CUDA versions present at install time, the package you get is tied to that environment. Installing into a different torch build later means reinstalling rather than moving the environment.
Quantization runs are driven from configs and scripts. The README points at scripts/ptq for FP8-Static and SmoothQuant on Hy3, and at configs/ for per-model settings such as configs/flux and configs/seed_oss. A first run therefore means picking an existing config for a supported model, adjusting the model path and calibration data, and invoking the script for that algorithm:
ls configs/
ls scripts/ptqWhat you should see is a directory of per-model config folders and a set of PTQ entry scripts. If your model is not represented in either place, you are in unsupported territory and should expect to write the integration yourself rather than tune an existing config. The dataset/ and evaluation/ directories exist for calibration data and for measuring the damage afterwards.
Where AngelSlim stops being the right tool
The clearest limitation is support surface. Model coverage is enumerated in release notes and configs rather than derived from a general interface. Qwen3, Qwen3-VL, Qwen3-Omni, GLM-4.6, DeepSeek-R1/V3, Kimi-K2, Hunyuan 0.5B through 7B, Seed-OSS, FLUX, Hunyuan-MT-7B and Hy3 appear by name. A model outside that list is not a configuration change away from working.
The second limitation is that AngelSlim does not serve models. The 214 GiB Hy4 preview was run through Prima.cpp, a separate project, and the vLLM work referenced in the notes is a pull request to vLLM, not code in this repository. Quantized output has to load in whatever runtime you already use, and the README does not promise compatibility with any particular one.
The third is operational. Quantization-aware distillation on Megatron-Core needs distributed training with TP/EP/CP/SP parallelism. That is a cluster job, not a laptop job. The hardware note in the Hy4 announcement, a laptop plus a four-GPU server, describes the deployment target after compression, not the resources needed to produce the compressed model.
Finally, the licence. The repository metadata reports NOASSERTION even though setup.py carries an Apache License 2.0 header. Those two facts disagree, and the LICENSE file at the repository root is the thing to read before you depend on the output.
How it differs from llama.cpp and from SpecForge
The closest comparison depends on which half of AngelSlim you need.
For quantization, llama.cpp occupies adjacent ground but approaches it from the inference side. llama.cpp defines its own GGUF container format and its own quantization types, and the runtime is the product. AngelSlim produces weights for other runtimes and treats the runtime as someone else's problem. The STQ1_0 kernel for 1.25-bit models shows the boundary clearly: AngelSlim released the kernel and opened a pull request to llama.cpp rather than shipping an inference engine. If you want a single binary that quantizes and runs a model, llama.cpp is the shorter path. If you want to apply research algorithms such as DAQ, TEQUILA or Sherry to a specific checkpoint and then hand the result to your existing stack, AngelSlim is aimed at that.
For speculative decoding, SpecForge is the natural comparison, and it appears in the related searches for this project. The difference the README makes visible is scope and coupling. AngelSlim's speculative-decoding work is one feature among quantization, distillation and sparse attention, and its drafter training is tied to the AngelSpec framework with DFly and MTP + TTT. A team that only wants to train an Eagle3 drafter may find the surrounding toolkit heavier than necessary; a team that wants to quantize a model and then attach a drafter gets both from one dependency tree.
Maintenance, release cadence and upgrade cost
The last push to the default branch was on 2026-09-04, and the repository is not archived. Releases are not frequent in the version-number sense: v0.2.0 in November 2025, v0.3.0 in January 2026, v0.5.0 in June 2026. Between those, the news list shows a steady stream of model and algorithm additions, several per quarter through 2026.
That cadence is the upgrade cost. The project grows by adding support for named models and named algorithms, and each addition can bring its own config directory, its own script and its own documentation page. Upgrading means re-checking whether the config you depend on still matches the current script interface, and whether the algorithm you use has been superseded. The version scheme compounds this: because setup.py derives the version from git and the local torch and CUDA build, two environments installed from the same commit can report different package versions.
On licensing, setup.py's header states Apache License 2.0, while the repository metadata reports NOASSERTION. The Hy-MT1.5-1.8B and Qwen3 quantized weights are published on Hugging Face and ModelScope under the AngelSlim organization, and model weights typically carry their own terms separate from the toolkit. Read the LICENSE file and the model card for the specific checkpoint you intend to ship. This is a description of what the repository states, not legal advice.
Editorial conclusion
Adopt AngelSlim if you are compressing a model family it already lists in configs/ or scripts/ptq, such as Qwen3, DeepSeek, Hunyuan or GLM, and you have the GPU memory to run the calibration and training stages yourself. Do not adopt it as a generic drop-in quantizer for an arbitrary checkpoint, and do not expect the README to walk you through a first run: the docs at angelslim.readthedocs.io and the per-feature pages under docs/source/features/ are where the actual procedures live, and the README does not document rollback or how to undo a quantization run. Before committing, check three things against your own setup: that your target model appears in configs/ or the PTQ scripts, that the licence terms are acceptable, since the repository LICENSE is not a recognized SPDX identifier, and that the quantized weights load in your serving stack, because AngelSlim produces checkpoints and does not ship an inference server of its own.
Frequently asked questions
What is AngelSlim used for?
It is a Python toolkit for compressing large models. The README covers quantization (FP8, INT4, NVFP4, 2-bit and 1.25-bit among others), distillation, speculative decoding with Eagle3 and DFlare, and sparse attention with Stem.
Does AngelSlim support Qwen3 and Eagle3 together?
The related searches for the project include Qwen3 Eagle3 variants, and the v0.3.0 release notes state that training and deployment of Eagle3 for all-scale LLMs, VLMs and audio models is supported, with guidance documentation on the AngelSlim docs site. The README does not list a specific prebuilt Qwen3 Eagle3 configuration by name.
How do I install AngelSlim?
The README does not give a pip command. The repository contains setup.py and a requirements/ directory, and setup.py derives its version from git plus the local torch and CUDA build, so a source install from a git checkout is the path it is written for. Check angelslim.readthedocs.io for the current instructions.
What is the 1.25-bit quantization in AngelSlim?
The README describes Sherry as a hardware-efficient 1.25-bit quantization algorithm and references an STQ1_0 kernel for 1.25-bit models, along with a pull request to llama.cpp. A 1.25-bit Hy-MT1.5-1.8B translation model is published on Hugging Face.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/tencent-angelslim)