DistillKit: offline LLM distillation with compressed teacher logits
An Open Source Toolkit For LLM Distillation
At a glance
- What is it?
- Arcee's Apache-2.0 toolkit trains students either against a live teacher or against pre-captured, bit-packed logit datasets. The compression knobs are the interesting part, and they are also where the sharp edges are.
- Who is it for?
- Adopt DistillKit if you are distilling a student from a teacher whose logits you want to capture once and reuse, or if you want to train against Arcee's published packed logit datasets without building a capture pipeline. Do not adopt it if you need a stable, versioned API: the package is at 1.0.0, there are no retrieved releases, and the CLI is a single config-file entry point.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 141 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What DistillKit solves, and for whom
Standard supervised fine-tuning gives a student one signal per token: the correct token. Knowledge distillation gives it the teacher's whole distribution over the vocabulary, which is a denser target. The README states the project's own framing directly: the student learns from the teacher's probability distribution, which it calls "a much richer learning signal."
The problem DistillKit actually attacks is not the distillation loss. It is the storage bill. The README says storing top-k token-logit pairs becomes prohibitively expensive when distilling on billions of tokens. If a teacher's full distribution over a 151,936-token vocabulary were kept in float32, every single token position would cost roughly 600 KB. Multiply that by a corpus and the workflow stops being practical.
So the audience is narrow and specific. It is teams that want to distil at corpus scale, want to reuse one teacher capture across several students, or are VRAM-limited and cannot hold teacher and student at once. The README also names the models this infrastructure has been used for: Virtuoso, SuperNova Medius, and Blitz. That is the strongest claim in the repository, and it is a claim about Arcee's own releases rather than about general users.
Online versus offline: the decision the README makes for you
DistillKit ships two workflows, and the README gives a rule of thumb rather than a feature matrix. Online distillation runs the teacher in real time during student training, with no storage overhead. Offline distillation reads pre-captured teacher outputs, which allows training multiple students from the same teacher and works when VRAM is tight.
The stated rule: if you can fit both teacher and student with dense distributions into VRAM, use online. Otherwise use offline with the compression system. That is a memory-bound decision, not a quality decision, and it is worth reading as such. Online avoids the compression pipeline entirely, so it also avoids the reconstruction step and the config-matching constraint that offline imposes.
The offline path is where the engineering lives. It requires a capture step, a compression configuration, and a training configuration whose compressor block matches the capture. The README is blunt about the failure mode: "the configuration that was used to capture the logits must be reflected in the distillation configuration. Mixing and matching isn't gonna work out so hot." There is no documented validation that catches a mismatch for you.
How logit compression actually works
The pipeline has five documented stages. Select top-k logits from the teacher output. Sort by log-probability, optionally applying delta encoding. Fit a polynomial to the distribution curve. Quantize the residuals, with optional error diffusion. Bitpack everything into byte vectors.
The knobs map onto those stages. k controls how many top logits are kept. exact_k and exact_dtype control how many of the leading values are stored losslessly and in what precision. polynomial_terms is a list that can hold integers and the string "sqrt", which is how the curve fit is parameterised. residual_bins and error_diffusion control the quantisation of what the polynomial does not capture. delta_encoding applies to the sorted values.
The README gives two concrete configurations. The recommended one, described as what Arcee uses for new captures, sets k to 128, exact_k to 16, exact_dtype to bfloat16, polynomial_terms to [0, 1, 2, 3, 4, "sqrt"], and leaves residual_bins empty and both delta_encoding and error_diffusion false. The README states this takes about 300 bytes per token, or 0.15% of the uncompressed distribution size. The budget pick uses k of 50, exact_k of 1, polynomial_terms of [0, 1, "sqrt"], and comes in around 114 bytes per token. The README claims that is smaller and reconstructs better than storing the top 32 logprobs in bfloat16.
Those two figures are the only size numbers the README publishes. Treat them as the project's own measurements, not as independent verification. What is clear from the design is that the tradeoff is explicit and user-owned: the README says there are "lots of knobs you can twiddle here."
Installing DistillKit and running a first offline job
The README's installation path is a clone plus an editable install. The package requires Python 3.10 or newer. Run this from the parent directory where you want the checkout to live:
git clone https://github.com/arcee-ai/distillkit.git
cd distillkit
pip install -e .The core dependencies pulled in are TRL pinned at 0.25.1, Accelerate at 1.11.0, Transformers at 4.50.3 or newer, Datasets, Pydantic 2.12.4, Click 8.3.0, PyTorch 2.0.0 or newer, wandb and tqdm. If you intend to capture your own teacher outputs rather than use a published dataset, the README points to an extra:
pip install -e ".[capture]"That extra adds vLLM and pyarrow. The README recommends most users start with the pre-captured datasets instead, so treat the capture extra as the second step, not the first.
The entry point is a single CLI that takes a config file. The README's quick start is a YAML file named config.yaml, with a dataset block pointing at the packed logit dataset arcee-ai/Qwen3-235B-Logits-Packed-8192, a teacher block whose kind is dataset, a compressor block, a list of loss functions, and training arguments. The compressor block in that example sets d to 151936 for the vocabulary size, k to 128, exact_k to 32, polynomial_terms to [0, 1, 2], delta_encoding true, error_diffusion false, and exact_dtype float32. The loss list pairs cross_entropy at weight 0.5 with kl at weight 0.5. Then:
distillkit config.yamlThat command is the console script declared in pyproject.toml, mapping to distillkit.main:main. Expect training to start against the dataset, with progress reported through tqdm and metrics through wandb. The README does not document what the command prints on a bad config, so check the compressor block against your dataset before launching rather than after.
Loss functions, sparse distributions, and the chunking detail
Losses are composable with independent weights, which is the part of the config most users will tune first. The README lists distribution-based losses kl, jsd and tvd, plus ranking losses including hinge, and hidden state alignment. The quick-start example combines cross_entropy and kl, which is the conventional arrangement: hard labels keep the student anchored while the teacher distribution shapes it.
The kl entry in the example carries three extra keys worth noting. temperature is 1.0. missing_probability_handling is set to zero, which is how the loss treats tokens outside the stored top-k. sparse_chunk_length is 1024, and the README mentions automatic chunking for memory-efficient processing of long sequences. That last key is the practical lever when a long sequence length causes memory trouble.
Sparse versus dense is a separate axis from online versus offline. Dense distributions cover the full vocabulary and are more accurate but memory-intensive. Sparse distributions keep only the top-k and the README calls them a lossy but useful approximation, adding that with sufficient training data sparse distillation can reach equivalent performance to dense. That is a claim about training data volume, and it implies the sparse path is a bet on corpus size rather than a free saving.
Where DistillKit is the wrong tool
The clearest limitation is the coupling between capture and training configuration. Nothing in the README describes a schema check, a version stamp on captured datasets, or a migration path if you change compressor settings after capturing. If you capture with one configuration and train with another, the README says it will not work out well, and the failure is silent enough that you may only notice it in the resulting model quality.
The second limitation is the dependency surface. TRL is pinned at 0.25.1 and Accelerate at 1.11.0 with compatible-release specifiers, which means minor updates within those series but not major ones. Transformers is a floor, not a pin. In practice that means a DistillKit install constrains the rest of your training environment, and upgrading TRL for an unrelated reason risks breaking the toolkit. There is no retrieved release history to check against.
The third is scale. The quick-start example uses per_device_train_batch_size of 1 with gradient_accumulation_steps of 8 and gradient_checkpointing enabled, at a sequence length of 8192. Those settings are what a memory-constrained setup looks like. If your goal is a small experiment with a modest teacher, the compression pipeline, the capture step and the config-matching constraint are overhead you do not need; plain supervised fine-tuning or online distillation alone would be simpler.
Finally, offline distillation is the wrong choice if you have only one student and enough VRAM for both models. The README's own rule of thumb routes you to online in that case.
The alternative: plain TRL fine-tuning, and when it wins
DistillKit is built on Transformers, TRL and Accelerate, so the honest alternative is the underlying stack used directly. A TRL trainer with a standard causal language modelling objective gives you hard-label fine-tuning with no teacher, no capture step, and no compressor configuration. The difference in approach is the training signal: one correct token per position instead of a distribution over the vocabulary.
That difference matters most when you have no teacher logits and no budget to produce them. It matters less when you do, because the richer signal is the whole point of distillation and it is what DistillKit exists to make affordable. The choice is therefore not really DistillKit versus TRL. It is whether the distillation signal is worth a capture pipeline and a compression configuration.
If it is, and you want to avoid the capture work, the README points at pre-captured datasets such as arcee-ai/Qwen3-235B-Logits-Packed-8192. Using those means accepting Arcee's compressor settings as given, which sidesteps the mismatch risk entirely.
Licence, maintenance and upgrade cost
DistillKit is Apache-2.0 and the repository is not archived. The last push was on 2026-05-12, which is more than four months before today, so this is not a project to describe as actively developed on the strength of that date alone. There are no retrieved releases, so there is no changelog to read before upgrading and no published compatibility matrix beyond the version constraints in pyproject.toml. The package version itself is 1.0.0.
Apache-2.0 is permissive and includes an explicit patent grant, which matters for a component that may end up inside a commercial training pipeline. It does not settle the licensing status of the teacher model whose logits you capture, and the README says nothing about that. That question belongs to whoever owns the teacher weights, not to DistillKit's licence file.
The upgrade cost is concentrated in the pinned dependencies. TRL at 0.25.1 and Accelerate at 1.11.0 are the two that will fight you if your environment has moved on. Since there are no releases, the practical upgrade path is pulling main and re-reading the config schema, which means any compressor key rename is something you discover by running the CLI rather than by reading a migration note.
Editorial conclusion
Adopt DistillKit if you are distilling a student from a teacher whose logits you want to capture once and reuse, or if you want to train against Arcee's published packed logit datasets without building a capture pipeline. Do not adopt it if you need a stable, versioned API: the package is at 1.0.0, there are no retrieved releases, and the CLI is a single config-file entry point. Verify first that the compressor block in your config matches the one used at capture time, that your GPU memory fits the batch size and sequence length you intend, and that your target student architecture is supported by the pinned Transformers and TRL versions.
Frequently asked questions
How do I install DistillKit?
Clone the repository, change into it, and run pip install -e . with Python 3.10 or newer. If you want to capture your own teacher outputs, install the capture extra with pip install -e ".[capture]", which adds vLLM and pyarrow.
How do I distill an LLM model with DistillKit?
Write a YAML config with a dataset block, a teacher block, a logprob_compressor block, loss_functions, and training_args, then run the distillkit CLI against that file. The README's quick start uses a pre-captured dataset and a cross_entropy plus kl loss pair.
What is the teacher-student distillation model in DistillKit?
A teacher model produces a probability distribution over the vocabulary, and the student trains against that distribution rather than only against the correct token. DistillKit supports running the teacher live during training or reading pre-captured teacher outputs from a dataset.
Why must the compressor configuration match between capture and training in DistillKit?
The README states that the configuration used to capture the logits must be reflected in the distillation configuration, and that mixing and matching will not work out well. The compressor block controls how the stored logits are reconstructed, so a mismatch changes what the loss is actually computed against.
Does DistillKit support online distillation as well as offline?
Yes. Online distillation runs teacher inference in real time during student training with no storage overhead, while offline distillation reads pre-captured outputs. The README's rule of thumb is to use online if both models fit in VRAM with dense distributions, and offline otherwise.
What is DistillKit's licence?
The repository is licensed under Apache-2.0. That covers the toolkit itself, not the licence terms of any teacher model whose logits you capture, which the README does not address.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/arcee-ai-distillkit)