# UniSD: five self-distillation mechanisms, and an install order that carries the instructions

> The official implementation for a paper on a unified self-distillation framework for large language models. The framework is spelled out flag by flag, the environment is pinned to the patch level, and the one loose joint is a vLLM pin that sits outside the range trl supports.

**Ahren09/UniSD** — Official implementation for "Towards a Unified Self-Distillation Framework for Large Language Models" (https://arxiv.org/abs/2605.06597).

- Repository: https://github.com/Ahren09/UniSD
- Stars: 789 · Forks: 44
- Language: Python
- License: Apache-2.0
- Published: 2026-09-19 · Updated: 2026-09-19 · Language: en
- Canonical page: https://hysenlabs.com/projects/ahren09-unisd

## Five mechanisms, seven mode names, one integrated recipe

The framework is a table of components, the mode string that selects each one, and the flags it reads. Multi-teacher agreement comes in two forms: sequence level, with modes agreement_seq_random, agreement_seq_retrieval and agreement_seq_induction, and token level, with agreement_tok_random, agreement_tok_retrieval and agreement_tok_induction. Both read --num-auxiliary-contexts and --gamma_agreement, and the token level variant adds --agreement_stat.

The remaining three are single modes. EMA teacher stabilization is ema, driven by --ref_model_sync_steps and --ref_model_mixup_beta. Token level contrastive learning is contrastive, with --contrastive_weight and --contrastive_margin. Feature matching has two modes, match_joint and match_repr, sharing --final_layer_distill_weight. Divergence clipping, called JSD-Clip, is clip, with --alpha and --token_clip.

The integrated recipe is unisd_star, and the table says it combines EMA, matching, contrastive and agreement. Notably absent from that combination is divergence clipping, so the strongest configuration named in the table leaves one of the five mechanisms out.

## The install procedure is an order, not a requirements file

The environment targets Python 3.12 with CUDA 12.8 and cu128 wheels, and the README explains why a single pip command cannot do it: PyTorch's cu128 build lives on the PyTorch wheel index, and flash-attention must be compiled against the installed torch. The sequence is fixed.

```bash
# 1) Create and activate the env
conda create -n unisd python=3.12 -y
conda activate unisd
pip install -U pip setuptools wheel packaging ninja

# 2) Install cu128 PyTorch from the PyTorch wheel index (must precede flash-attn build)
pip install --index-url https://download.pytorch.org/whl/cu128 \
    torch==2.11.0 torchvision==0.26.0 torchaudio==2.11.0

# 3) Point flash-attn's CUDA build at a 12.x toolkit
#    (on many hosts /usr/local/cuda → 13.x, which mismatches torch's cu128 ABI)
export CUDA_HOME=/usr/local/cuda-12.6

# 4) Install everything else — flash-attn builds from source here
```

Step 4 has no command in the block. The missing piece appears as a comment at the top of requirements.txt, which says to apply the file last with pip install -r requirements.txt --no-build-isolation. That flag, which is what makes the source build of flash-attn possible, is documented only in the file it applies to. Note also that the target is CUDA 12.8 while the export points at cuda-12.6, and a note says any 12.x toolkit from 12.4 to 12.8 works.

## A pinned vLLM that sits outside the range trl supports

requirements.txt pins vllm==0.20.2 and trl==1.4.0 in the same environment. The README then states the problem plainly: trl 1.4.0 officially supports vLLM 0.12.0 through 0.18.0, so the pinned vLLM is outside the supported range.

The advice given is to expect a warning from trl at import time, and to pin vllm<0.19 if a runtime error surfaces from VLLMClient. The README says the combination works in the project's own smoke tests, and the direct training command does pass --use_vllm, so the fast inference path is part of the normal run rather than an optional extra.

That leaves an unresolved choice for a reader. The pin file installs a version trl does not claim to support, the documentation offers a lower pin as a remedy without saying which features break, and nothing explains what in the setup requires 0.20.2 rather than a version inside the range.

The rest of the pin set is broader than the visible feature set. deepspeed 0.19.0, peft 0.19.1, openai 2.36.0, deepspeed and peft among them, appear in requirements.txt with no flag or configuration file named for any of them anywhere in the README.

## The same setting has two names depending on the entry point

There are two launch paths, and they do not speak the same flag language.

The preset orchestrator is scripts/run_experiments.py, which handles GPU scheduling, dependency aware sweeps and defaults. Its documented example is python scripts/run_experiments.py contrastive --weight 0.1 --margin 0.5, and a --dry-run flag previews every job before launch. Note the names: weight and margin.

The direct path is python -m src.train.train_unisd, which exposes every UniSD flag. Its template passes --mode, --dataset, --model_name, --per_device_train_batch_size, --num-auxiliary-contexts and --use_vllm. In the component table the contrastive mechanism is configured by --contrastive_weight and --contrastive_margin.

So one setting has a short name in the orchestrator and a prefixed name in the trainer, and the README does not say the orchestrator maps one onto the other or whether they can be mixed. The orchestrator subcommand name, contrastive, matches the mode string, which is the one naming convention the two paths do share.

## The second worked example stops part way through its own flag

The direct command example is the shortest complete-looking run in the document, and it ends mid-flag:

```bash
# Example: token-level contrastive on MBPP with Qwen2.5-7B
python -m src.train.train_unisd \
    --mode contrastive --dataset mbpp \
    --model_name Qwen/Qwen2.5-7B-Instruct \
    --per_device_train_batch_size 4 \
    --contra
```

The template block above it is complete, with --num-auxiliary-contexts and --use_vllm in place, so the shape of a run is clear. The concrete example, which is the part a first-time user would copy, is the part that stops.

Two facts survive from it and they are the only concrete ones in the document: the dataset name mbpp and the model identifier Qwen/Qwen2.5-7B-Instruct. Everything else about scale, GPU count and expected duration is left to the reader, even though the orchestrator is explicitly built around GPU scheduling with a --gpus flag and the --dry-run preview exists precisely to answer those questions before a run starts.

## Six benchmarks and six models, none of them named

The abstract claims evaluation across six benchmarks and six models from three model families, and reports the integrated recipe improving over the base model by +5.4 and over the strongest baseline by +2.8. The highlights restate the grid as 6 benchmarks multiplied by 6 models multiplied by 3 model families.

None of those is enumerated. No benchmark name appears other than mbpp in the one truncated example command, no model family is named beyond the Qwen2.5-7B-Instruct identifier in that same line, and no metric is stated, so the +5.4 and +2.8 figures have no scale, no dataset split and no averaging method attached to them. The mechanism table is precise to the level of individual flag names, which makes the missing evaluation grid the largest gap in the document.

What is named outside the abstract is the project surface: a project page at unifiedsd.github.io, the arXiv entry 2605.06597, and the same identifier on the Hugging Face papers index, with ten authors across Georgia Tech, UCLA, Carnegie Mellon and William and Mary.

## A repository of six top-level entries and no tagged release

The tree is short: .gitignore, LICENSE, README.md, requirements.txt, scripts/ and src/. There is no tests directory, no examples directory, no checkpoints, no configs and no results directory. Apache-2.0 is the declared licence and the repository carries no homepage of its own, while the README points at a separate project page.

There are no GitHub releases, so there is no tagged version to pin. The default branch was last pushed on 2026-06-13, which puts the code roughly three and a half months behind the date this article is written on.

The environment is pinned far more tightly than the source tree is versioned. Every dependency in requirements.txt is fixed with an exact version, including local build identifiers such as torch==2.11.0+cu128 and flashinfer-python==0.6.8.post1, across four groups covering inference, the training stack and utilities. Verification of the install is a four line import of torch, vllm, flash_attn and flashinfer that prints their versions and whether CUDA is available, and the only optional variables named are WANDB_API_KEY for logging and HF_TOKEN for gated models.

## Conclusion

UniSD is a research repository with an unusually careful install procedure, and anyone running it should follow the order rather than a single requirements file, because torch has to be in place before flash-attn compiles. Before you plan a run, decide whether the vLLM warning matters for your hardware, check which benchmark and model the paper's figures came from since the README names none, and expect to work from the default branch dated 2026-06-13 rather than a tagged release.

## FAQ

### What is UniSD and what does it add to self-distillation?

UniSD is the official implementation for a paper on a unified self-distillation framework for large language models. It studies five mechanisms, multi-teacher agreement at sequence and token level, EMA teacher stabilization, token-level contrastive learning, feature matching and divergence clipping, then combines four of them into an integrated recipe named unisd_star.

### What does UniSD require before pip install works?

Python 3.12 with CUDA 12.8 and cu128 wheels. Torch 2.11.0 has to be installed from the PyTorch wheel index first, flash-attn 2.8.3 is compiled against that torch with CUDA_HOME pointing at a 12.x toolkit, and requirements.txt is applied last with --no-build-isolation, a flag documented only in that file's header comment.

### Why does UniSD warn about trl and vLLM?

requirements.txt pins trl==1.4.0 together with vllm==0.20.2, while trl 1.4.0 officially supports vLLM 0.12.0 to 0.18.0. The README says to expect a warning from trl at import time and to pin vllm<0.19 if a runtime error appears from VLLMClient.

### How do I launch a UniSD training run?

Two ways. scripts/run_experiments.py is a preset orchestrator that handles GPU scheduling and dependency-aware sweeps, accepts a --gpus list and has a --dry-run flag to preview jobs. python -m src.train.train_unisd exposes the flags directly, including --mode, --dataset, --model_name, --per_device_train_batch_size, --num-auxiliary-contexts and --use_vllm.

### Does UniSD publish tagged releases?

No. The repository has no GitHub releases, so there is no tagged version to pin, and the default branch was last pushed on 2026-06-13. The README links a separate project page at unifiedsd.github.io, the arXiv entry 2605.06597 and the same identifier on the Hugging Face papers index.

## Sources

- [Ahren09/UniSD on GitHub](https://github.com/Ahren09/UniSD)
- [Issues](https://github.com/Ahren09/UniSD/issues)
- [License: Apache-2.0](https://github.com/Ahren09/UniSD/blob/main/LICENSE)
- [README](https://github.com/Ahren09/UniSD/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/ahren09-unisd
