Model or dataset
yoshitomo-matsubara/torchdistill avatar
yoshitomo-matsubara/torchdistill

A torchdistill config is a list of import paths, and the pip install drops the configs directory

A coding-free framework built on PyTorch for reproducible deep learning studies. PyTorch Ecosystem. 🏆26 knowledge distillation methods presented at TPAMI, CVPR, ICLR, ECCV, NeurIPS, ICCV, AAAI, etc are implemented so far. 🎁 Trained models, training logs and configurations are available for ensuring the reproducibiliy and benchmark.

1,633 stars143 forksPythonMIT

At a glance

What is it?
yoshitomo-matsubara/torchdistill is an MIT licensed PyTorch framework that moves a deep learning experiment into a declarative YAML file, with 26 distillation methods implemented and trained models published for benchmarking. Two packaging facts shape what you actually get: the wheel excludes the configs and notebooks the documentation points at, and the runtime dependency list carries Cython.
Who is it for?
torchdistill suits a researcher who wants an experiment to be a reviewable file rather than a script, and who will clone the repository rather than install from a package index, because the example configs and the five demo notebooks are excluded from the built distribution and they are the documentation. Verify four things.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 13 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The wheel excludes configs, examples, demo, docs, and tests

The build configuration ends with an exclusion list:

code
[tool.setuptools.packages.find]
exclude = ["tests*", "examples*", "demo*", "docs*", "configs*"]

Five top level directories, all of which exist in the repository, are kept out of the built package. Two of them are the ones the documentation sends you to. The README tells you to look at configs/ to see which modules are abstracted and how they are defined, and points at configs/sample/ for training without a teacher. The five notebooks under demo/ are the worked route in, covering CIFAR knowledge distillation, plain CIFAR training, intermediate representation extraction, and two GLUE notebooks for fine tuning and for distillation with submission. The release archive is a different package again, with a MANIFEST.in beside the pyproject.toml and a setup.cfg still in the tree. So an index install gives you the library, and everything that explains how to use it stays in the repository.

A config entry is an import instruction rather than a schema

The YAML extension is a custom tag called !import_call, and it takes a dotted path and a set of keyword arguments. A dataset entry looks like this:

yaml
datasets:
  cifar10/train: !import_call
    key: 'torchvision.datasets.CIFAR10'
    init:
      kwargs:
        root: &root_dir '~/datasets/cifar10'
        train: True
        download: True
        transform: !import_call
          key: 'torchvision.transforms.Compose'
          init:
            kwargs:
              transforms:
                - !import_call
                  key: 'torchvision.transforms.RandomCrop'
                  init:
                    kwargs:
                      size: 32
                      padding: 4
                - !import_call
                  key

The tag nests, so transforms compose inside transforms, and a YAML anchor reuses one directory string. Note what the mechanism is: nothing validates the path or restricts it to a registry of known components. The loader is told which object to import and which arguments to pass, which is what lets a user bring their own models and losses without editing the package, and is also why a config file should be reviewed like code before it is run. Dataset keys carry a slash to name a split, cifar10/train.

Cython is a runtime dependency of a framework whose pitch is writing no code

The dependency list is six entries: torch at 2.12.0 or newer, torchvision at 0.27.0 or newer, numpy, pyyaml at 6.0 or newer, scipy, and cython. Five of those are pure Python wheels you would expect from a configuration driven research tool. The sixth is a compiler front end and C extension toolchain, listed as a hard requirement rather than an extra, and the build backend is plain setuptools with no extension configuration visible alongside it. The README's central claim runs the other way. It says the framework helps you design and perform general deep learning experiments without coding, that many components and PyTorch modules are abstracted so you can define them in a declarative config file, and that in many cases you will not need to write Python code at all. Both statements can hold, since a compiled component does not require the user to write code, but a reader budgeting an install will find a build tool in the list before they find that out.

Two of the five examples target transformers, which is not a declared dependency

Five example scripts are named. Three are for torchvision work: image classification over ImageNet ILSVRC 2012, CIFAR-10 and CIFAR-100; object detection over COCO 2017; and semantic segmentation over COCO 2017 and PASCAL VOC. Two live under a directory named hf_transformers: a text classification script for GoEmotions, and a general language understanding script covering the GLUE tasks CoLA, SST-2, MRPC, STS-B, QQP, MNLI, QNLI, RTE, WNLI, plus AX. Neither transformers nor tokenizers appears in the dependency list or in the two optional groups, which are pytest for tests and sphinx, sphinx_rtd_theme, and sphinx_sitemap for docs. So one third of the published examples assume a package the project does not install for you, and the GLUE fine tuned models it publishes to a model hub come from that path. Two architecture families, one dependency set, and the gap falls on the text side.

The version lives in the package, and the last two releases each add one method

pyproject.toml declares no version. It marks the version dynamic and resolves it from an attribute: `version = { attr = "torchdistill.__version__" }`. So the single source of the version number is the package itself, and a file in the repository is the only place to read it. The release history gives the pace. v1.1.3 landed 2025-05-11 with a new text classification example and interface updates. v1.1.4 landed 2025-12-24 with a new method, bug fixes, and the end of Python 3.9 support, which is what `requires-python = ">=3.10"` now reflects. v1.1.5 landed 2026-08-05 with a new method, FSDP and FSDP2 support, and experiment tracking. Two releases roughly seven and a half months apart, each adding one method, against a stated total of 26 methods implemented from venues including TPAMI, CVPR, ICLR, ECCV, NeurIPS, ICCV, and AAAI.

Pretrained CIFAR models are announced with a link to the v0.1.1 release

Some models for CIFAR-10 and CIFAR-100 are reimplemented and shipped as pretrained models inside the project, and the sentence offering them ends by pointing at a release page for v0.1.1. The current line is 1.1.5, so the pointer to the pretrained weights and the pointer to the current version are more than a major version apart, and the announcement itself sits in the README section on examples rather than in a release note of its own. The same section carries two other distribution claims of different reliability. Trained models are available for reproducibility and benchmarking, with trained models, training logs, and configurations all named as published artifacts. And a single benchmark line is given in the README itself, top one validation accuracy for ILSVRC 2012, linked out to a hosted page, while the GLUE results are described only as samples held in the examples directory.

Cite the papers instead of the repository, and the citation is what funds maintenance

The header carries two DOI badges, one for a Springer volume and one for an NLP systems workshop paper, and the text asks that a paper referring to torchdistill cite those papers rather than the repository. A CITATION.bib sits at the root to hold them. The stated reason is unusual and worth quoting in substance: a citation is appreciated and motivates the author to maintain and upgrade the framework. So the project's own continuity argument is bibliographic rather than technical. Alongside that, the framework's history is a rename, since torchdistill was formerly known as kdkit, which means configs, scripts, or notes written against the old name need adjusting, and the package index link in the header points at the current name only.

Training without a teacher is the same config with the teacher entries deleted

The distillation machinery is optional by construction. The claim is that you can train models without teachers by excluding teacher entries from a declarative config file, and the repository carries two notebooks for the same dataset to show it, one for CIFAR knowledge distillation and one for plain CIFAR training. Reading a dataset out of a config is three lines:

python
from torchdistill.common import yaml_util
config = yaml_util.load_yaml_file('./test.yaml')
train_dataset = config['datasets']['cifar10/train']
test_dataset = config['datasets']['cifar10/test']

The intermediate representation story is handled separately, by a forward hook manager that attaches to named modules without touching a model's forward signature, registering hooks by dotted path such as conv1, layer1.0.bn2, or fc, with requires_input and requires_output chosen per hook. That is the mechanism behind the claim that extracting intermediate representations does not require reimplementing models whose forward interface keeps changing.

Editorial conclusion

torchdistill suits a researcher who wants an experiment to be a reviewable file rather than a script, and who will clone the repository rather than install from a package index, because the example configs and the five demo notebooks are excluded from the built distribution and they are the documentation. Verify four things. That the framework still imports what you need for a Hugging Face model, since two of the five example scripts are written for transformers and transformers appears in neither the dependency list nor the extras. That your build environment has what Cython needs, since it is a hard runtime dependency. Which version you are on, because the version number lives in the package rather than in pyproject.toml and the two newest releases each add a single method. And that the pretrained CIFAR models are still where the v0.1.1 release notes point, since the current line is 1.1.5.

Frequently asked questions

What is torchdistill and what does its declarative config replace?

It is a PyTorch framework where models, datasets, optimizers, and losses are defined in a declarative PyYAML file instead of in Python code. Custom modules can be used without editing the local torchdistill package.

Can torchdistill train a model without a teacher?

Yes. You train without teachers by excluding teacher entries from the declarative config file. The repository ships separate demo notebooks for CIFAR knowledge distillation and for plain CIFAR training.

How do I extract intermediate representations without changing a model's forward function?

Use ForwardHookManager and register hooks by module path, for example conv1, layer1.0.bn2, or fc, setting requires_input and requires_output per hook. A demo notebook covers this and is also offered on Colab and SageMaker Studio Lab.

What does pip install for torchdistill leave out?

The package configuration excludes five top level directories: tests, examples, demo, docs, and configs. The example configs and the demo notebooks the documentation points at are therefore in the repository rather than in the installed package.

Which Python versions and libraries does torchdistill require?

Python 3.10 or newer, with torch at 2.12.0 or newer, torchvision at 0.27.0 or newer, numpy, pyyaml at 6.0 or newer, scipy, and cython. Optional groups add pytest for tests and sphinx with its read the docs theme and sitemap extension for documentation.

Official sources

  1. License: MIT
  2. Project website
  3. README
  4. Releases
  5. yoshitomo-matsubara/torchdistill on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/yoshitomo-matsubara-torchdistill.svg)](https://hysenlabs.com/projects/yoshitomo-matsubara-torchdistill)