# timm: the changelog admits precision fixes change low-precision outputs

> timm, the pytorch-image-models package, is a collection of image encoders and backbones with pretrained weights and its own training, evaluation, inference, and ONNX export scripts. It moves fast, adding Qwen, DeepSeek, and Sapiens vision towers in September 2026, and its own release notes warn that a RoPE precision fix changes half precision outputs.

**huggingface/pytorch-image-models** — The largest collection of PyTorch image encoders / backbones. Including train, eval, inference, export scripts, and pretrained weights -- ResNet, ResNeXT, EfficientNet, NFNet, Vision Transformer (ViT), MobileNetV4, MobileNet-V3 & V2, RegNet, DPN, CSPNet, Swin Transformer, MaxViT, CoAtNet, ConvNeXt, and more

- Repository: https://github.com/huggingface/pytorch-image-models
- Website: https://huggingface.co/docs/timm
- Stars: 37,172 · Forks: 5,203
- Language: Python
- License: Apache-2.0
- Published: 2026-08-17 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/huggingface-pytorch-image-models

## Precision fixes are documented as output-changing

The RoPE refactor in the September 25, 2026 entry says two things that matter more than the feature list around them. Non-learned RoPE frequency and coordinate buffers are now kept at least float32 in half and bfloat16 models, and mixed RoPE phases are computed in float32 under autocast. The entry then states plainly that these precision fixes can change low-precision model outputs.

A second entry in the same block says DINOv3 with grid_indexing set to 'xy' now swaps coordinate channels as requested, and that outputs for these previously incorrect configurations change. The refactor also added grid_type and normalize_coords options to the Fourier and rotary builders, and it consolidated two axial RoPE modules behind a private base, so the bands buffer is no longer None in cached mode even though existing calls stay compatible.

The consequence is that upgrading timm can change your numbers without a single line of your code changing, which means a benchmark table comparing two timm versions is comparing two different computations. It also means the old and new behaviour are both reachable, so reproducing a specific result requires knowing which timm version produced it rather than which one is current.

## Multi-label arrives with three target encodings and a different smoothing default

Multi-label classification support arrived on September 25, 2026 behind a task flag added to train.py and validate.py, trained with BCE through a task class called MultiLabelClassificationTask. Targets can be written three ways: as lists of class indices, as dense multi-hot vectors using a target format flag, or as one binary field per class using a target key that takes a comma separated list, with forward slash paths allowed into nested fields. The hfds, hfids, tfds, and wds readers support it, and Mixup, CutMix, and the NaFlex loaders all handle multi-label targets.

Two defaults move with the task. Label smoothing now moves each target towards 0.5 and defaults to 0 for multi-label, and validation reports mAP as the default evaluation metric alongside micro, macro, and sample F1, where the F1 threshold comes from its own flag. A separate addition computes dataset-level metrics across distributed ranks while excluding distributed sampler padding.

The consequence is that a config written for single-label classification carries over with a silently different smoothing value, and the three F1 numbers are not a single quantity because they depend on a threshold you have to choose. The padding exclusion matters for distributed runs, where the same script can report different denominators on different world sizes.

## Checkpoint loading was hardened twice, and the allowlist is yours to extend

The March 23, 2026 entry is about security rather than accuracy. Pickle checkpoint handling was improved, all loading defaulted to weights_only set to True, and a safe_global mechanism was added for ArgParse. The July 10, 2026 entry adds a second pass, described as hardening pickle loading again alongside custom-label inference improvements. Two scripts at the repository root exist to support this work, clean_checkpoint.py and class_weights.py.

The consequence is a change in the default contract. A checkpoint written by an older version that pickled more than tensors will now be refused unless the weights_only default is turned off, and where you do need the extra globals, safe_global is an allowlist you extend per checkpoint rather than a general permission. Doing this twice in four months also means the behaviour you tested against an intermediate version is not the behaviour you get at the head of main.

## The wheel metadata and the requirements file list different dependencies

The pyproject declares five dependencies: torch, torchvision, pyyaml, huggingface_hub, and safetensors, with no version constraints on any of them. The requirements file at the root lists the same five with floors, torch at 1.7 or higher, huggingface_hub at 0.17.0 or higher, and safetensors at 0.2 or higher, and adds numpy as a sixth entry.

The consequence is that installing the wheel does not bring numpy, and installing from the requirements file does. The torch floor of 1.7 exists only in the text file, so a pip install against whatever torch happens to be in the environment raises no objection, however old it is. Two lists that differ by one package is the kind of gap that only shows up as an ImportError in someone else's environment, and nothing in the root listing says which file is authoritative for a normal install.

## The version is read from timm/version.py and Python support stops at 3.12

The build uses pdm-backend and marks the version as dynamic, sourcing it from a file at timm/version.py. The package name is timm, the author is Ross Wightman, the licence is Apache-2.0, and the project URL points at the documentation on huggingface.co rather than at a site hosted in the repository. The test extras are pytest with timeout, xdist, forked, and expecttest, and the pytest configuration sets testpaths to the tests directory with markers for the base and cfg model test setups.

The requires-python floor is 3.8, and the classifiers enumerate 3.8 through 3.12. The most recent releases are v1.0.28, v1.0.29, and v1.0.30, dated 2026-07-11, 2026-08-28, and 2026-09-22.

The consequence is that a build from a dirty working tree reports whatever that single file says, so an unreleased build and the last tag can carry the same string. The declared interpreter support also stops at 3.12 in the classifiers even though the newest model families being added are from 2026, so the metadata is describing the floor and the tested range rather than the present.

## Two training scripts sit in the root and the answer is in UPGRADING.md

The root is a script directory: train.py, validate.py, inference.py, benchmark.py, onnx_export.py, onnx_validate.py, bulk_runner.py, avg_checkpoints.py, class_weights.py, clean_checkpoint.py, and hubconf.py, alongside distributed_train.sh and two files carrying the legacy prefix, legacy_train.py and legacy_distributed_train.sh. A convert directory holds conversion scripts, hfdocs holds documentation sources, and results holds outputs. A file named UPGRADING.md sits at the root next to CONTRIBUTING.md, AGENTS.md, and CLAUDE.md.

The consequence is that nothing in the file listing tells you which training entry point is current. A new user has two plausible scripts and no marker on either, and the answer lives in a separate document that has to be found first. The legacy scripts are kept rather than deleted, which is kind to anyone with an existing pipeline and unhelpful to anyone learning the flags, because the option surface of the two scripts is not something the root listing lets you compare.

## New vision towers and changed defaults ship in the same release

The September 10 to 11, 2026 block added Qwen3-VL, Qwen3.5, and Qwen3.8 vision transformer classifier and encoder variants, including classifiers with and without the native spatial merger, plus Sapiens2 vision transformers with EVA and NaFlexViT support, iFormer, EfficientViM, and DeepSeek-V4 and V4.1 classifiers and encoders, each with native weights on the timm Hub. The same block switched the default NaFlex SigLIP position interpolation, changed how inference.py selects input size, improved model kwargs parsing, and fixed meta-device construction for a few models.

Earlier entries follow the same shape. A Qwen-Drive vision tower was added on September 22. On August 11, CPUBone, PP-LCNetV2, and LingBot-Vision arrived with per-batch image and batch size scheduling for non-NaFlex training, including progressive small-to-large resolution schedules.

The consequence is that a default changed in the same release as new models, so a run that reproduced last week can diverge this week with no change on your side. The other consequence is what these families are: vision towers and encoders taken from language model releases, which makes timm a distribution point for third-party checkpoints as much as a zoo of published architectures.

## Conclusion

timm fits a vision researcher who wants a model family and its pretrained weights in one place, with a real training script rather than a notebook, and who can pin a commit when numbers have to hold. It does not fit someone who needs identical results across an upgrade, because the project states outright that precision fixes alter low-precision outputs, and it does not fit a team that treats the requirements file and the wheel metadata as interchangeable, because they list different dependencies. Before you adopt a version, pin it and record it, read UPGRADING.md rather than guessing which training script is current, and check whether the default NaFlex position interpolation or inference input-size selection changed in the release you are taking.

## FAQ

### What is a PyTorch model?

In this repository a model is one of the image encoders and backbones the timm package provides, created through its model factory and usually loaded with pretrained weights from the timm Hub, with hubconf.py at the repository root. The collection covers ResNet, ResNeXT, EfficientNet, NFNet, ViT, MobileNetV4, RegNet, DPN, CSPNet, Swin, MaxViT, CoAtNet, ConvNeXt and more.

### What is the main purpose of PyTorch?

The repository answers that question only for its own scope. pytorch-image-models exists to collect PyTorch image encoders and backbones together with training, evaluation, inference, and export scripts and pretrained weights, and the timm package depends on torch, torchvision, pyyaml, huggingface_hub, and safetensors.

### Is ChatGPT built on PyTorch?

The repository does not cover that. It is a collection of image encoders and backbones, and its newest additions are vision towers and encoders such as Qwen3-VL and DeepSeek-V4 variants rather than language models. Nothing in the project addresses how ChatGPT is built.

### Is Yolo TensorFlow or PyTorch?

The repository does not include YOLO, so the comparison does not apply to it. Its scope is a different set of architectures, named in the repository description as ResNet, ResNeXT, EfficientNet, NFNet, Vision Transformer, MobileNetV4, MobileNet-V3 and V2, RegNet, DPN, CSPNet, Swin Transformer, MaxViT, CoAtNet, and ConvNeXt.

## Sources

- [Official documentation](https://huggingface.co/docs/timm)
- [Official README](https://github.com/huggingface/pytorch-image-models#readme)
- [Project repository](https://github.com/huggingface/pytorch-image-models)
- [Release notes](https://github.com/huggingface/pytorch-image-models/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/huggingface-pytorch-image-models
