ResNeSt: Split-Attention Backbones for PyTorch and Gluon
ResNeSt: Split-Attention Networks
At a glance
- What is it?
- ResNeSt is a ResNet variant that adds split-attention blocks to the standard bottleneck. It ships pretrained ImageNet weights and drop-in backbones for Detectron2, but its training recipes live outside this repository.
- Who is it for?
- Adopt ResNeSt if you need a drop-in ResNet backbone with published ImageNet weights and a Detectron2 wrapper for detection or segmentation fine-tuning. Do not adopt it if you need an actively developed training framework: the last push was on 2026-09-11, but the newest release on the release page is weights_step2 from 2021-05-17, and the README points PyTorch ImageNet training at the external PyTorch Encoding Toolkit rather than this repo.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 23 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What ResNeSt Is and Who It Is For
ResNeSt stands for Split-Attention Network, described in the README as "A New ResNet Variant" that "significantly boosts the performance of downstream models such as Mask R-CNN, Cascade R-CNN and DeepLabV3." The project is a backbone library, not an end-to-end training framework. It targets engineers who already have a detection, segmentation or classification pipeline and want to swap the feature extractor without rewriting the rest of the model.
The repository ships four classification variants with published ImageNet accuracy in the README table: ResNeSt-50 at crop size 224, ResNeSt-101 at 256, ResNeSt-200 at 320 and ResNeSt-269 at 416. The README lists PyTorch and Gluon numbers side by side, for example 81.03 and 81.04 for ResNeSt-50. If your work is downstream of ImageNet classification, the transfer learning section points at a Detectron2 wrapper in the d2 directory and notes that MMDetection has adopted the ResNeSt backbone in its own configs.
It is the wrong tool if you want a maintained training codebase. The README sends PyTorch ImageNet training to the external PyTorch Encoding Toolkit and describes that path as "slightly worse than Gluon implementation." The repository itself is the model definition, the weights and the verification scripts.
Split-Attention Inside the Bottleneck
The mechanism is a modification of the standard ResNet bottleneck. Instead of one convolutional path per block, the feature channels are split into groups, and each group is processed separately before the outputs are combined with attention weights. This is the split-attention idea the paper title refers to. The README does not walk through the block internals, so the architecture details are best read from the arXiv paper linked at the top of the repository.
What the repository does show is the interface layer. There is a hubconf.py at the top level for Torch Hub loading, a resnest package with separate torch and gluon submodules, a d2 directory for the Detectron2 wrapper, and a configs directory. The scripts directory holds dataset preparation and verification code, split into scripts/dataset, scripts/torch and scripts/gluon. That layout tells you the intended data flow: prepare ImageNet once, then verify the PyTorch or Gluon weights against the published accuracy numbers.
One design decision is worth flagging. The README notes that "the inference speed reported in the paper are tested using Gluon implementation with RecordIO data." If you plan to use the PyTorch path, the speed figures in the paper do not directly describe your setup. The accuracy table is separate and lists both frameworks, so you can compare accuracy across implementations but not assume the same throughput.
Installing resnest and Loading ResNeSt-50
The README gives two install options and says you only need to choose one. The PyPI path installs the pre-release, which the badge at the top of the README labels v0.0.6. The setup.py in the repository confirms the base version string as '0.0.6' and appends a date stamp when the RELEASE environment variable is not set.
pip install resnest --preAfter installing, the README shows loading a pretrained model through the package. ResNeSt-50 is the example used throughout.
from resnest.torch import resnest50
net = resnest50(pretrained=True)The alternative is Torch Hub, which pulls the model definition from the GitHub repository rather than the installed package. The README uses force_reload=True when listing models, which is what you want after a fresh clone or a hub cache change.
import torch
torch.hub.list('zhanghang1989/ResNeSt', force_reload=True)
net = torch.hub.load('zhanghang1989/ResNeSt', 'resnest50', pretrained=True)To check the weights against the published number, the README provides a verification script. You need ImageNet prepared first, and the README offers a download helper in scripts/dataset.
cd scripts/dataset/
python prepare_imagenet.py --download-dir ./Then, from scripts/torch, the verification command takes the model name and crop size as flags. For ResNeSt-50 the README uses a crop size of 224, matching the table.
cd scripts/torch/
python verify.py --model resnest50 --crop-size 224The README does not document what the script prints on success, so treat the published 81.03 figure as the reference point rather than an expected console string.
Where the Repository Stops and Other Projects Begin
The clearest limitation is scope. Training ImageNet models with PyTorch is not handled here; the README directs you to the PyTorch Encoding Toolkit. Training with MXNet Gluon is handled in the scripts/gluon folder, but the README does not present that as a general-purpose trainer either.
For detection and segmentation, the README points to a separate fork, detectron2-ResNeSt, and to the d2 directory in this repository for the wrapper. That means your upgrade path spans at least two repositories. If the fork lags behind upstream Detectron2, you inherit that gap. The README does not document a compatibility matrix between the d2 wrapper and Detectron2 versions.
Release cadence reinforces this. The release list shows v0.0.5 in June 2020, then two weight-release steps in May 2021, and nothing after that. The last push to the repository was on 2026-09-11, so the code has seen activity, but the release page has not produced a tagged model release in years. Anyone who pins to a release tag is pinning to something old. Anyone who installs from PyPI gets the pre-release channel, which is a different kind of risk: pre-release packages can change without a stable version bump.
There is also a licensing question the README does not answer. The code is Apache-2.0, and setup.py carries the same identifier. The README does not state the licence or terms attached to the pretrained weights themselves, and those are downloaded separately through Torch Hub or the package. That distinction matters if you ship a product.
ResNeSt Versus Plain torchvision ResNet
The obvious alternative is torchvision's ResNet family, which most PyTorch users already have installed. The difference is in the block. A torchvision ResNet-50 uses the standard bottleneck with no channel grouping or attention. ResNeSt replaces that with the split-attention block and reports higher ImageNet accuracy for the same depth: the README table gives ResNeSt-50 at 81.03 with PyTorch, while the ResNet-50 baseline is not listed in this repository at all.
The trade-off is not just accuracy. ResNeSt adds parameters and computation per block, and the README's own note about inference speed being measured on the Gluon implementation means you cannot read a PyTorch latency number off this page. If your constraint is latency on a fixed budget, the split-attention block is a cost you have to measure yourself.
The second alternative is to skip this repository and use the backbone through MMDetection, which the README says has adopted it. That gets you a maintained detection framework with ResNeSt as a config option, at the price of adopting MMDetection's training conventions. If you only need classification, the package route shown above is shorter.
Maintenance, Versioning and Upgrade Cost
The repository is not archived, and the last push was on 2026-09-11. That is recent enough that the code is not abandoned, but it says nothing about whether new features are landing. The release history is the better signal for that, and it is thin: v0.0.5 in June 2020 and two weight releases in May 2021.
Versioning has a quirk that affects upgrades. In setup.py, when the RELEASE environment variable is absent, the version gets a date suffix appended in the form bYYYYMMDD. That means installing from source without RELEASE set produces a version string that changes every day. Pin your dependency if you build from source, or install from PyPI and accept the pre-release channel.
Dependencies are listed in setup.py and are ordinary: numpy, tqdm, nose, torch>=1.0.0, Pillow, scipy, requests, iopath and fvcore. The nose test dependency is a legacy choice; nose has not been the default in most Python test setups for a long time, and it is pulled in unconditionally by install_requires. That is a small but real cost if you are building a minimal container.
On licensing, the repository and setup.py both carry Apache-2.0. That covers the code in this tree. The README does not describe the terms for the pretrained weights, and the third-party implementations it links to (TensorFlow, Caffe, JAX) are separate projects with their own licences. Check those individually before redistribution; this is not legal advice.
Editorial conclusion
Adopt ResNeSt if you need a drop-in ResNet backbone with published ImageNet weights and a Detectron2 wrapper for detection or segmentation fine-tuning. Do not adopt it if you need an actively developed training framework: the last push was on 2026-09-11, but the newest release on the release page is weights_step2 from 2021-05-17, and the README points PyTorch ImageNet training at the external PyTorch Encoding Toolkit rather than this repo. Before committing, verify that the pretrained checkpoint for your chosen variant downloads through torch.hub or resnest50(pretrained=True), and confirm which licence terms attach to the weights you pull.
Frequently asked questions
What exactly is ResNeSt?
ResNeSt is a ResNet variant built around split-attention blocks, described in the README as a new ResNet variant that boosts downstream models such as Mask R-CNN, Cascade R-CNN and DeepLabV3. The repository provides the model definitions, pretrained ImageNet weights and verification scripts.
What does ResNeSt stand for?
The name stands for Split-Attention Network, as stated in the README heading and the paper title on arXiv. The split-attention block is the architectural change relative to a standard ResNet bottleneck.
Is ResNeSt-50 still relevant?
The repository still publishes ResNeSt-50 weights and the README lists 81.03 ImageNet accuracy for the PyTorch implementation at crop size 224. Whether it fits your work depends on your latency budget, since the README reports inference speed only for the Gluon implementation with RecordIO data.
Is ResNeSt a CNN or a Transformer?
ResNeSt is a convolutional network. It is described as a ResNet variant, and the transfer learning section points to CNN-based detectors and segmentation models such as Mask R-CNN and DeepLabV3.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/zhanghang1989-resnest)