# PaddleClas: Baidu's industrial toolbox for seeing and sorting

> PaddleClas is the PaddlePaddle image recognition and classification suite, an Apache-2.0 toolbox for industry and academia spanning the PP-HGNet and PP-LCNet backbones, SSLD semi-supervised distillation, ten PULC ultra-light classifiers and the PP-ShiTu image retrieval system. Its 2024 update exposed 98 models through the PaddleX low-code platform, with deployment paths from Python engines to Android and domestic Chinese accelerators.

**PaddlePaddle/PaddleClas** — A treasure chest for visual classification and recognition powered by PaddlePaddle

- Repository: https://github.com/PaddlePaddle/PaddleClas
- Stars: 5,848 · Forks: 1,193
- Language: Python
- License: Apache-2.0
- Published: 2026-09-22 · Updated: 2026-09-22 · Language: en
- Canonical page: https://hysenlabs.com/projects/paddlepaddle-paddleclas

## A suite shaped by industrial deployment

PaddleClas describes itself as the image recognition and classification toolkit that PaddlePaddle prepares for industry and academia, helping users train better vision models and ship applications, and its character shows in what it optimizes for. The flagship backbones are each aimed at a deployment niche rather than a leaderboard, PP-HGNet for higher accuracy at equal inference time on GPUs, PP-LCNet tailored for Intel CPUs with the MKLDNN acceleration library, and PP-LCNetV2 adapted for OpenVINO, choices that only make sense when your users run models on cash registers and gate cameras rather than research clusters. On top sits the SSLD semi-supervised knowledge distillation scheme, the training recipe that turns those backbones into the lighter models the suite ships. The default branch is release/2.6, the last tagged release is v2.6.0 from November 2024, and the repository was last pushed on 2026-09-15.

## PULC: ten classifiers for fixed questions

PULC, the practical ultra-light image classification scheme, addresses the most common industrial shape of the problem, a fixed set of questions about an image that must be answered cheaply. The documented model library lists ten of them, whether a person exists, person attribute recognition, safety helmet wearing, traffic sign classification, vehicle attribute recognition, whether a car exists, text image orientation, text line orientation, language classification and table attribute recognition, each a small model solving a task that recurs across security, retail and document pipelines. The framing is instructive, these are not thousand-class ImageNet generalists but yes-or-no and small-vocabulary specialists, the shape of classification that production systems actually consume, and the quick start guide lets a developer try each before training anything.

## PP-ShiTu: detect, embed, search

The PP-ShiTu image recognition system implements retrieval-based recognition as a pipeline, a mainbody detection model first localizes the object, a feature extraction model embeds it, and a vector search stage matches the embedding against a managed gallery, with deep hashing documented as a compact alternative. Version two, PP-ShiTuV2 from September 2022, raised recall-at-one by eight points and covers more than twenty application scenarios including product recognition, waste sorting and aerial imagery, and the system ships with a gallery management tool plus an Android demo application. The feature models now include CLIP-based variants at ViT-Base and ViT-Large scale alongside the efficient PPLCNetV2 base, and the faiss-cpu dependency in the requirements confirms the vector search machinery. Retrieval rather than fixed classes is what lets the same system recognize a supermarket's changing inventory without retraining.

## 98 models behind a low-code front

The headline of the November 2024 update is the low-code full-process capability through PaddleX, Baidu's development platform built on the suite's technology, which organizes 98 PaddleClas core models into six pipelines callable through one Python API, covering general classification, multi-label classification, general recognition and face recognition. The same API reaches a wider catalog of more than 200 models across detection, segmentation, OCR and time series, composable into modules, with both a unified command line and a graphical interface for usage and customization. Deployment options span high-performance inference, serving and edge, and the hardware story is deliberately broad, NVIDIA GPUs alongside the domestic Kunlunxin, Ascend, Cambricon and Haiguang accelerators, switchable during development. The update also added newer classification backbones, MobileNetV4, StarNet and FasterNet, plus multi-label and face recognition models including MobileFaceNet.

## Deployment from Python to Paddle Lite

The documentation's deployment tree enumerates the exit ramps, inference through the Python prediction engine, through a C++ engine for tighter integration, service deployment through Paddle Serving, and edge deployment through Paddle Lite, the last being what backs the Android demo. The repository supports the same breadth in its tooling, a hubconf.py for model hub integration, a benchmark directory for performance measurement, and the test_tipc suite, the training-inference pipeline checks that validate the whole loop from training configs through exported inference models across combinations of hardware and precision. For a team deciding whether to bet on this stack, the depth of the deployment documentation is the strongest signal, suites that only measure accuracy leave the hard eighty percent of industrial vision work undiscovered, and this one documents the C++ and edge paths with the same care as the Python quick start.

## Four recipes from the supermarket floor

The industrial examples section grounds the abstractions in complete worked scenarios, fresh food self-checkout built on PP-ShiTuV2 recognition, personnel access video management built on PULC person classifiers, smart supermarket goods recognition on PP-ShiTu, and electric scooter detection inside elevators, the distinctly Chinese property-management scenario of keeping battery bikes out of lifts. Each example links a full walkthrough, and together they demonstrate the intended division of labor, retrieval systems where categories change often, fixed ultra-light classifiers where the question is constant, and a bias toward deployment on ordinary hardware near the camera. Reading the recipes before the model lists is the right order for a newcomer, they establish which tool fits which shape of problem, so the subsequent choice among backbones and pipelines is driven by the scenario rather than by benchmark curiosity.

## A package whose keywords are its history

The Python package installs as paddleclas with a matching command-line entry point, requires Python 3.8 or newer, and pins its environment carefully, numpy at 1.24.4 below Python 3.13 and 1.26.4 above, OpenCV held at 4.6, and faiss-cpu for the retrieval stack. The project keywords read like a timeline of vision research absorbed into the suite, knowledge distillation, AutoAugment, CutMix, RandAugment, GridMask, DeiT, RepVGG and Swin Transformer, a catalog of techniques that were once papers and are now configuration options. The README ships in Chinese and English with a separate legacy Chinese edition, the version history documents the release line, and the tree separates the ppcls core from deployment, tools and datasets. It is a production-stable project by its own classifier, and its recent push activity indicates maintenance continues on the release branch between major versions.

## Conclusion

Use PaddleClas when your vision work sits on the PaddlePaddle stack and you want a complete industrial pipeline, training recipes, tuned backbones, ultra-light classifiers, retrieval-based recognition and edge deployment, rather than assembling those from a bare model zoo. Use a PyTorch model collection such as timm when your infrastructure is already PyTorch and you primarily need pretrained backbones. Verify first that your target hardware is supported, since the tuning work concentrates on Intel CPUs, NVIDIA GPUs and the Chinese accelerator lineup, and start from the PaddleX quick start if the low-code path covers your scenario before descending into raw training configs.

## FAQ

### What is PaddleClas?

PaddleClas is the PaddlePaddle image recognition and classification suite, an Apache-2.0 toolbox of models, training methods and deployment paths for industry and research. It contributes the PP-HGNet and PP-LCNet backbones, the PULC ultra-light classifiers, the PP-ShiTu image retrieval system and the SSLD semi-supervised distillation recipe.

### How many models does PaddleClas provide?

Through the PaddleX low-code integration, 98 PaddleClas core models are callable with one Python API across six pipelines, within a wider catalog of more than 200 models spanning detection, segmentation and OCR. The PULC scheme alone documents ten practical classifiers, from person detection to table attribute recognition.

### Which hardware does PaddleClas support?

The PaddleX integration supports NVIDIA GPUs alongside Chinese domestic accelerators including Kunlunxin, Ascend, Cambricon and Haiguang, with switching during model development. The PP-LCNet backbone is tailored for Intel CPUs with MKLDNN, and PP-LCNetV2 adapts to OpenVINO.

## Sources

- [Issues](https://github.com/PaddlePaddle/PaddleClas/issues)
- [License: Apache-2.0](https://github.com/PaddlePaddle/PaddleClas/blob/release/2.6/LICENSE)
- [PaddlePaddle/PaddleClas on GitHub](https://github.com/PaddlePaddle/PaddleClas)
- [README](https://github.com/PaddlePaddle/PaddleClas/blob/release/2.6/README.md)
- [Releases](https://github.com/PaddlePaddle/PaddleClas/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/paddlepaddle-paddleclas
