Library / SDK
wuji3/Doraemon avatar
wuji3/Doraemon

Doraemon: a visual modelling library whose PyPI build is still versioned 0.0.10a

A powerful baseline for image classification, face recognition and image retrieval with Pytorch

807 stars101 forksPythonGPL-3.0

At a glance

What is it?
Doraemon wraps timm backbones, metric-learning heads, training utilities and Hugging Face deployment into one visual modelling library for classification, retrieval and face recognition. The packaging tells a different story from the changelog, and a third of the headline tasks is only usable internally.
Who is it for?
Doraemon is aimed at someone who wants metric-learning and GradCAM tooling wired into one training loop rather than assembling it from separate repos, and who is already using timm. Four things to check first.
Can I use it commercially?
Yes, with conditions. GPL-3.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
Is it still maintained?
Yes. The repository last received commits 154 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 6, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The published version has never caught up with the release tag

Two version numbers describe the same project and they disagree. The GitHub release is v0.1.0, dated 2025-03-19, and the changelog records the same date for the v0.1.0 release. The packaging script says `version='0.0.10a'`, which is a pre-release string, and it was never advanced. So anyone installing from an index gets a build stamped 0.0.10a indefinitely, and a version check in your own environment will not tell you which commit you are on. The distribution name diverges too. The repository is Doraemon, the setup script publishes `name='doraemon-torch'`, and that is the string you install:

bash
# Create and activate environment
python -m venv doraemon
source doraemon/bin/activate

# Install from PyPI
pip install doraemon-torch

# Or install in editable mode for development
pip install -e .

The editable route is the one the documentation offers for development, and given the version mismatch it is also the only way to be certain what you have.

Only the doraemon package is published, so configs and tools stay in the repo

The setup script restricts discovery to `find_packages(include=['doraemon', 'doraemon.*'])`, which means only the library itself lands in the distribution. The repository root also carries `configs/`, `data/`, `deploy/`, `misc/`, `scripts/` and `tools/`, and none of those are inside the included pattern. This matters more than a normal packaging detail, because the deployment documentation lives at `deploy/README.md` and the classification and retrieval walkthroughs live under `doraemon/models/`. The parts of the repository that explain how to run anything are split across the included package and the excluded directories, so an installed copy gives you the code and not the surrounding material. The dependency list also requires Python 3.10 or newer, which is the floor you get alongside that arrangement.

torchaudio and a CPU-only vector index sit in a vision library

Sixteen packages are declared with lower bounds and no upper bounds, and two of them do not obviously belong. `torchaudio>=2.5.1` is a hard requirement of a library whose documented tasks are image classification, content-based image retrieval and face recognition, with no audio task anywhere in the task table. `faiss-cpu>=1.7.2` is the other: it pins the retrieval index to the CPU build specifically, so on a machine with a GPU the nearest-neighbour search for content-based retrieval still runs on the processor unless you replace that dependency yourself. The rest is a conventional vision stack: torch, torchvision, timm, transformers, opencv-python, Pillow, numpy, torchmetrics, tensorboard, tqdm, grad-cam, datasets, imagehash and prettytable. That last pair, image hashing for perceptual comparison and prettytable for console output, tells you the library expects duplicate-oriented work and readable training logs, though neither is explained in the highlights.

The same file is cited for optimizers and for metric-learning heads

The highlights index points `doraemon/engine/optimizer.py` twice. The first entry attributes SGD, Adam, SAM and layer-specific learning rates to it. The second attributes label smoothing, OHEM, Focal Loss, ArcFace, CircleLoss and MagFace to the same path. The first three losses are training objectives; the last three are margin-based metric-learning heads, which is what the retrieval task needs and is a different kind of object from an optimizer. Either the heads live beside the optimizer or the highlights table is grouping unrelated things under one filename, and the text does not say which. Other paths are unambiguous: `doraemon/engine/vision_engine.py` holds the unified training flow, `doraemon/dataset/transforms.py` holds CutOut, ColorJitter, Copy-Paste, Mixup and class-specific augmentation, and `doraemon/utils/cam.py` holds the GradCAM interpretation used for bad-case analysis.

Face recognition is listed as in progress and its tutorial says stay tuned

The task table has three rows and two statuses. Image classification is supported, content-based image retrieval is supported, and face recognition is in progress, with the explanation that training on MS-Celeb-1M-v1c and evaluation on LFW are supported internally while the end-to-end public pipeline is still being organised. The tutorials section repeats the same gap in a shorter form, listing a documentation page for classification, one for retrieval, and for face recognition simply that it is to come. So one of the three tasks in the library's own description does not have public documentation, and the training and evaluation capability that does exist is described as internal. That is an unusually candid status line, and it is also a clear boundary for anyone planning around this library: budget for the retrieval and classification paths, and treat face recognition as something to verify rather than something to install.

Every example dataset is one Hugging Face repository under a single account

The three dataset entry points are all Hugging Face datasets published by the same account: `wuji3/oxford-iiit-pet` for classification, `wuji3/image-retrieval` for retrieval, and `wuji3/face-recognition` for face recognition. The retrieval set is described in the changelog as product data collected from Kaggle and TianChi, released 2024-10-01. That gives you a working start for each task without assembling data yourself, and it also means every example depends on one namespace being available. The backbone story is broader by comparison: the library claims support for more than a thousand visual backbones through timm, everything returned by `timm.list_models(pretrained=True)`, with CLIP, SigLIP, DeiT, BEiT, MAE, EVA, DINO, ResNet, Swin Transformer and ViT named as examples. For choosing between them it points outward, at the benchmark results in the pytorch-image-models repository and at issue 1933 there, rather than publishing its own numbers.

One tag in March 2025, a paper in November, and pushes into May 2026

The changelog dates run from 2023-06-01, when image classification support shipped with Oxford-IIIT Pet examples, hard example mining, GradCAM visualisation, auto-labelling and class-specific augmentation. Face recognition followed on 2024-04-01, retrieval on 2024-10-01, the v0.1.0 release on 2025-03-16, and a paper on arXiv on 2025-11-07. The repository was last pushed on 2026-05-05 and has exactly one GitHub release. Deployment has two routes and they carry different obligations: locally, run a trained `*.pt` weight file together with its configuration and inference code, or publish to Hugging Face and load through `AutoModel.from_pretrained()` and `AutoProcessor.from_pretrained()`. The first is self-contained. The second leans on the transformers dependency and puts your weights behind a schema the library defines. The licence is GPL-3.0, which is the part to read first if either route ends up inside a product.

Editorial conclusion

Doraemon is aimed at someone who wants metric-learning and GradCAM tooling wired into one training loop rather than assembling it from separate repos, and who is already using timm. Four things to check first. Pin your install to a commit if you care about reproducibility, because the PyPI version string has never reached the release tag and the changelog has entries with no corresponding version bump. Expect the editable install, since only the doraemon package is published and configs, deploy, tools and data stay in the repository. Plan for the CPU in retrieval, because the vector search dependency is the CPU build. And if face recognition is your target, verify what you actually get, because the task table marks it in progress and the tutorial section says only that it is to come. The GPL-3.0 licence is the other thing to read before this enters anything commercial.

Frequently asked questions

What package name do I install for Doraemon?

`pip install doraemon-torch`. The setup script publishes the distribution under `doraemon-torch` even though the repository is named Doraemon, and it declares `version='0.0.10a'` while the GitHub release tag is v0.1.0. For development the repository also offers `pip install -e .`.

Is face recognition usable in Doraemon today?

The task table marks it in progress. Training on MS-Celeb-1M-v1c and evaluation on LFW are described as supported internally, while the end-to-end public pipeline is still being organised, and the tutorials section lists face recognition as to come rather than linking a document.

Which datasets does Doraemon ship entry points for?

Three, all on Hugging Face under the wuji3 account: oxford-iiit-pet for image classification, image-retrieval for content-based retrieval, and face-recognition. The retrieval dataset is described as product data collected from Kaggle and TianChi.

What does Doraemon need installed alongside it?

Sixteen declared dependencies with lower bounds and no upper bounds, including torch, torchvision, timm, transformers, opencv-python, grad-cam, datasets, imagehash and prettytable, plus torchaudio and faiss-cpu. Python 3.10 or newer is required. Because the vector search dependency is the CPU build, retrieval indexing does not use a GPU.

How do I deploy a model trained with Doraemon?

Two ways. Locally, run a trained `*.pt` weight file with the deployment configuration and inference code. Or publish it to Hugging Face and load it with `AutoModel.from_pretrained()` and `AutoProcessor.from_pretrained()`. Detailed instructions are in the deployment guide under deploy/.

Official sources

  1. Issues
  2. License: GPL-3.0
  3. README
  4. Releases
  5. wuji3/Doraemon on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/wuji3-doraemon.svg)](https://hysenlabs.com/projects/wuji3-doraemon)