# VAE-CVAE-MNIST: two papers, three files and one flag

> timbmg/VAE-CVAE-MNIST is a small PyTorch reference implementation of a variational autoencoder and its conditional form on MNIST, with the model chosen by a single command-line flag and every hyperparameter left as an option in the code. The plots are ten epochs at untuned defaults, which makes this a repository to read rather than to benchmark, and the requirements file is where the practical friction sits.

**timbmg/VAE-CVAE-MNIST** — Variational Autoencoder and Conditional Variational Autoencoder on MNIST in PyTorch

- Repository: https://github.com/timbmg/VAE-CVAE-MNIST
- Stars: 664 · Forks: 111
- Language: Python
- License: not declared
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/timbmg-vae-cvae-mnist

## Two named papers, one dataset and a single switch between them

The scope is deliberately narrow and it is stated in the first line: a variational autoencoder and a conditional variational autoencoder, both on MNIST, both in PyTorch. Each is tied to a specific paper rather than to a general idea. The plain model follows Auto-Encoding Variational Bayes, and the conditional one follows Semi-supervised Learning with Deep Generative Models. The interesting engineering decision is that the two models are not two projects. Running the conditional version means adding a single flag to the same command, and everything else about the invocation stays the same. That is a better design for a teaching repository than a pair of near-duplicate scripts, because the difference you are meant to notice is the conditioning, not the plumbing. The corollary is where the documentation stops. Hyperparameters such as the learning rate, the batch size, and the depth and size of the encoder and decoder layers are not in a table or a config file; they are command-line options in the code, and the instructions simply point you at them. The defaults are the lesson, so leaving them in the argument parser is defensible, and it also means no number in this repository can be reproduced without reading the source.

## Ten epochs, default settings, and nothing tuned

The results section states its own conditions before showing anything: every plot was produced after ten epochs of training, using the default settings in the code, with no tuning applied. That sentence is the most useful thing in the readme, because it tells you exactly what not to conclude. Ten epochs on handwritten digits is enough for the latent space to organise itself by digit class, which is the visible result in the first figure, and it is nowhere near enough to say anything about model quality. The first figure makes the sampling explicit: the modelled latent distribution after those ten epochs, drawn with a hundred samples per digit, so a thousand points in a two-dimensional latent space with ten clusters in it. That is a demonstration of the encoder's behaviour, not a measurement of it. There is no held-out likelihood, no reconstruction error table, no comparison against another method and no seed count anywhere in the repository. So read this as a reference implementation whose figures exist to show that the code runs and produces the expected structure. If you need to know whether one of these models is better than something else, nothing here will answer that, and the absence is deliberate rather than an oversight.

## The second figure asks the decoder to invert the first one

The two figures cover the two directions of the model, which is the right way to present a variational autoencoder. The first asks whether the encoder produced a structured latent space. The second samples z at random and shows what comes back, and for the conditional model each class label has been supplied exactly once. So the pair reads as a round trip: structure on the way in, plausible images on the way out. The conditional variant is where the two columns diverge, because a conditional model can be asked for one class at a time instead of sampling the label as part of the latent draw, and that is visible in a grid of digits grouped by label. Worth noting how the comparison is presented: each figure is a side-by-side pair of images, one for the plain model and one for the conditional, rather than a table of numbers. Images are the right medium for a qualitative check, because clustering and sample quality are visual properties, and they are the wrong medium if you want to argue that one column is better. Nothing in the repository quantifies that comparison, so any claim beyond it is yours to support. The sample count for the generative figure is also unstated, unlike the hundred per digit in the latent figure, which tells you the two plots were not produced under a shared sampling protocol.

## Three Python files, and the hyperparameters live inside them

The repository is models.py, train.py and utils.py, a figures directory, a requirements file and the readme. That is the conventional split for this kind of project: model definitions, the training loop, and helpers, with generated figures kept out of the source. There is no configuration directory, no notebook, no test directory, no pretrained checkpoint and no dataset download script in the listing, so the data pipeline is whatever the training script does on its own, and the only record of a successful run is a picture. Three files is the constraint that shapes every other property here. It is small enough to read end to end in an afternoon, which is the entire reason to keep a repository this size, and it is also small enough that nothing is checked automatically. If you extend it, you are reading source to find out what the defaults are, because there is no other record of them, and the natural first step is to copy the argument parser out of the training script into something you can version. That is a reasonable thing to do and worth doing deliberately rather than by accident, since the defaults are the part of this repository that carries meaning.

## The requirements file is a frozen plotting environment plus a recent torch

```
cycler==0.10.0
kiwisolver==1.0.1
matplotlib==3.0.2
numpy==1.22.0
pandas==0.24.1
Pillow>=6.2.2
pyparsing==2.3.1
python-dateutil==2.8.0
pytz==2018.9
scipy==1.10.0
seaborn==0.9.0
six==1.12.0
torch==2.2.0
torchvision==0.2.1
``` Fourteen entries, thirteen of them pinned to an exact version, and the single exception is the imaging library, which has a floor and no ceiling. What is pinned is dominated by the transitive closure of a plotting stack rather than by anything a variational autoencoder needs: a solver used to render mathematical text inside figures, a cycler, a parser, a date utility, a timezone database, a compatibility shim for Python 2 style imports, a dataframe library and a statistical plotting library. None of those are training dependencies. They are what an environment accumulates when someone installs a plotting library once and freezes whatever the resolver produced, which is the most likely history of this file. The deep learning side is pinned separately and much more tightly, with the framework and its vision library on exact versions. So the file was assembled in two eras and never revisited, and that is worth knowing before you spend an evening on an install failure that has nothing to do with the model. If you want a working environment, the sensible move is to let a resolver pick the plotting stack and pin only the tensor pair.

## torch 2.2.0 with torchvision 0.2.1 is the line to change first

The vision library is how the dataset arrives here, since MNIST is loaded through it, so the pair matters more than most of the other pins. The file asks for the framework at 2.2.0 and the vision library at 0.2.1, and those two are not the combination that ships together. The vision release that accompanies the framework's 2.2 line sits in the 0.17 series; the 0.2 series belongs to the framework's 1.x era. When a compiled extension is built against a different set of internals than the framework it is imported by, you get an import-time failure or an operator error the first time a convolution runs, and neither error mentions version numbers in a way that points at this file. It is a one-line edit, and it is the most likely single reason this repository does not run on a current setup. The broader lesson is about frozen requirement files in general: an exact pin is only useful while it stays true, and a pin that was correct when written becomes a liability as the ecosystem moves. Everything else in the list, from the plotting stack to the imaging floor, is a smaller version of the same problem.

## No releases, no licence file, and a master branch

Two facts about the repository's packaging are worth knowing before you build on it. There are no published releases, so there is no tagged version to pin, no changelog to read and no artefact to download; you clone a branch. The default branch is called master rather than main, which places the project in the period before that convention changed and is a small but reliable signal about when the code was last touched. The last commit is dated 17 July 2026, so this is not abandoned work, and the freshness of that date is the strongest signal available about intent. The licence is the real gap. The repository's licence field is not set and there is no licence file in the top-level listing, which means the code is readable and copyable but carries no grant. For a repository whose value is that you can read it and take the structure, that is an awkward omission, and for anyone who wants to lift the model definitions into their own project it is a blocker rather than a detail. If that matters to you, ask the author. Do not infer permission from the absence of a restriction, because that is not what an unlicensed repository means.

## Conclusion

Use this repository if you want to read a working variational autoencoder and a conditional one side by side on the same dataset, because three files and one flag is the fastest honest way to see what the conditioning actually changes, and the two papers it implements are linked rather than paraphrased. Do not use it if you need a tuned result, a benchmark number, or a component you can drop into a pipeline: there is no test suite, no checkpoint, no configuration file, and the reported figures come from ten epochs at whatever the defaults happen to be. Before you run anything, fix the dependency file, since the framework and the vision library are pinned to versions that were not released as a pair and that mismatch is the first reason an install fails. And settle the licence question before you copy any of it, because the repository carries no licence file and no licence field, so nothing in it grants you permission to reuse it. The code was last touched on 17 July 2026, so the intent is current even where the pins are not.

## FAQ

### How do I run the conditional version of timbmg/VAE-CVAE-MNIST?

Add the conditional flag to the same command. Every other setting, including the learning rate, the batch size and the depth and size of the encoder and decoder layers, is a command-line option inside the code rather than something listed in the documentation.

### Which papers does this VAE and CVAE MNIST repository implement?

The plain model follows Auto-Encoding Variational Bayes and the conditional model follows Semi-supervised Learning with Deep Generative Models. Both links are given at the top of the readme rather than paraphrased.

### How were the sample images in this repository produced?

Every plot comes from ten epochs of training with the default settings in the code and no tuning. The latent distribution figure uses a hundred samples per digit, so a thousand points across ten classes.

### What files does timbmg/VAE-CVAE-MNIST contain?

Three Python files, a models file, a training script and a utilities file, plus a figures directory, a requirements file and the readme. There is no configuration file, notebook, test directory or pretrained checkpoint in the repository listing.

### Does the repository report benchmark numbers for the VAE or CVAE?

No. The readme presents side-by-side images for the latent distribution and for sampled outputs, and states that they come from ten epochs at default settings with nothing tuned. There is no held-out score and no comparison against another method.

### Which versions does the requirements file pin?

The deep learning framework at 2.2.0 and its vision library at 0.2.1, plus exact pins across a plotting stack including the dataframe, plotting and numerical libraries. The imaging library is the only entry given as a range rather than an exact version.

## Sources

- [Issues](https://github.com/timbmg/VAE-CVAE-MNIST/issues)
- [README](https://github.com/timbmg/VAE-CVAE-MNIST/blob/master/README.md)
- [timbmg/VAE-CVAE-MNIST on GitHub](https://github.com/timbmg/VAE-CVAE-MNIST)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/timbmg-vae-cvae-mnist
