# machine-learning-visualized is a book build, and the notebooks live in seven other repositories

> gavinkhung/machine-learning-visualized builds a Jupyter Book whose notebooks are downloaded at build time from seven sibling repositories. The website is regenerated by a workflow after every commit, and the EPUB path lists its four chapters by hand.

**gavinkhung/machine-learning-visualized** — ML algorithms implemented and derived from first-principles in Jupyter Notebooks and NumPy

- Repository: https://github.com/gavinkhung/machine-learning-visualized
- Website: https://ml-visualized.com/
- Stars: 1,982 · Forks: 187
- Language: TeX
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/gavinkhung-machine-learning-visualized

## The repository is configuration, and the notebooks are fetched

The most important thing about this repository is what it does not contain. It is a Jupyter Book, a tool for building a website from Markdown and notebooks, and the notebooks themselves live in seven separate GitHub repositories: one each for neural networks, autoencoders, logistic regression, the perceptron, principal component analysis, K Means clustering and gradient descent. None of those notebooks are in this tree. A shell script at the root downloads the relevant notebooks from those repositories, and only then can the book be built. What lives here instead is the scaffolding: `_toc.yml` for the table of contents and structure, `_config.yml` for configuration, the chapter1 through chapter4 directories the notebooks land in, a book directory for combined output, gifs and static assets, an index page, two Dockerfiles and a compose file. The result is a build repository for a teaching site rather than a source repository for the teaching material itself.

## The site is rebuilt after every commit, from whatever the siblings hold

The automation is described in one sentence and it has consequences. The website is updated by the GitHub Action at .github/workflows/ci.yml after every commit or pull request. That action runs a download step and a build step, so the published site is a function of two repositories rather than one. If a sibling repository changes, the next build here picks up the change without any commit in this repository recording it, and if a sibling is renamed, renamed or made private, the build stops finding content. Nothing in the download step pins a commit, a tag or a release of the notebooks it fetches. The practical reading is that the site is a live composite: useful for following the material as it improves, unsuited to citing a specific version of a derivation, and impossible to rebuild byte for byte from a given commit of this repository alone. Anyone who wants a fixed copy has to save the built HTML or the notebooks themselves.

## Three build paths, and they do not run the same steps

The usage section offers a local CLI, Docker Compose and plain Docker, and the difference between them matters. The CLI route is two commands:

```sh
pip install -U jupyter-book
jupyter-book build .
```

That path builds whatever is already on disk, so the download step has to have happened first by hand:

```sh
chmod +x ./download_notebooks.sh
./download_notebooks.sh
```

The Compose route is the one that automates it. The compose file defines a service whose command is chmod on the script, then runs it, then runs jupyter-book build, with the working directory bind mounted into the container, so a single command produces the whole book:

```sh
docker compose run --rm jupyter-book
docker compose down --remove-orphans --volumes --rmi local
```

The plain Docker route is the compose route with the steps spelled out: build from Dockerfile.book, run with the directory mounted, and then stop, remove and remove the image by hand. The teardown commands are four separate steps because a container and an image are two things, which is a reminder that this path is the manual one.

## The EPUB path names its four chapters by hand

The print and ebook build is a separate pipeline, and it starts by concatenating notebooks listed explicitly:

```sh
brew install --cask mactex
nbmerge $(ls chapter1/*.ipynb chapter2/*.ipynb chapter3/*.ipynb chapter4/*.ipynb | sort) -o book/combined.ipynb
jupyter nbconvert --to latex book/combined.ipynb
```

The glob enumerates four chapter directories by name and sorts the result before merging. That is the whole fragility of this path in one line: if the book ever gains a chapter5, or renames one of these four, the EPUB silently loses it, because nothing is discovered automatically and nothing fails loudly. The repository tree does carry exactly chapter1 through chapter4 today, so the two are currently in agreement. The LaTeX output is then converted in a container with pandoc, using mathml output with embedded resources and standalone mode, which is what lets the mathematics render in readers that have no TeX installed.

## The pandoc service is pinned to amd64 while the TeX step is a Mac cask

The compose file defines two services, and their platform handling is inconsistent. The book service builds from Dockerfile.book and mounts the working directory. The pandoc service sets platform: linux/amd64 explicitly, which forces the amd64 image even on an ARM machine, where it then runs under emulation. Meanwhile the LaTeX step is installed with brew install --cask mactex, a macOS package manager. So the EPUB path is macOS first, in that the LaTeX toolchain is a Homebrew cask, and the conversion step is an emulated amd64 container, which is a combination that will be slower than it looks and will not work at all on Linux without replacing that cask. The pandoc invocation itself is otherwise carefully specified, mounting the directory at /data and writing the finished ebook into the book directory.

## Two upper bounds and six packages with none

The requirements file is short, and the shape of it is the point. Two entries carry upper bounds: jupyter-book is held below 2, and the pydata sphinx theme is held below 0.15.3. The rest carry no version constraint at all: matplotlib, numpy, nbconvert, nbformat, celluloid and scienceplots. Both bounded packages are the ones where a major or theme release would break a build rather than improve it, which is defensible, but it means the two halves of the file have different intentions and the unpinned half is where a surprise will come from. The unpinned pair also does real work: matplotlib draws the training curves the whole project exists to show, and scienceplots supplies the styles for them. celluloid is there for animation, which is consistent with the gifs directory in the tree.

## Two visualization toolkits, and an output section that is two headings

Two kinds of notebook sit behind the site, and they answer different questions. The Jupyter Book notebooks implement and mathematically derive each algorithm from first principles, and each one's output is a series of visualisations across the training phase until the weights converge. Alongside them are interactive notebooks built with Marimo, whose stated purpose is narrower: letting a reader see how the weights influence the loss functions directly rather than watching a fixed run. That split is a good fit for the subject, since a derivation is not something you can poke at and a loss surface is. What is missing is the signposting. The output section at the end of the README carries two subheadings, one for the Marimo notebooks and one for the mathematical explanations, and neither has any text or example beneath it, so the two outputs the project is proudest of are announced and then not shown.

## Conclusion

This project fits a reader who wants the derivation and the animation in the same page, and who is willing to read a book assembled from several repositories rather than one. It does not fit a curriculum you intend to pin, because the notebooks are fetched at build time and the site is rebuilt after every commit, so what a given commit produced depends on what the sibling repositories held at that moment. Before relying on it, check four things. Which chapter set you are reading, since the EPUB pipeline enumerates chapter1 through chapter4 explicitly and would silently skip a fifth. Which build path you use, because the CLI route and the Compose route do not run the same steps. What the dependency pins imply, since two packages carry upper bounds and six carry none. And whether the interactive notebooks and the mathematical explanations are reachable, since the output section of the README is two headings and nothing else.

## FAQ

### Where are the notebooks for machine-learning-visualized?

They are not in this repository. Each algorithm has its own GitHub repository: neural networks, autoencoders, logistic regression, the perceptron, principal component analysis, K Means clustering and gradient descent. A shell script in this repository downloads the relevant notebooks, and only then can the book be built.

### How do I build machine-learning-visualized locally?

Run `chmod +x ./download_notebooks.sh` and then `./download_notebooks.sh` to fetch the notebooks, then install jupyter-book and run `jupyter-book build .`, or use `docker compose run --rm jupyter-book`, which performs both steps in one command. The finished book is written to `_build/html/index.html`.

### How does the machine-learning-visualized website get updated?

A GitHub Action at .github/workflows/ci.yml rebuilds it after every commit or pull request. Because the notebooks are fetched from other repositories at build time and no commit or tag is pinned, the published site reflects whatever those repositories held when the build ran.

### Can I get machine-learning-visualized as an EPUB?

There is a documented EPUB path: install the mactex cask, merge the notebooks listed as chapter1 through chapter4 with nbmerge, convert to LaTeX with jupyter nbconvert, and convert that to EPUB with pandoc in a container. The chapter list is written out by hand, so a new chapter directory would not be picked up automatically.

## Sources

- [gavinkhung/machine-learning-visualized on GitHub](https://github.com/gavinkhung/machine-learning-visualized)
- [Issues](https://github.com/gavinkhung/machine-learning-visualized/issues)
- [License: MIT](https://github.com/gavinkhung/machine-learning-visualized/blob/main/LICENSE)
- [Project website](https://ml-visualized.com/)
- [README](https://github.com/gavinkhung/machine-learning-visualized/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/gavinkhung-machine-learning-visualized
