# scikit-learn-mooc's notebooks are generated from python_scripts by a Makefile rule

> The source for an INRIA machine learning course, where the notebooks in the site are build output, the quizzes need a repository that is not in the clone, and the only tagged release is a session from 2022.

**INRIA/scikit-learn-mooc** — Machine learning in Python with scikit-learn MOOC

- Repository: https://github.com/INRIA/scikit-learn-mooc
- Website: https://inria.github.io/scikit-learn-mooc
- Stars: 1,423 · Forks: 604
- Language: Jupyter Notebook
- License: CC-BY-4.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/inria-scikit-learn-mooc

## The notebooks are a build artefact of the Python scripts

The notebooks directory is not the source. A Makefile pattern rule generates each one from a script of the same name:

```
$(NOTEBOOKS_DIR)/%.ipynb: $(PYTHON_SCRIPTS_DIR)/%.py
	python build_tools/convert-python-script-to-notebook.py $< $@
```

The list of files to build is itself derived at parse time by globbing python_scripts/*.py and rewriting the paths with two perl substitutions, so adding a script is enough to add a notebook.

Two more steps run before conversion. A matplotlibrc is copied out of python_scripts into the notebooks directory, which is where the figures' appearance is decided, and a sanity check script is run over both directories. So the published notebook, its figure styling and its build validation are one chain, and editing a notebook directly means your change is overwritten on the next build.

## Quiz generation needs a repository that is not in the clone

One target reaches outside the repository. The quizzes target passes two directories to a generator script:

```
quizzes:
	python build_tools/generate-quizzes.py $(GITLAB_REPO_JUPYTERBOOK_DIR) $(JUPYTER_BOOK_DIR)
```

The first of those is hard coded to a sibling checkout on another host:

```
GITLAB_REPO_JUPYTERBOOK_DIR = ../mooc-scikit-learn-coordination/jupyter-book
```

The Makefile comments explain the assumption, that the coordination repository and this one are siblings on the same machine, and that it should hold in most development setups. If it does not, you are told to override the variable by passing GITLAB_REPO_JUPYTERBOOK_DIR with make -e.

That is an honest piece of documentation and a real obstacle. A fresh clone cannot build the quizzes without also obtaining a second repository that is not referenced from the README.

## One exact pin, six packages with no bound at all

The dependency list is nine lines long:

```bash
scikit-learn==1.8
pandas >= 1
matplotlib>=3.10
seaborn >= 0.13
plotly
skrub
jupyterlab
notebook
IPython
```

Only scikit-learn is pinned exactly, at 1.8. plotly, skrub, jupyterlab, notebook and IPython carry no version constraint at all, which means a rebuild months from now resolves to whatever is current. The two data tools have floors but no ceilings, so a breaking release in pandas or matplotlib lands in the course notebooks rather than in a build you control.

That is a defensible choice for course material that wants current tooling, and it is a poor one for a course whose exercises teach specific library behaviour, since the code being taught can change underneath the teaching.

## Four dependency files and one tool nobody declares

The root holds environment.yml, environment-dev.yml, requirements-dev.txt and requirements.txt. Four places describe what to install, and which one applies depends on whether you want the runtime or the development environment and whether you prefer conda or pip.

There is also check_env.py at the root, which suggests the divergence between those declarations was noticed at some point and given a script.

Meanwhile one tool the build genuinely needs is missing from every list. The wrap-up target executes its Python files into notebooks:

```
	jupytext --execute --to notebook $(WRAP_UP_DIR)/*.py
```

jupytext appears in no dependency file in the repository root, so that target fails on a clean environment until you install it yourself.

## The only tagged release is a session from 2022

There is exactly one release, and it is not a version number. It is tagged session-3, named Third MOOC session, published on 2022-10-18.

The last push to main was on 2026-09-11. So the tag namespace is organised by teaching session rather than by release, and nothing after October 2022 has been tagged at all. That is fine for a course that evolves continuously and wants one snapshot per cohort, and it means you cannot pin a stable state of the material the way you would pin a library.

For citation the project uses a Zenodo archive instead of the tags, with a DOI of 10.5281/zenodo.7220306, plus a CITATION.cff file at the root. The DOI is the stable handle here, not the repository.

## The linter ignores import order, unused imports and line length

The project configuration is two tools and nothing else. Black is set to a line length of 79 with target versions from py38 through py311 and preview mode enabled. Ruff's lint section ignores four rules:

```
ignore = [
    'E402',  # module level import not at top of file
    'F401',  # imported but unused
    'E501',  # line too long
    'E203',  # whitespace before ':'
]
```

For notebook-shaped teaching code that is reasonable. Scripts that build a figure layer by layer need imports after code, unused imports appear when a cell is demonstrated and then simplified, and long lines come from readable plotting calls.

Two details are worth noting. There is no project table in this file, so it configures tooling rather than defining a package, and requirements.txt does that job. And Black's target list stops at py311 while scikit-learn is pinned to 1.8, so the formatting target and the library floor were last reviewed at different times.

## Two index files and a Binder link bound to main

There are two entry points into the material and no statement about which is current. full-index.ipynb is the Binder target and the local build source, and one-day-course-index.md sits beside it as a separate index for a compressed version of the course.

The online environment is a Binder link that resolves the main branch directly:

```
https://mybinder.org/v2/gh/INRIA/scikit-learn-mooc/main?filepath=full-index.ipynb
```

Because the branch is named in the URL rather than a commit, the same link gives you different notebooks over time. That is the right trade for a course that wants to be current, and the wrong one for reproducing an exercise that behaved a certain way.

The static build is a Jupyter Book site, with local setup deferred to a separate file rather than described in the README, which in 148 words does the job of naming the course, the host platform and the licence.

## Conclusion

This is a course repository rather than a library, and reading it as documentation has a specific consequence: the notebooks you see online are generated from Python scripts, so an error in a cell is usually an error in the script the cell came from, and a fix belongs upstream of the build. Expect the static site to be more reliable than the local build. Two things will stop you before that. Generating quizzes needs a second repository that is not part of this clone, and running the wrap-up targets needs jupytext, which no dependency file declares. If you only want the material, read the published site or the static build rather than trying to reproduce it locally.

## FAQ

### What is scikit-learn-mooc?

It is the source for the Machine learning in Python with scikit-learn MOOC, a free course hosted on the FUN-MOOC platform. The README asks you to enrol there for quiz solutions, executable notebooks and the discussion forum, and offers the static site for reading without enrolling.

### How do I run the scikit-learn-mooc notebooks locally?

The README defers to local-install-instructions.md at the repository root rather than repeating the steps. The notebooks are generated from the scripts in python_scripts, so a local build also needs the build tools and the Makefile targets rather than only a Python environment.

### Which scikit-learn version do the scikit-learn-mooc exercises use?

requirements.txt pins scikit-learn to exactly 1.8, which is the only exact pin in the file. pandas, matplotlib and seaborn carry floors without ceilings, and plotly, skrub, jupyterlab, notebook and IPython carry no constraint at all.

### Can I use the scikit-learn-mooc material in my own course?

The material is developed publicly under a CC-BY licence, with the full text in the LICENSE file at the repository root. Cite it through the project's Zenodo archive using the DOI it publishes, or through the CITATION.cff file in the root.

### Why is there no newer release of scikit-learn-mooc?

The repository has a single release, tagged session-3 and named Third MOOC session, published on 2022-10-18. The last push to main was on 2026-09-11, so the material continues to change without new tags.

## Sources

- [INRIA/scikit-learn-mooc on GitHub](https://github.com/INRIA/scikit-learn-mooc)
- [License: CC-BY-4.0](https://github.com/INRIA/scikit-learn-mooc/blob/main/LICENSE)
- [Project website](https://inria.github.io/scikit-learn-mooc)
- [README](https://github.com/INRIA/scikit-learn-mooc/blob/main/README.md)
- [Releases](https://github.com/INRIA/scikit-learn-mooc/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/inria-scikit-learn-mooc
