# dataflowr notebooks are ordered by teaching session, not by module number

> The Jupyter material for a deep learning course taught at école polytechnique in 2023, where the module numbers jump around between sessions, several exercises ship as empty files to fill in, and the tree has grown directories the course schedule never mentions.

**dataflowr/notebooks** — code for deep learning courses

- Repository: https://github.com/dataflowr/notebooks
- Website: https://dataflowr.github.io/website/
- Stars: 1,270 · Forks: 333
- Language: Jupyter Notebook
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/dataflowr-notebooks

## The schedule jumps from module 6 straight to module 15

The course is presented as a sequence of sessions rather than a numbered progression, and the module numbers inside those sessions do not run in order. Session one is finetuning VGG, with module 1 covering the general overview and a dogs and cats notebook. Session two is PyTorch tensors and automatic differentiation, modules 2a and 2b. Session three then runs modules 3, 4, 5, and 6, and jumps to module 15 for dropout and MC dropout before session four starts again at module 7 with dataloading. Session four goes 7, 8a, 8b, 8c, then 16 for batchnorm and 17 for resnets. Session five covers autoencoders and GANs, session six recurrent networks, session seven attention and transformers. Anyone navigating by directory name will notice the same gaps, since the tree carries Module1 through Module19 with holes where the numbering skipped.

## Four exercises ship empty, and filling them in is the point

Several notebooks are deliberately incomplete, and the naming marks them. The collaborative filtering exercise is 08_collaborative_filtering_empty.ipynb, with a companion on a larger dataset for when the small one is finished. Word2vec has 08_Word2vec_pytorch_empty.ipynb. The ResNet session includes ODIN_mobilenet_empty.ipynb, whose stated purpose is to transform a classifier into an out-of-distribution detector. The recurrent network session has 11_predictions_RNN_empty.ipynb for predicting engine failure. The pattern is consistent: the repository ships the scaffolding and the surrounding code, and the part that demonstrates the concept is left blank. That makes the material usable as a course and much less useful as a reference, since there is no finished version of these four exercises to compare against. Other notebooks in the same directories are complete, so the distinction has to be made file by file.

## Automatic differentiation gets a second implementation in Julia

The autodiff session is the one place where the material leaves Python. Alongside the PyTorch modules there is a notebook that takes another look at automatic differentiation using dual numbers and Julia, hosted in the Module2 directory. The stated reason for looking twice is a correction rather than a repetition: automatic differentiation is not only the chain rule, since backpropagation and dual numbers are each described as clever algorithms for implementing it. That framing matters for the rest of the course, because later sessions use gradients without re-deriving them. The reminder notes in the same session include the observation that PyTorch tensors are NumPy on GPU plus gradients, and that broadcasting follows the same rules as NumPy and is used everywhere in deep learning.

## Dropout is taught before embeddings, batchnorm after them

The ordering encodes an argument about what a student should hold in their head at each point. Regularization and uncertainty estimation with MC dropout arrive in session three, attached to the convolutional network material, which is earlier than the embedding and dataloading sessions that follow. Batchnorm is not taught there either; it lands in session four alongside resnets, immediately before the out-of-distribution exercise that uses a trained MobileNet. The reminders treat dropout and batchnorm as separate items to know rather than as one regularization topic, and one of the session four lines is about architectures with skip connections allowing deeper models, which is the connection between the resnet module and what follows it. Module 15 and module 16 having their own directories next to Module5 and Module6 is the same fact visible in the file listing.

## The word2vec exercise turns an unsupervised task into a supervised one

The notes attached to the embedding session contain the clearest statement of intent in the repository. Word embeddings start in an unsupervised setting, and the exercise reframes that as a supervised task: predicting central and context words within a window, then learning the representation through negative sampling. That description is a compressed account of the whole derivation, and it is placed next to the collaborative filtering material because both use embeddings, with the reminder that categorical variables in deep learning are handled with embeddings. The repository also carries a directory for graph material and another for large language model material, neither of which appears in the session schedule shown here, so the tree is broader than the 2023 course it documents.

## One notebook exists to correct the official PyTorch attention tutorial

The attention session pairs module 12 with a specific erratum. The notebook 12_seq2seq_attention.ipynb is described as correcting the PyTorch tutorial on attention in seq2seq, which is a stronger statement than supplementing it. For a course that expects students to read vendor tutorials alongside the slides, that distinction matters, because a student following the official version and the course version would otherwise get different behaviour from code that looks similar. The same pattern shows up elsewhere in the sessions: a dogs and cats notebook for the finetuning session with practicals that add more dogs and cats, an MLP from scratch exercise that opens the first homework, and the start of the second homework on class activation maps and adversarial examples placed at the end of session four.

## The dependency list is flat and unpinned, and the tree outgrew the course

requirements.txt names fourteen packages with no version constraints: jupyter, numpy, pandas, requests, opencv-python, scikit-image, seaborn, matplotlib, tqdm, annoy, torch, torchvision, and scikit-learn. Nothing pins a version, so a fresh environment resolves to whatever is current, which for a teaching repository is a reasonable trade and for reproducing a 2023 result is not. Two details in the dependency choice are informative about the exercises rather than about packaging: opencv-python and scikit-image point at image work, and annoy is an approximate nearest neighbour library, which is what the collaborative filtering exercise on a larger dataset would need. The homeworks sit in four directories, HW1 through HW4, and the last commit on master is dated 2026-05-29 with no tagged releases, so there is no version to pin the notebooks to.

## Conclusion

These notebooks fit a learner who wants the exercises of a specific introductory deep learning course rather than a reference implementation, because the interesting files are the unfinished ones and the notes assume someone teaching in person. They do not fit someone looking for maintained library code, since the schedule describes a 2023 course and the requirements file pins nothing. Before starting, check which of the four homeworks you want to attempt, since each begins inside a session rather than standing alone, and read the per-session reminders, because several of them flag concepts the notebooks use before they are explained.

## FAQ

### What does dataflowr/notebooks contain?

Notebooks and code for the dataflowr deep learning course, following the schedule taught at école polytechnique in 2023. The tree holds one directory per module from Module1 to Module19 and four homework directories, HW1 through HW4, alongside newer material for graphs, large language models, and NeRF.

### Which dataflowr notebooks are meant to be filled in by the student?

Four are shipped deliberately incomplete: 08_collaborative_filtering_empty.ipynb, 08_Word2vec_pytorch_empty.ipynb, ODIN_mobilenet_empty.ipynb, and 11_predictions_RNN_empty.ipynb. The collaborative filtering one also has a larger dataset companion to move on to afterwards.

### What packages do the dataflowr notebooks require?

requirements.txt lists jupyter, numpy, pandas, requests, opencv-python, scikit-image, seaborn, matplotlib, tqdm, annoy, torch, torchvision, and scikit-learn, with no version pins on any of them.

### How does the dataflowr course present word2vec?

Module 8c covers it, and the session notes describe starting from an unsupervised setting and turning it into a supervised task of predicting central and context words in a window, learning the representation through negative sampling. The exercise notebook is 08_Word2vec_pytorch_empty.ipynb.

### Does dataflowr/notebooks use anything other than Python?

Most of it is Python notebooks built on PyTorch. One exception is the Module2 notebook on automatic differentiation with dual numbers and Julia, which takes a second look at autodiff alongside the PyTorch modules.

## Sources

- [dataflowr/notebooks on GitHub](https://github.com/dataflowr/notebooks)
- [Issues](https://github.com/dataflowr/notebooks/issues)
- [License: Apache-2.0](https://github.com/dataflowr/notebooks/blob/master/LICENSE)
- [Project website](https://dataflowr.github.io/website/)
- [README](https://github.com/dataflowr/notebooks/blob/master/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/dataflowr-notebooks
