Open-source project
girafe-ai/ml-course avatar
girafe-ai/ml-course

girafe-ai/ml-course: A University-Style Open Machine Learning Curriculum

Open Machine Learning course

3,571 stars1,309 forksJupyter NotebookMIT

At a glance

What is it?
girafe-ai/ml-course is an MIT-licensed collection of Jupyter notebooks, slides, and homework assignments covering the first semester of a machine learning course taught at MIPT. It progresses from Naive Bayes and linear models through gradient boosting and sequence-to-sequence deep learning, with three archived semester releases and an active Yandex ML Trainings branch.
Who is it for?
This course works best for students who want a structured, semester-long curriculum starting from classical models and progressing to deep learning, and who have at least working knowledge of Python and linear algebra. The requirement to set up Poetry and the full dependency stack, which includes PyTorch, XGBoost, and several NLP libraries, makes it unsuitable as a first introduction to Python.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 2 days ago.
What is it written in?
Mainly Jupyter Notebook, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What girafe-ai/ml-course Covers and Who It Is For

This repository holds the first semester of the girafe-ai machine learning program. The pyproject.toml identifies it as 'Machine learning course at MIPT,' confirming its origin as a university course at the Moscow Institute of Physics and Technology. The primary authors are Radoslav Neychev and Vladislav Goncharenko, with additional contributions from six other instructors listed in the README.

The intended audience is students who already have some programming background and want a structured introduction to machine learning that progresses from classical statistical models to modern deep learning. The course does not teach Python basics: it assumes fluency with Python and enough linear algebra to follow the recap session in week 0. The README links to a separate prerequisites.md file, which specifies what background is needed.

The course runs over approximately eleven weeks of active lecture content, covering Naive Bayes, k-nearest neighbors, linear regression, linear classification (logistic regression), support vector machines, principal component analysis, decision trees, ensemble methods, gradient boosting, and deep learning topics including backpropagation, dropout, batch normalization, embeddings, and sequence-to-sequence models.

Repository Layout: Twelve Week Folders and a Homework Directory

The top-level directory structure reflects the course schedule directly. Folders named week0_00 through week0_11 each correspond to a course week. Within each week folder, the README references PDF slides and lecture recordings hosted on YouTube. The homeworks/ directory holds graded assignments and at least one lab: the README table references assignments on kNN, linear regression, SVM kernels, and a decision tree implementation from scratch, as well as Lab01 on ML pipeline construction.

Three semester releases are tagged in the repository: Spring 2020, Fall 2020, and Spring 2021. The main branch reflects the most recently updated content. A separate branch, 23f_yandex_ml_trainings, holds materials from the Yandex ML Trainings 2023 program and is linked prominently at the top of the README.

All notebooks are in Jupyter format (.ipynb), which means running them requires a Jupyter server. The repository also includes a poetry.lock file and a pyproject.toml, so the intended workflow is to install the environment with Poetry before opening any notebooks.

Setting Up the Python Environment

The repository uses Poetry for dependency management. The pyproject.toml file shows the package name as 'ml-mipt' and lists a full stack of scientific computing libraries:

- scikit-learn for classical ML algorithms - pandas, numpy, scipy, and statsmodels for data handling and statistics - matplotlib and seaborn for visualization - torch and torchvision for deep learning - xgboost for gradient boosting - opencv-python for image processing

The NLP sections require nltk, gensim, spacy, pytorch-transformers, and torchtext. These are listed as optional extras in the pyproject.toml under the nlp group. The reinforcement learning extras include gym and graphviz. ipywidgets is also listed as a direct dependency, used in a week covering MNIST downloads via torchvision.

The README does not include explicit shell commands for cloning or installing the environment. It directs readers to prerequisites.md for setup guidance. Given the size of the dependency tree, first-time setup will require downloading several hundred megabytes of packages, including PyTorch, which varies in size by platform and CUDA version.

The repository also includes a .pre-commit-config.yaml, indicating that the development workflow uses pre-commit hooks to enforce code quality. Students contributing back to the repository need to install pre-commit separately. The setup.cfg and pyproject.toml together form the project configuration; the pyproject.toml requires Python 3.8 or later.

Lecture Recordings, Slides, and the Teaching Format

Each week in the README table links to at least one lecture video and one seminar video on YouTube, along with a PDF slide deck in the corresponding week folder. The recordings are from the 2021 and 2022 cohorts, with notes in the table when a session was not recorded. One seminar (week 9) was replaced by a backpropagation seminar due to an instructor illness. One class (week 7) was a review session rather than new content.

The course follows a lecture-plus-seminar structure, where lectures introduce theory and seminars focus on implementation exercises. Homework assignments appear in the schedule with deadlines given in AOE time (Anywhere on Earth), suggesting the course was designed to accommodate students in different time zones.

For additional reading, the README lists three books: the YSDA ML Book (available in Russian only), 'Probabilistic Machine Learning: An Introduction' by Kevin Murphy (with both an English and a Russian translation link), and the 'Deep Learning Book' by Goodfellow et al., with a note that Part I is strongly recommended.

What the Course Leaves Out and Where It Constrains You

The course covers the first semester only. Topics that typically appear in a second semester, such as transformers, large language models, computer vision architectures beyond the introduction, and reinforcement learning beyond what the gym extras provide, are not documented in the main branch README.

The course assumes access to lecture recordings, which are hosted on external YouTube links. If those links become unavailable, the textual materials (slides and notebooks) remain, but the video instruction disappears. The README does not document any offline-first distribution mechanism.

The homework assignments are structured for a graded course environment: deadlines are given and the format assumes a student who submits work for evaluation. There is no automated grader or answer key in the repository, so self-study learners working through the assignments alone have no mechanism for feedback beyond comparing their implementations to the provided code structure.

The course does not cover MLOps, deployment, or model serving. It ends at the model training and evaluation level. The exam program is available in the approximate_program.pdf file at the repository root, giving students a complete list of topics they are expected to know by the end of the semester. The extra_materials.md file links to additional reading beyond the three recommended textbooks.

fast.ai as an Alternative and the MIT License

fast.ai's Practical Deep Learning course takes the opposite teaching approach: it starts with working deep learning models and works backward to the underlying theory. girafe-ai/ml-course starts from statistical foundations and builds toward deep learning. Students who want to train a model in the first session and understand it gradually should look at fast.ai. Students who want to understand why a linear classifier works before touching a neural network will find the bottom-up structure here more aligned with that goal.

Both resources use Jupyter notebooks and are freely available. The primary structural difference is that fast.ai courses are designed to be completed by following a pre-recorded online format, while girafe-ai/ml-course is structured around a university semester with live seminars and timed homework assignments. fast.ai focuses on deep learning, whereas girafe-ai/ml-course spends its first several weeks on classical models before reaching neural networks.

The repository is licensed under MIT, which permits use, copying, modification, and redistribution without restriction beyond attribution. The last push was on 2026-09-22, and three named semester releases (Spring 2020, Fall 2020, Spring 2021) serve as stable snapshots.

Editorial conclusion

This course works best for students who want a structured, semester-long curriculum starting from classical models and progressing to deep learning, and who have at least working knowledge of Python and linear algebra. The requirement to set up Poetry and the full dependency stack, which includes PyTorch, XGBoost, and several NLP libraries, makes it unsuitable as a first introduction to Python. The course does not document a speaking or oral exam component, but the approximate_program.pdf covers the exam scope. The last push was on 2026-09-22, which means the main branch is being maintained.

Frequently asked questions

What Python libraries does girafe-ai/ml-course require?

The pyproject.toml lists scikit-learn, pandas, numpy, scipy, matplotlib, seaborn, PyTorch, torchvision, XGBoost, and OpenCV as core dependencies. NLP sections add nltk, gensim, spaCy, and pytorch-transformers as optional extras.

Does girafe-ai/ml-course include homework assignments?

Yes. The homeworks/ directory contains weekly assignments covering kNN, linear regression, SVM kernels, and a decision tree implementation, as well as at least one lab on ML pipeline construction. Deadlines are listed in the README schedule for each semester cohort.

Is there a version of girafe-ai/ml-course for Yandex ML Trainings?

Yes. The README links to a separate branch, 23f_yandex_ml_trainings, which holds materials from the 2023 Yandex ML Trainings program. It is distinct from the main semester course content.

Official sources

  1. girafe-ai/ml-course on GitHub
  2. Issues
  3. License: MIT
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/girafe-ai-ml-course.svg)](https://hysenlabs.com/projects/girafe-ai-ml-course)