Self-hosted service
Yorko/mlcourse.ai avatar
Yorko/mlcourse.ai

mlcourse.ai: an open machine learning course you run from your own machine

Open Machine Learning Course

10,715 stars5,690 forksPythonNOASSERTION

At a glance

What is it?
Yorko/mlcourse.ai is a self-paced, ten-week machine learning curriculum by Yury Kashnitsky, shipped as Jupyter notebooks in five languages and built with Jupyter Book. It is strong on theory-to-practice balance and weak on anything resembling a maintained software product.
Who is it for?
Adopt mlcourse.ai if you want a structured ten-week path from Pandas to gradient boosting and you are willing to read math before code; the repository gives you the notebooks, the articles and the Kaggle Inclass assignments in one place. Do not adopt it if you need a maintained library, a stable API or an English-only pipeline, because the content is a course and the bonus assignments are copyrighted and sold separately.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 27 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 17, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What mlcourse.ai actually is, and who it is written for

This is not a library and not a framework. It is a course. The README describes mlcourse.ai as "an open Machine Learning course by OpenDataScience (ods.ai)", led by Yury Kashnitsky, who the README says holds a Ph.D. in applied math and reached Kaggle Competitions Master tier. The stated design goal is a balance between theory and practice: math formulae in the lectures, assignments and Kaggle Inclass competitions for the practice half.

The audience is narrow and specific. You are expected to already write Python and to tolerate derivations. The course runs ten weeks, from Pandas through Gradient Boosting, and the README says it is currently in self-paced mode, meaning there is no cohort, no schedule and no instructor feedback loop. If you want a live class with deadlines and grading, this is the wrong shape. If you want a fixed syllabus you can work through on your own schedule, the repository gives you the whole thing at once.

One structural detail matters more than it looks: the material exists in several parallel trees. The top level of the repository contains jupyter_english/, jupyter_russian/, jupyter_chinese/, jupyter_french/ and jupyter_arabic/, plus mlcourse_ai_jupyter_book/, which is the Jupyter Book build target. The README lists mirrors on the main site, a Kaggle Dataset of the same notebooks, and an Arabic version. That redundancy is deliberate, and it is also the source of most maintenance overhead.

How the course is assembled: notebooks, Jupyter Book, and a Makefile

The mechanism is simpler than the topic list suggests. Content lives as notebooks and Markdown inside mlcourse_ai_jupyter_book/, and Jupyter Book compiles that directory into a static site. The Makefile exposes four targets: install, build, check_links and deploy. The build target runs jb build against mlcourse_ai_jupyter_book, and the deploy target pushes the generated HTML with ghp-import under the custom domain mlcourse.ai.

So the data flow is: notebook sources in the repository, Jupyter Book as the renderer, a _build/html directory as the artifact, ghp-import as the publisher. There is no server component, no database and no API. The course site is a static output of the repo, which is why the README can point readers at both the site and the Kaggle mirrors without any synchronization layer between them.

Dependencies are declared in pyproject.toml under the project name mlcourse-ai, version 1.0.0, with requires-python set to >=3.13. The runtime list is long and pinned loosely with lower bounds: numpy>=2.4.1, pandas>=2.3.3, scikit-learn>=1.8.0, matplotlib>=3.10.8, seaborn>=0.13.2, plotly>=6.5.2, statsmodels>=0.14.6, xgboost>=3.1.3, prophet>=1.2.1, wordcloud>=1.9.5, and jupyter-book==1.0.4.post1, the one exact pin in the list. A separate dev group carries black, pylint, pre-commit, gdown and sphinx-jupyterbook-latex. The single exact pin on jupyter-book is a sensible choice, since Jupyter Book 1.x changed build behavior across releases and the deploy target depends on the output path staying stable.

Installing mlcourse.ai and building the book locally

The repository uses uv for dependency management, and the Makefile's install target is a one-line wrapper around it. You need Python 3.13 or newer, because pyproject.toml sets requires-python to >=3.13. From a clone of the repository, run the install target:

bash
make install

That executes uv sync, which resolves pyproject.toml and uv.lock into a local virtual environment. The lockfile is committed, so the resolved versions should match what the project itself builds against, even though most entries in pyproject.toml are lower bounds rather than pins.

To render the course as a static site, use the build target. It invokes Jupyter Book on the book directory:

bash
make build

The output lands in mlcourse_ai_jupyter_book/_build/html. Open the index page there in a browser and you should see the table of contents for the ten topics, from topic01 on Pandas through the later gradient boosting material. If a notebook fails to execute, the failure surfaces during this step, not later, which is the main reason to build locally before reading anything.

If you only want to read, the README points at three mirrors that require no local setup: the main site at mlcourse.ai, a Kaggle Dataset containing the same notebooks as Kaggle Notebooks, and the Arabic version under jupyter_arabic/. The local build is worth the effort only if you intend to modify notebooks or regenerate the site. There is also a link-checking target:

bash
make check_links

This runs the Jupyter Book linkcheck builder. Expect it to be slow and expect it to fail on external links that have rotted, because the course links out to Medium, Habr and Kaggle across five language trees.

The bonus assignments are the part that is not open

The README is unusually direct about this, and it is the most important licensing boundary in the project. The main course is described as remaining open and free, created in the ODS.ai community. The Bonus Assignments pack is separate: ten assignments, sold through a Patreon tier or a Boosty tier for a stated contribution of $17/month, containing the non-demo versions of the assignments, including beating a baseline in the Alice and Medium Kaggle competitions and implementing stochastic gradient descent and gradient boosting from scratch.

The README states plainly that unlike the rest of the course content, the Bonus Assignments are copyrighted, and that public sharing of the pack is prohibited, while informally the author is fine with sharing with two or three friends. Treat that as the author's stated position, not as a licence you can rely on. The repository's own LICENSE.md is reported as NOASSERTION by the hosting platform, while the README badge points at CC BY-NC-SA 4.0. Those two signals do not obviously agree, and anyone planning to reuse the notebooks in a commercial training product should read LICENSE.md directly rather than trusting the badge. This is a description of what the files say, not legal advice.

A secondary cost sits in the pricing mechanics the README describes: the first payment is charged at the moment you join the tier, and the next on the first day of the following month, so the README itself recommends purchasing in the first half of a month. That is a billing detail most course pages omit, and it is worth knowing before you click.

Where mlcourse.ai breaks down

The first limitation is that this is a course with a version number, not a maintained library. The only release listed is v1.0.0, tagged "Self-paced mlcourse.ai", from 2022-01-16. The last push to the repository was on 2026-09-03, so the repository is not abandoned, but a single release tag from 2022 tells you the content is not versioned in any way you can depend on. If you build a pipeline on top of these notebooks, a later push can change them without a version bump.

The second limitation is Python 3.13. The requires-python constraint is a hard floor, and the dependency set includes numpy 2.x, pandas 2.x and scikit-learn 1.8, which are major versions that changed behavior relative to the 0.x and 1.x lines most older tutorials assume. If your environment is pinned to an older stack, you will be resolving conflicts rather than learning machine learning.

The third is multilingual drift. Five language trees and a Kaggle mirror of the same notebooks mean corrections do not propagate automatically. The README links Chinese notebooks through nbviewer and Kaggle notebooks in English, and the Arabic version lives in its own directory. Nothing in the repository enforces that a fix in jupyter_english/ reaches jupyter_chinese/.

Finally, the wrong-tool case is clear. If you want a maintained implementation of an algorithm, use scikit-learn or xgboost directly. If you want a benchmark suite, this is not one. mlcourse.ai teaches, and the notebooks are pedagogical artifacts with narrative, plots and commentary around the code. Pulling a function out of a lecture notebook and shipping it is a misuse of the material, and the licence may not permit it anyway.

How it compares with a documentation-first resource like scikit-learn

The obvious alternative is not another course, it is the scikit-learn user guide and its example gallery. The difference in approach is structural. scikit-learn documents an API: each estimator has a reference page, a parameter list, and examples that are tested against the current release. mlcourse.ai documents a sequence of ideas: topic03 covers classification, decision trees and k nearest neighbors, topic05 covers bagging and random forests, and the value is in the ordering and the derivations, not in any single function signature.

That means the two solve different problems. If your question is "what does this parameter do", scikit-learn answers it and mlcourse.ai does not. If your question is "why does bagging reduce variance and when does it not", the course material is built for exactly that, and the reference documentation is not. The course also reaches into territory the library docs deliberately avoid: the README lists a Kaggle Inclass competition component and assignments that ask you to beat a baseline under guidance, which is a workflow question rather than an API question.

There is a practical overlap worth naming. The course depends on scikit-learn, xgboost, statsmodels and prophet as its runtime, so you end up reading both. The sensible split is to work the course for the sequence and the intuition, and keep the library documentation open in a second tab for the exact signatures, because the notebooks are pinned to library versions that will keep moving after the course text stops being updated.

Who should adopt mlcourse.ai, and what to verify first

Adopt it if you are a working developer or analyst who already knows Python and wants a ten-week structure that goes from Pandas and visual analysis through linear models, bagging, and gradient boosting, with math included rather than hidden. The self-paced mode suits people who cannot commit to a cohort schedule, and the five language trees plus the Kaggle mirrors mean you can read in Russian, Chinese, French or Arabic, or run the notebooks on Kaggle without installing anything.

Do not adopt it if you need a supported library, if you are restricted to Python 3.12 or older, or if you want the full assignment set without paying, since the ten bonus assignments are the monetized portion and the README describes them as copyrighted. Do not adopt it as a citation source for production decisions either; a course notebook is not a specification.

What to verify before you invest time: run make install and make build on your own machine and confirm the notebooks execute with your Python version, because requires-python is >=3.13 and the dependency list spans numpy 2.x and scikit-learn 1.8. Then read LICENSE.md itself rather than the README badge, since the platform reports the licence as NOASSERTION while the badge says CC BY-NC-SA 4.0, and the non-commercial clause is the part that decides whether you can reuse the material at work. If you only need the reading, skip the build entirely and start at mlcourse.ai.

Editorial conclusion

Adopt mlcourse.ai if you want a structured ten-week path from Pandas to gradient boosting and you are willing to read math before code; the repository gives you the notebooks, the articles and the Kaggle Inclass assignments in one place. Do not adopt it if you need a maintained library, a stable API or an English-only pipeline, because the content is a course and the bonus assignments are copyrighted and sold separately. Before you commit, run uv sync and uv run jb build mlcourse_ai_jupyter_book locally and confirm the notebooks execute against the pinned dependencies in pyproject.toml, which requires Python 3.13 or newer.

Frequently asked questions

Is there a 3-month AI ML course available?

mlcourse.ai is structured as ten weeks of guided study, which is roughly two and a half months, and the README describes it as being in self-paced mode. Because it is self-paced, you set the actual duration rather than following a fixed three-month schedule.

What is the AI & ML course?

In this repository it refers to mlcourse.ai, an open Machine Learning course by OpenDataScience led by Yury Kashnitsky. It covers ten topics, from exploratory data analysis with Pandas to gradient boosting, with lectures, assignments and Kaggle Inclass competitions.

Which AI and ML course is best?

The repository does not compare itself with other courses, so it offers no ranking. The README states the design goal instead: a balance between theory and practice, with math formulae in the lectures and practice through assignments and Kaggle Inclass competitions.

What is a good website for practicing machine learning?

The README points to three places for this course: the main site at mlcourse.ai, a Kaggle Dataset holding the same notebooks as Kaggle Notebooks, and the Arabic version under jupyter_arabic/. The Kaggle mirror is the one designed for running the notebooks rather than just reading them.

Official sources

  1. Issues
  2. Project website
  3. README
  4. Releases
  5. Yorko/mlcourse.ai on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/yorko-mlcourse-ai.svg)](https://hysenlabs.com/projects/yorko-mlcourse-ai)