Self-hosted service
DataTalksClub/machine-learning-zoomcamp avatar
DataTalksClub/machine-learning-zoomcamp

Machine Learning Zoomcamp: a free 4-month ML engineering course you take through a GitHub repository

Learn ML engineering for free in 4 months! Register here 👇🏼

14,454 stars3,194 forksJupyter NotebookLicense varies

At a glance

What is it?
DataTalksClub's Machine Learning Zoomcamp is a free, cohort-based course that walks from regression and classification to FastAPI, Docker, Kubernetes and serverless deployment. It is aimed at people who can already program, and the certificate only exists inside the live cohort.
Who is it for?
Adopt Machine Learning Zoomcamp if you already program, can use a command line, and want graded homework plus a certificate that requires two qualifying projects and peer review inside the live cohort. Skip it if you have no programming experience, want a non-technical introduction to AI, or want advanced machine learning research topics, since the README names those as poor fits.
Can I use it commercially?
Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
Is it still maintained?
Yes. The repository last received commits 14 days ago.
What is it written in?
Mainly Jupyter Notebook, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 17, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gap Machine Learning Zoomcamp fills: notebooks that never reach a service

Most introductory machine learning material stops at a fitted model and a metric. The README describes the course as following the complete path from a machine learning problem to a production service, and spells that path out as framing the problem, preparing the data, training and evaluating the model, exposing it through an API, packaging it, and deploying it. That last half is the part that is usually missing, and it is the part the repository is organised around.

The intended audience is stated plainly: people with programming experience who want to move into machine learning engineering, plus data analysts, data engineers and technical students. The prerequisites are one year of programming, comfort on the command line, and a laptop with an internet connection. Prior machine learning experience is explicitly not required. The README also names who should stay away: people with no programming experience, people who want a non-technical introduction to AI, and people looking for advanced machine learning research topics. Those three exclusions are worth taking at face value rather than treating as modesty.

How the course is structured: numbered modules, homework, and a repository you clone

The top-level layout is the syllabus. Directories run from 01-intro and 02-regression through 03-classification, 04-evaluation, 05-deployment, 06-trees, 08-deep-learning, 09-serverless, 10-kubernetes and 11-kserve. The numbering makes the progression visible without opening a single file: modelling first, then evaluation, then the deployment chain that ends in Kubernetes and KServe. Note that 07 does not appear in the top-level listing, so anyone reading the directory names should not assume a contiguous sequence.

The tooling named in the README is Python, NumPy, pandas and scikit-learn for the modelling half, then TensorFlow and PyTorch for deep learning, then FastAPI, Docker, Kubernetes and AWS Lambda for serving. The repository is mostly Jupyter Notebook, which matches the teaching style, and it also carries course.yaml, a cohorts/ directory, docs/ and projects/ alongside the module folders.

Delivery is split into two tracks with different consequences. In the live cohort, homework is graded, there is a leaderboard, peer review, and certificate eligibility. Self-paced learners get the same pre-recorded lectures and the same homework, but the homework is not scored, there is no leaderboard, no peer review and no certificate. The README is careful about one thing that trips people up: live cohort does not mean mandatory live classes. Lectures are pre-recorded in both tracks. What the cohort adds is shared deadlines and the certificate path. It runs once a year, with modules from September through December and final projects concluding in January.

Setting up and completing your first module from the repository

There is no package to install. The course is consumed by cloning the materials repository and working through a module folder, so the first real step is getting the files locally. The command below is the standard clone for a repository whose default branch is main.

bash
git clone https://github.com/DataTalksClub/machine-learning-zoomcamp.git
cd machine-learning-zoomcamp

After that, list the top level to confirm the module layout matches what the README describes. You should see the numbered directories, plus cohorts/, docs/, projects/ and course.yaml.

bash
ls

For a first real use, open the introductory module and read its lesson notes in order. The README does not specify a filename inside 01-intro, so check the directory contents before assuming a naming convention.

bash
ls 01-intro

The lecture videos are not in the repository. The README points to a YouTube playlist and to the course platform at courses.datatalks.club, and registration for the 2026 cohort is a separate step at courses.datatalks.club/register/ml-zoomcamp/. Questions go to the DataTalks.Club Slack, with a dedicated course channel listed in the README's quick links table. If you intend to earn a certificate, register before working through the modules, because self-paced completion does not count toward it.

Where Machine Learning Zoomcamp is the wrong choice

The certificate is the sharpest limitation. It requires two qualifying projects and the required peer reviews, and both must happen during a live cohort. You may submit the midterm and one capstone, or both capstone projects, but a self-paced learner who finishes every module and every homework assignment still ends with no certificate and no graded feedback. If the credential is the reason you are considering the course, the self-paced track does not deliver it.

The second constraint is the annual schedule. Modules run from September through December and final projects conclude in January, and the cohort runs once a year. Someone who discovers the course in February and wants graded work has to wait for the next cycle or accept the unscored track.

Third, the deep learning modules use cloud resources because the README states that a powerful computer or local GPU is not required. That is convenient, but it means the heavier exercises depend on external infrastructure rather than your laptop, and the README does not document what happens if that access is unavailable to you.

Finally, the material is introductory by design. The README points people seeking advanced machine learning research topics elsewhere, so this is not the place to go deeper into model architecture or training theory. And on maintenance: the repository is not archived, and the last push was on 2026-09-10, which is recent enough that the materials are being touched, but that says nothing about the depth of any individual module.

How it compares with a self-directed scikit-learn and FastAPI path

The obvious alternative is assembling the same skills yourself: work through the scikit-learn user guide for modelling, then a FastAPI tutorial for serving, then the Docker documentation for packaging. That path is free, self-scheduling and unlimited in depth. The difference is not the content, it is the forcing function. Independently, nothing tells you that your evaluation metric was chosen badly, and nothing makes you finish the deployment step after the model already works.

Machine Learning Zoomcamp adds graded homework, a leaderboard, peer review and deadlines inside the live cohort, and it sequences the deployment chain for you: FastAPI, then Docker, then Kubernetes and KServe, then AWS Lambda for the serverless module. The trade-off is that you inherit someone else's order and pace. If you already deploy services for a living, the 05-deployment and 10-kubernetes modules will cover ground you know, and you cannot skip ahead to the certificate because two qualifying projects are still required.

A second alternative is a paid, cohort-based bootcamp. The README states the course is free in both tracks, so the comparison comes down to support and structure rather than cost. What you give up here is the smaller group and the direct instructor access that paid programmes typically sell. What you keep is the material, the Slack community and the YouTube lectures, all of which remain available whether or not you register.

Licence, maintenance and what upgrading actually costs

The repository does not state a licence in the README, and the GitHub metadata lists the licence as unknown. That matters if you plan to reuse the notebooks, the homework or the project scaffolding in your own teaching material or product. Without an explicit licence, you should not assume permission to redistribute. The README does welcome pull requests, which suggests contributions are expected, but a contribution policy is not a licence grant. If reuse matters to you, treat the absence of a licence as a question to resolve with the maintainers rather than as an implicit yes.

Upgrade cost is low in the conventional sense. There is no version to pin and no dependency manifest to migrate, because you clone the repository and read it. The cost that does exist is schedule drift: the cohort runs once a year, and the materials move with it. If you start in one cycle and finish in the next, the module numbering, the homework and the deadlines you were working against may have changed. The repository carries a cohorts/ directory precisely because each year has its own syllabus, so check which cohort folder your notes came from before assuming the deadlines still apply.

One practical consequence: because the live cohort is what produces the certificate, and the cohort is annual, a missed submission window is not something you can patch by working harder in the self-paced track. The work is recoverable, the credential is not.

Editorial conclusion

Adopt Machine Learning Zoomcamp if you already program, can use a command line, and want graded homework plus a certificate that requires two qualifying projects and peer review inside the live cohort. Skip it if you have no programming experience, want a non-technical introduction to AI, or want advanced machine learning research topics, since the README names those as poor fits. Before committing, verify the 2026 syllabus and deadlines under cohorts/2026/ on the main branch, and confirm on the course platform that registration for the September 14, 2026 cohort is still open, because certificate eligibility and the leaderboard exist only in the live cohort and not in the self-paced track.

Frequently asked questions

Can I learn machine learning in 3 months?

Machine Learning Zoomcamp is described as a 4-month course, with modules running from September through December and final projects concluding in January. The README does not offer a 3-month variant, so the stated schedule is longer than that.

Can I learn machine learning for free?

Yes. The README lists the cost as free in both the live cohort and the self-paced track. The lectures are pre-recorded and the course materials are in a public GitHub repository.

What are the 7 types of machine learning?

The README does not enumerate seven types of machine learning, so this question cannot be answered from it. What the README does describe is the course's own progression: regression and classification models, tree-based models, deep learning, and then deployment.

Which platform is best for machine learning?

The README does not rank platforms. It names the tools the course uses: Python, NumPy, pandas, scikit-learn, TensorFlow, PyTorch, FastAPI, Docker, Kubernetes and AWS Lambda, with cloud resources used for the deep learning modules.

Official sources

  1. DataTalksClub/machine-learning-zoomcamp on GitHub
  2. Issues
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/datatalksclub-machine-learning-zoomcamp.svg)](https://hysenlabs.com/projects/datatalksclub-machine-learning-zoomcamp)