Open-source project
DataTalksClub/data-engineering-zoomcamp avatar
DataTalksClub/data-engineering-zoomcamp

Data Engineering Zoomcamp: seven modules, one pipeline, deadlines not classes

GitHub describes it as Data Engineering Zoomcamp is a free 9-week course on building production-ready data pipelines. The next cohort starts in January 2026. Join the course here 👇🏼. The repository metadata lists Jupyter Notebook as its primary language. This article stays within the project description and details documented in the GitHub repository README.

45,883 stars8,932 forksJupyter NotebookLicense varies

At a glance

What is it?
A free nine-week data engineering course whose live option means deadlines and graded homework rather than live teaching, with every lecture pre-recorded either way. The materials are Jupyter notebooks in numbered module folders, and the repository contradicts itself about which cohort the current files serve.
Who is it for?
Take the self-paced route if you want the notebooks and nothing else, and register for a live cohort only if the scored homework, peer review and certificate matter to you, since that is the entire difference between the two options. Before you start, read Module 1 and Module 5 to see how much of the pipeline you are expected to deploy to Google Cloud, and note that the repository carries no LICENSE file, so treat the material as reference rather than something to republish.
Can I use it commercially?
Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
Is it still maintained?
Yes. The repository last received commits 15 days ago.
What is it written in?
Mainly Jupyter Notebook, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

A live cohort means deadlines and a leaderboard, not live classes

The most quoted feature of this course is the one that does not exist. An important note in the README states that live cohort does not mean live classes, that all lectures are pre-recorded, and that live means working alongside others with deadlines, scored homework, a leaderboard, peer review and a certificate at the end. The comparison table makes the same split in six rows: lectures are pre-recorded on both tracks, homework is graded in the live cohort and available but not scored self-paced, and the leaderboard, peer review and certificate exist only for the live track. Both cost nothing. So the decision is not about scheduling or teaching, it is about whether external structure and a certificate are worth registering for, and a self-paced learner gets exactly the same notebooks.

The description says January 2026 and the table says January 2027

Two dates for the same cohort live in the same repository. The repository description states that the next cohort starts in January 2026, while the live cohort table gives the start as January 2027. Nothing in the README explains which is current, and the last push was on 2026-09-15 to the main branch, so the files on main are being edited either way. The practical effect is that a reader cannot tell whether the material on main targets the cohort that already ran or the one that has not started, and a self-paced learner working through the notebooks in order may hit changes written for a different group. Registration is the arbiter: the README routes you to courses.datatalks.club for the live cohort and tells the self-paced reader to just start learning.

Module folders run 01 to 07, and the ingestion workshop is filed under cohorts/2026

The learning material is a directory of numbered folders, one per module:

bash
01-docker-terraform/
02-workflow-orchestration/
03-data-warehouse/
04-analytics-engineering/
05-data-platforms/
06-batch/
07-streaming/

The exception is Workshop 1: Data Ingestion, which does not live at the top level. Its link points into cohorts/2026/workshops/dlt.md, so one required piece of the pipeline sits in a folder named for a single cohort year while the seven modules sit in version-neutral directories. That asymmetry is a maintenance liability in a repository with no releases: a future cohort either reuses the 2026 path or a link in the README goes dead, and nothing in the repository checks it. The final project, which the README says applies all the concepts in a real scenario with peer review, sits in its own projects/ directory alongside course.yaml and the per-cohort material.

Registration lives on courses.datatalks.club, not on the homepage GitHub shows

The repository metadata points somewhere different from the README. The homepage recorded for the repository is an Airtable share link, while the registration button in the README goes to courses.datatalks.club/register/de-zoomcamp/ and the course platform for deadlines and homework is courses.datatalks.club. If you arrive from the repository card rather than from the README, you land on a form that is not the one the project documents. The rest of the entry points are consistent and worth bookmarking as a set: the YouTube playlist holds the video lectures, Slack carries questions in #course-data-engineering, Telegram carries announcements, and a separate FAQ document and two documentation pages cover logistics and the course itself. Six instructors and six past instructors are listed, so the material has had rotating ownership across cohorts.

No LICENSE file sits in the tree, so the reuse terms are unstated

The top-level listing of the repository contains no LICENSE file, and the repository metadata reports no licence. The files are .gitignore, README.md, course.yaml, the seven module directories, cohorts/, projects/, scripts/, images/, and a set of loose documents including after-sign-up.md, certificates.md, learning-in-public.md, workshop-best-practices.md and awesome-data-engineering.md. The distinction matters because the course asks for something in return: the self-paced path ends with a project for your portfolio, and the live path adds peer review of that project. Learners are being asked to publish work built on these notebooks while the repository states no terms for the material itself. Read the certification and logistics pages before you assume either direction, and if you plan to teach from these files, ask first.

Module 1 opens on Google Cloud and Module 5 ends on deployment, with no cost stated

The syllabus starts on someone else's infrastructure. Module 1, Containerization and Infrastructure as Code, opens with an introduction to GCP, then covers Docker and Docker Compose, running PostgreSQL with Docker, and infrastructure setup with Terraform. Module 5, Data Platforms, ends with deployment to cloud on BigQuery, after Bruin for end-to-end pipelines and a session on ingestion, transformation and quality. The course is free, and the README says so in the table and in the heading, but it does not state what the cloud components cost, and BigQuery bills by query volume. A learner who finishes every module has therefore run a pipeline in Google Cloud. Check what a free tier covers for your region before Module 1, because a stuck Terraform apply or an idle BigQuery dataset is a cost, and the course materials are not the place that will warn you.

Seven modules, seven toolchains, and the pipeline is the only thing that carries over

Read the tool column and breadth is the plan. Module 2 uses Kestra for orchestration, the ingestion workshop uses dlt, Module 3 is BigQuery, Module 4 is dbt with DuckDB and BigQuery, Module 5 is Bruin, Module 6 is Apache Spark with DataFrames, SQL and the internals of GroupBy and joins, and Module 7 is Kafka, Kafka Streams, KSQL and Avro schema management. No tool appears in two modules, and the only thing carried from one to the next is the pipeline itself. That is a defensible design for a survey of a stack, and it has a cost: you will not arrive fluent in any of them, and each is one you then have to learn properly somewhere else. The prerequisites are honest about the entry point, asking for basic coding experience, familiarity with SQL, and Python experience described as helpful but not required, while stating that no prior data engineering experience is necessary.

Editorial conclusion

Take the self-paced route if you want the notebooks and nothing else, and register for a live cohort only if the scored homework, peer review and certificate matter to you, since that is the entire difference between the two options. Before you start, read Module 1 and Module 5 to see how much of the pipeline you are expected to deploy to Google Cloud, and note that the repository carries no LICENSE file, so treat the material as reference rather than something to republish.

Frequently asked questions

What is the Data Engineering Zoomcamp?

A free 9-week course on data engineering that has you build an end-to-end data pipeline from scratch, structured as modules, hands-on workshops and a final project. It is aimed at developers, analysts and data scientists, and the repository says no prior data engineering experience is necessary.

Is the Data Engineering Zoomcamp good?

All lectures are pre-recorded on both the live and self-paced tracks, so the difference is structure rather than teaching. The live cohort adds deadlines, scored homework, a leaderboard, peer review and a certificate, while the self-paced track gives you the same materials with homework available but not scored.

Is the Data Engineering Zoomcamp worth it?

It costs nothing on either track, and the self-paced steps are to follow the materials on GitHub, ask questions in Slack, and do self-checked homework plus a project for your portfolio. A certificate is awarded only to learners who complete the final project during a live cohort, and that final project includes a peer review and feedback process.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/datatalksclub-data-engineering-zoomcamp.svg)](https://hysenlabs.com/projects/datatalksclub-data-engineering-zoomcamp)