Open-source project
DataExpert-io/data-engineer-handbook avatar
DataExpert-io/data-engineer-handbook

The Data Engineering Handbook is a link list, and one signup URL is malformed

GitHub describes it as This is a repo with links to everything you'd ever want to learn about data engineering. The repository metadata lists Jupyter Notebook as its primary language. This article stays within the project description and details documented in the GitHub repository README.

44,276 stars9,337 forksJupyter NotebookLicense varies

At a glance

What is it?
A repository whose value is curation rather than instruction: a README of companies, books, communities, blogs and whitepapers, plus separate Markdown files for projects, interviews, newsletters and data cleaning, and folders for two boot camps. Nothing is versioned, dated or licensed, so the reader inherits the maintenance burden of every link.
Who is it for?
Use the handbook as a map of the field, then commit to one tool per category and read that tool's own documentation, because the entries are names and links with no version, no cost and no maturity signal. Check the two boot camp folders before following the cohort links, since the dated signup pages are the part most likely to have moved, and note that the root carries no LICENSE file, so republishing the curated list inside a company wiki has no stated basis.
Can I use it commercially?
Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
Is it still maintained?
Yes. The repository last received commits 58 days ago.
What is it written in?
Mainly Jupyter Notebook, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The getting-started path is a 2024 roadmap, two boot camps and a broken signup URL

Three different onramds sit in the opening section and they age at different speeds. A newcomer is sent to a roadmap whose title carries the year 2024, hosted on a third-party blog. The two boot camps are separate: a 4-week free beginner camp with an introduction and a software list, and a 6-week free intermediate community camp with the same pair of documents. The third route is a Databricks AI boot camp advertised with a date of August 3rd, next to a Databricks Free Edition signup. That signup link is where the section stops being usable. Its query string reads provider=DB_FREE_TIERutm_source=github&utm_medium=video&utm_campaign=DataExpert, with the campaign parameters glued directly onto the value of provider and no separator. The link still lands, but the attribution parameters are part of the provider value, so the campaign tracking is broken. Nobody noticed because the repository is not archived and its last push was 2026-08-03. That date is worth pausing on as well, since a boot camp advertised for a specific August belongs to one cohort, and a reader arriving months later has no way to tell from the repository whether the cohort is still running or whether the material was updated for a new one. The three routes also disagree on cost: two are described as free boot camps with documents in the repository, while the third is a signup for a vendor's free edition. Nothing in the repository reconciles them.

Cube is filed under Data Integration and again under Semantic Layers

The company directory is grouped into categories, and one project occupies two of them. Cube appears in the Data Integration group alongside Fivetran, Airbyte, dlt, Sling, Meltano, Estuary and Arpe.io, and it appears again at the head of the Semantic Layers group next to dbt Semantic Layer. Everything else in the list appears once. Two things follow. The first is that the categories are not a partition, so a reader scanning a single group cannot assume it holds everything relevant, and the overlap is a real product statement rather than a mistake: a semantic layer is treated both as an integration surface and as its own category. The second is that duplicate entries are never marked, so nobody maintains a rule about them. A curated list without an inclusion rule drifts, and the first visible symptom of drift is a project filed twice with no comment explaining why.

Data Lake and Cloud holds three table formats, three startups and Microsoft

The Data Lake/Cloud category contains items at three different levels of the stack. Delta Lake, DuckLake, Apache Iceberg and Apache Polaris are open table and catalog formats, which you adopt by writing configuration. Onehouse, Ilum and Lakekeeper are commercial products built around that layer. Tabular and Microsoft are companies, and Microsoft is a company the size of the rest of the list combined. Then the category stops and a Data Warehouse group begins with Snowflake, Firebolt and Databend. The consequence for a reader is that the heading does not tell you what kind of thing you are linking to, and choosing by category alone can put you in a product page when you needed a file format spec. The same flatness shows up in Analytics/Visualization, which holds eleven entries from Looker Studio and Tableau down to Evidence and Redash, with no indication of which are hosted and which you run yourself.

Two of the three must-read books are not about data engineering tooling

The books file is described as a list of over 25 titles, and the README nominates three as must-read. Fundamentals of Data Engineering is the obvious one. Designing Data-Intensive Applications is a general book about data systems, storage engines and distributed design, which is why it appears on most engineering reading lists regardless of job title. Designing Machine Learning Systems is about model deployment and monitoring, which is a machine learning concern rather than a pipeline concern. Read the three as a sequence, and it works: the first gives you the discipline, the second the storage and consistency vocabulary, the third the part of the stack that sits after your transforms. Read them as tooling references and none of the three will tell you how to configure anything. The tooling names are elsewhere in the README, in the company categories, with no book attached to any of them. A reader who wants to compare orchestrators has eight names in one group and no criteria, and a reader who wants to learn one has to already know which of the twenty-five books to open. The handbook optimises for the moment before that decision, when you want to see the whole field at once, and it is honest about being a starting point rather than a path.

Nothing in the list carries a date, from a 2011 BI paper to current lakehouse work

The whitepapers section opens with A Five-Layered Business Intelligence Architecture, published in 2011, and then moves to a Lakehouse paper on unifying data warehousing with advanced analytics. That is a fifteen-year span inside one section, and neither entry is dated in the text. The blog list has the same problem, mixing Netflix TechBlog, Uber, Databricks, Airbnb, the AWS big data blog, Microsoft Data Architecture, Microsoft Fabric, Oracle, Meta, Onehouse and Estuary as a flat set of names, with no indication of which ones are still publishing. A link list with no dates is not wrong, but it transfers a specific cost to the reader: every recommendation has to be re-validated against the present, and there is nothing in the repository that tells you when that last happened. The last push, 2026-08-03, is the only timestamp available and it reflects edits, not verification.

A top-level directory is named read_this_for_application_fundamentals_for_python

The root of the repository contains .gitignore, README.md, six Markdown files, four folders and one entry that is a sentence with no file extension: read_this_for_application_fundamentals_for_python. It is a directory name written as an instruction to the reader, which tells you how the repository is used. This is a place where people dump what they think is worth keeping, not a structured publication. The other names follow the same logic in a milder form. books.md, communities.md, projects.md, interviews.md, newsletters.md and data_cleaning.md are all topic files rather than chapters, and beginner-bootcamp/, intermediate-bootcamp/ and databricks-ai-bootcamp/ are cohort folders. Sorting the root alphabetically puts the sentence first among the entries after README.md, which is a small reminder that the layout was designed for a person browsing, not for anything automated.

No LICENSE file in the root, while every link points at someone else's work

The repository metadata reports no licence, and the top-level listing contains no LICENSE file. What the project actually contains is other people's property: book listings, vendor homepages, community invitations, company engineering blogs and two whitepapers. The distinction matters more here than in a code repository, because the value is the selection. Publishing this README inside a company wiki or a course as your own curated reading list has no stated basis in the repository, even though the material is only links. Read the terms yourself before you redistribute it, and note that the same caution applies one level down: a link to a paid book is an affiliate-free reference today and a link that may stop resolving tomorrow, with nothing in the repository recording when it last worked. There is no link checker in the top-level entries and no continuous integration configuration to hold one.

Editorial conclusion

Use the handbook as a map of the field, then commit to one tool per category and read that tool's own documentation, because the entries are names and links with no version, no cost and no maturity signal. Check the two boot camp folders before following the cohort links, since the dated signup pages are the part most likely to have moved, and note that the root carries no LICENSE file, so republishing the curated list inside a company wiki has no stated basis.

Frequently asked questions

What is the Data Engineering Handbook?

It is a curated repository of links for learning data engineering, described as everything you would want to learn about the field. The README groups companies by category and links books, communities, company engineering blogs and whitepapers, with separate files for projects, interviews, newsletters, data cleaning, books and communities.

Which orchestration tools does the Data Engineering Handbook list?

The Orchestration group holds Mage, Astronomer, Prefect, Dagster, Airflow, Kestra, Shipyard and Hamilton, with no further guidance on choosing between them. Other groups cover data lake and cloud, data warehouse, data quality, education, analytics and visualization, data integration, semantic layers, modern OLAP, LLM application libraries, real-time data and data lineage.

Does the Data Engineering Handbook teach the tools or only link to them?

It links. There is nothing to install and no code to run, since the repository is Markdown files and folders, including the book, community, project, interview, newsletter and data cleaning files. The closest thing to instruction is in the boot camp folders, which hold an introduction and a software list for the beginner and intermediate cohorts.

What is in the boot camp folders of the Data Engineering Handbook?

The README links beginner-bootcamp/introduction.md and beginner-bootcamp/software.md for the 4-week free beginner boot camp, and intermediate-bootcamp/introduction.md and intermediate-bootcamp/software.md for the 6-week free intermediate community boot camp. A separate databricks-ai-bootcamp folder exists, and the README advertises a Databricks AI boot camp alongside a Databricks Free Edition signup.

Official sources

  1. Official README
  2. Project repository