Open-source project
jakevdp/PythonDataScienceHandbook avatar
jakevdp/PythonDataScienceHandbook

The handbook's code is MIT and its prose is non-commercial with no derivatives

GitHub describes it as Python Data Science Handbook: full text in Jupyter Notebooks. The repository metadata lists Jupyter Notebook as its primary language. The metadata lists the MIT license. This article stays within the project description and details documented in the GitHub repository README.

50,066 stars19,137 forksJupyter NotebookMIT

At a glance

What is it?
The full text of a book sold by a major technical publisher, in free Jupyter notebooks, with two separate licence files. The environment pins NumPy 1.11 and pandas 0.18 from 2016 on Python 3.5, the readme also claims Python 2.7 compatibility, and the tree carries three separate notebook directories with no explanation of which is current.
Who is it for?
Read the Python Data Science Handbook for the explanations, because the prose on NumPy broadcasting, pandas reshaping and matplotlib internals is still clearer than most current material, and the online version is free with no account. Do not treat the repository as a runnable environment: the pins are from 2016, the interpreter is Python 3.5, and the readme itself warns the exact versions may not be available on your platform.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Probably not. The repository last received commits 27 months ago, on June 26, 2024.
What is it written in?
Mainly Jupyter Notebook, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Two licence files, and the prose is the restrictive one

The licensing section is split in two, and the split is the most consequential thing in the repository. The code, including every code sample in the notebooks, is released under the MIT licence, with its own licence file. The text of the book is released separately under a Creative Commons licence that is attribution plus non-commercial plus no-derivatives, with its own licence file again. The consequence is an asymmetry that catches people out. You may use, modify and sell the code, including inside a commercial product, with no obligation beyond the MIT notice. You may not use the prose commercially, you may not translate it, and you may not produce a derivative work from it. A notebook that combines both is exactly the case where you need to think, because the code half is permissive and the text half is not, and a derivative notebook is by definition a derivative work. If your plan is an internal course, a translated edition, or a commercial training package, the text licence is the constraint, not the code licence.

The requirements pin NumPy 1.11 and pandas 0.18, from 2016

The environment file is a museum piece, and the readme tells you to install it directly. It pins the numerical array library at 1.11.1, the dataframe library at 0.18.1, the statistics library and the machine learning library both at 0.17.1, the image library at 0.12.3, the imaging library at 3.4.2, the plotting library at 1.5.1 and the statistical plotting library at 0.7.0, followed by a list of unpinned tooling. The two documented commands are a plain install from that file and the creation of a named environment with Python 3.5.

bash
conda install --file requirements.txt
bash
conda create -n PDSH python=3.5 --file requirements.txt

The readme does warn that some of these exact versions may not be available on your platform and that you may have to tweak them. It should be taken at its word: pandas 0.18 and that array library version predate several breaking changes, so a notebook written against them will raise errors on a current install rather than merely warn.

It claims Python 2.7 compatibility in two separate places

The About section and the Software section both tell you the same thing, and what they tell you is historical. The About section says the book was written and tested with Python 3.5, though other versions including Python 2.7 should work in nearly all cases. The Software section repeats it with a hedge, saying most but not all of the code will also work correctly with Python 2.7 and other older versions. Neither section qualifies the claim with a date. The consequence is that a reader takes away a compatibility statement about an interpreter that stopped receiving security fixes years ago, and treats it as current information. It also sets an expectation about the notebooks that a modern reader will not meet, because the hedge is doing a lot of work: most but not all, and nearly all cases. For a book whose stated purpose is teaching, the older-interpreter support was a genuine kindness to readers at the time and is now just a sentence that has to be un-learned.

Three notebook directories, and the readme does not say which is current

The top level of the repository contains three directories of notebooks, one plainly named for the notebooks and two carrying version suffixes. The readme points only at the first, telling you to run the code using the notebooks in that directory, and it points at an index notebook as the table of contents for the whole book. Nothing in the readme says what the other two directories are, whether they are older drafts, or whether anything still refers to them. Alongside them sit two environment specifications, a conda environment file and the requirements file the readme quotes, which are two separate ways of describing the same stack and can drift apart without either being wrong. The consequence is a small navigational tax with a real failure mode. A reader who opens the wrong directory gets notebooks that look current and behave differently, and because the duplication is silent there is no signal that you have picked the stale copy. Check the path in the readme before you start, and keep your own environment file rather than trusting either.

The entire text of a commercially sold book is in the repository

The opening line states the arrangement plainly. The repository contains the entire handbook, in the form of free Jupyter notebooks, and a separate line further down invites you to buy the printed book through the technical publisher. That is a deliberate and reasonably common arrangement for an author who also sells a print edition, and it is why the two-licence split described earlier exists at all: the author keeps the prose under terms that permit reading and redistribution but not commercial exploitation, and puts the code under a permissive licence so that readers can do more with it. The consequence for a reader is simply that there is no reason to buy the book to read it, and a reason to read the licence before you reuse anything. The corollary is that the licence boundary is load-bearing. If the text were under the same permissive terms as the code, a derivative work would be unremarkable; because it is not, the same derivative work needs permission.

It assumes you already know the language, and says so

The About section sets the prerequisite in one sentence: familiarity with Python as a language is assumed. For anyone who needs a quick introduction to the language itself, it points to a separate free companion project, described as a fast-paced introduction to Python aimed at researchers and scientists. The libraries the book covers are named as the core set essential for working with data in Python, namely the interactive shell, the array library, the dataframe library, the plotting library, the machine learning library and related packages. The consequence is that this is a second book, not a first one, and the honest signal for a reader arriving with no Python experience is the companion project rather than an attempt to push through chapter one. That framing also dates the book in a useful way. It was written when that library set was the whole of Python data science; the set has since grown by a large factor, and the book's value now sits in the explanations of how those particular libraries work underneath rather than in the completeness of the survey.

The last push was 2024-06-26 and there are no releases

The repository's activity record is worth stating because it changes how you should treat every version claim in it. The last push to the master branch was on 2024-06-26, and the project has no GitHub releases, so there is no tag to pin and no changelog to read. That is consistent with what the repository is: a finished book, not software with a release cycle, and the right expectation is that it will not be revised against current library versions. The consequence is twofold. Nothing has validated the notebooks against a current scientific Python stack in more than two years, so any success you get from running them comes from your own environment work rather than from maintenance. And the compatibility statements in the readme, which are already about Python 3.5 and 2.7, describe a period rather than a supported configuration. Read it as a fixed artefact, and choose a different book if you need one that tracks the current libraries.

Editorial conclusion

Read the Python Data Science Handbook for the explanations, because the prose on NumPy broadcasting, pandas reshaping and matplotlib internals is still clearer than most current material, and the online version is free with no account. Do not treat the repository as a runnable environment: the pins are from 2016, the interpreter is Python 3.5, and the readme itself warns the exact versions may not be available on your platform. Two licensing points to settle before you build anything on it. The code is MIT and commercially usable, while the text is non-commercial and no-derivatives, so you may not translate it, adapt it into a derivative work, or sell material derived from the prose. And if you intend to run the notebooks, build your own environment from current libraries and expect to fix code, rather than following the pinned file.

Frequently asked questions

Is the Python Data Science Handbook free?

Yes. The repository contains the entire book as free Jupyter notebooks, and you can also read it online at the project's own site. The readme separately invites you to buy the printed edition through the technical publisher. The code is MIT licensed; the text is under a Creative Commons attribution, non-commercial, no-derivatives licence.

Is the Python Data Science Handbook for beginners?

Not at the language level. The readme states that familiarity with Python as a language is assumed, and points anyone needing a quick introduction to a separate free companion project aimed at researchers and scientists. It is a book about the scientific libraries for someone who can already write Python.

Is there a PDF version of the data science Python book available?

The repository ships the book as Jupyter notebooks rather than a PDF, and the readme points at a notebooks directory and at an index notebook serving as the table of contents. It also links to hosted notebook services so the examples can be run without installing anything, and to the printed edition sold through the technical publisher.

What is the best book for learning Python for data science?

This handbook covers the interactive shell, the array library, the dataframe library, the plotting library and the machine learning library, and its readme links to a separate free project as a language introduction. The trade is that it was written and tested against Python 3.5 with library versions from 2016, so it explains those libraries well and tracks nothing newer.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/jakevdp-pythondatasciencehandbook.svg)](https://hysenlabs.com/projects/jakevdp-pythondatasciencehandbook)