Open-source project
gedeck/practical-statistics-for-data-scientists avatar
gedeck/practical-statistics-for-data-scientists

gedeck/practical-statistics-for-data-scientists: the code repository behind the O'Reilly book, and why it is now frozen

Code repository for O'Reilly book

3,392 stars1,994 forksJupyter NotebookGPL-3.0

At a glance

What is it?
The repository holds the R and Python notebooks for Practical Statistics for Data Scientists, 2nd edition, plus a Docker setup for running them. Its README now points readers to a third-edition repository and says this one will no longer be maintained.
Who is it for?
Use this repository if you own the second edition and want its R and Python notebooks running locally, and accept that the README says it will no longer be maintained. If you are starting fresh, or you want material aligned with the third edition, go to gedeck/ai-assisted-statistics-for-data-scientists instead.
Can I use it commercially?
Yes, with conditions. GPL-3.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
Is it still maintained?
Yes. The repository last received commits 46 days ago.
What is it written in?
Mainly Jupyter Notebook, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What the repository actually contains, and who it is for

This is not a library you import. It is the companion code for Practical Statistics for Data Scientists: 50+ Essential Concepts Using R and Python by Peter Bruce, Andrew Bruce, and Peter Gedeck, published by O'Reilly in a second edition dated June 9, 2020. The top-level layout reflects that: a python/ directory, an R/ directory, a data/ directory, a docker/ directory, an environment.yml, an install.R, a requirements.txt, and a Makefile. The primary language recorded for the repository is Jupyter Notebook.

The intended reader is someone working through the book who wants to execute the examples rather than read them. That matters because the book covers both R and Python, and the repository keeps both tracks side by side rather than translating one into the other. The README also links to nbviewer for reading the notebooks online and to Binder for executing them in a hosted environment, noting that Binder can take some time if the environment needs to be rebuilt.

There is a second audience the README addresses directly, and it is not a happy one. A notice at the top states that the third edition is available under the name AI-Assisted Statistics for Data Scientists, links to that new repository, and says this repository will no longer be maintained, asking that issues and pull requests be filed against the new one. Anyone arriving here looking for current material should read that notice first.

How the Python track is wired: requirements.txt, environment.yml and dmba

The Python side is a flat list of packages in requirements.txt with no version pins: matplotlib, pandas, scikit-learn, scipy, statsmodels, wquantiles, seaborn, pygam, dmba, pydotplus, imblearn, xgboost, prince, numpy, adjustText. That is the whole file. Unpinned dependencies are a real constraint, not a stylistic one. A notebook written against a 2020 release of pandas or scikit-learn may behave differently under a later one, and nothing in the file stops a resolver from picking the newest available version.

One entry stands out: dmba. That is the companion package for the book's own material, and it is what the notebooks import for several of the worked examples rather than reimplementing them inline. If dmba is not installed, those cells fail before any statistics run. The other less obvious entries are pygam for generalized additive models, prince for correspondence analysis and PCA variants, imblearn for imbalanced sampling, wquantiles for weighted quantiles, and adjustText for label placement in matplotlib charts.

The R track is separate, with install.R and the R/ directory, and environment.yml covers the conda-style environment. The repository therefore carries three overlapping descriptions of its dependencies, which is convenient for a reader but means the three can drift apart.

Running the notebooks locally with the Makefile and Docker

The Makefile defines two image names, psfds and psfds_jupyter, and the targets that build and run them. The images target builds both from the Dockerfiles under docker/. Run it from the repository root.

bash
make images

That builds psfds from docker/Dockerfile.psfds and psfds_jupyter from docker/Dockerfile.jupyter. After that, the jupyter target starts Jupyter Lab inside the container, mounting the current directory at /src and publishing port 8893.

bash
make jupyter

The underlying command is a docker run with --rm, -v $(PWD):/src, -p 8893:8893, and jupyter lab --allow-root --port=8893 --ip 0.0.0.0 --no-browser. Because it binds to 0.0.0.0 with --allow-root and no token is set on the command line, treat this as a local-only convenience rather than something to expose on a shared network. Note that the mount is read-write, so notebooks you edit inside the container are written back into your working copy.

If you would rather not run Jupyter, the bash target drops you into the psfds image with the same mount, and bash-jupyter does the same with the Jupyter image. Both are useful for running a single script or checking which package versions actually landed in the image.

The maintenance question the README answers before you ask it

The banner at the top of the README is unusually direct for a book repository. It announces the third edition, gives the URL of the new repository, states that this repository will no longer be maintained, and asks that issues and pull requests go to the new location instead. The last push to this repository was on 2026-08-16, which is recent, but the README's own statement about maintenance is the thing to act on: an incoming issue here is likely to be closed with a pointer rather than a fix.

That has a practical consequence for anyone hitting an error. The failure you are most likely to see is a dependency resolution problem, since requirements.txt pins nothing. Filing it here is not the path the maintainer describes. The second consequence is that the R and Python notebooks will not be updated to match later package behaviour, so a cell that worked when the second edition shipped may need editing before it runs.

None of this makes the code wrong. It makes it fixed. For a statistics textbook companion, a fixed target is arguably the right design, because the numbers in the book should keep matching the numbers in the notebooks. The cost is that setup friction accumulates and nobody is going to remove it for you.

Where this repository is the wrong tool

If you want a statistics library to depend on, this is not it. There is no importable package here, no API surface, no release artefact, and the recent releases list is empty. The dmba package on PyPI is the closest thing to a reusable component, and it is a separate project that this repository happens to depend on.

If you are working from the third edition, this repository is the wrong edition. The README names the successor repository explicitly and describes it as covering AI-assisted statistics, which is a different scope from the 50+ concepts framing of the second edition. Reading the second-edition notebooks alongside third-edition text will produce mismatches in both the examples and the tooling.

If you need a hosted environment with no local setup, the Binder link is the intended route, but the README warns that it can take some time when the environment has to be rebuilt. Binder sessions also do not persist, so work done there is lost when the session ends. For anything you intend to keep, the Docker path or a local Python environment is the better choice.

Finally, if your environment is strict about dependency versions, the unpinned requirements.txt is a poor fit. You will need to pin versions yourself, and the repository gives you no guidance on which combination was current when the book was written.

How the R and Python tracks differ, and what to check before choosing one

The repository ships both languages because the book covers both, and the two tracks are not translations of one another. The R side carries install.R, the R/ directory, and the .Rproj file, which suggests the R material is meant to be opened as an RStudio project. The Python side carries requirements.txt and the notebooks under python/. The environment.yml sits above both, and the Docker images package the environment for either.

A reader who knows only one language should pick the matching track and stay there. The book's own framing is 50+ concepts illustrated in both languages, so the concept is the constant and the syntax is the variable. Working through both tracks in parallel doubles the setup work without adding statistical content.

The thing to verify first is whether the packages your chosen track needs are actually present. Run the make images target, then use the bash target to open a shell in the container and check what installed. If a notebook imports something that is not in requirements.txt or environment.yml, that is a gap you will have to fill yourself, and the README does not document a procedure for adding packages to the images.

Licence and the cost of keeping it running

The repository is licensed GPL-3.0. That is a copyleft licence, and it applies to the code in the repository rather than to the book's prose, which is O'Reilly's. If you copy notebook code into your own project, the GPL-3.0 terms travel with it, which is worth understanding before you paste a helper function into a proprietary codebase. This is a description of the licence identifier, not legal advice; read the LICENSE file and, if the distinction matters to your organisation, get proper counsel.

Upgrade cost is dominated by the unpinned dependencies. There is no lockfile, no changelog in the repository, and the README does not document a rollback procedure if a newer package version breaks a notebook. Your practical options are to let the resolver pick current versions and fix what breaks, or to pin everything yourself and accept that you are now maintaining a fork of the dependency list. The Docker images give you a third option: build them once, keep the images, and treat the built artefact as your pinned environment, which sidesteps the resolver entirely at the cost of a large local image.

Because the README states the repository will no longer be maintained, none of these costs will be reduced upstream. Budget for them as ongoing local work rather than as a one-time setup step.

Editorial conclusion

Use this repository if you own the second edition and want its R and Python notebooks running locally, and accept that the README says it will no longer be maintained. If you are starting fresh, or you want material aligned with the third edition, go to gedeck/ai-assisted-statistics-for-data-scientists instead. Before adopting it, verify that the pinned dependencies in requirements.txt still resolve for your Python version, and check whether the notebook you need is one of the ones whose packages are only listed by name with no version constraint.

Frequently asked questions

Is gedeck/practical-statistics-for-data-scientists still maintained?

The README states that the repository will no longer be maintained and asks that issues and pull requests be filed against the successor repository, gedeck/ai-assisted-statistics-for-data-scientists. The last push was on 2026-08-16, but the README's own notice is the clearer signal about where work happens now.

How do I run the notebooks from practical-statistics-for-data-scientists locally?

The Makefile has an images target that builds the psfds and psfds_jupyter Docker images from the Dockerfiles under docker/, and a jupyter target that runs Jupyter Lab on port 8893 with the current directory mounted at /src. The README also links to Binder for a hosted option, noting it can take some time if the environment needs rebuilding.

Does practical-statistics-for-data-scientists include both R and Python code?

Yes. The repository has separate R/ and python/ directories, plus install.R for the R track and requirements.txt for the Python track, with environment.yml covering the shared environment. The book it accompanies is subtitled 50+ Essential Concepts Using R and Python.

What dependencies does the Python side of practical-statistics-for-data-scientists need?

requirements.txt lists matplotlib, pandas, scikit-learn, scipy, statsmodels, wquantiles, seaborn, pygam, dmba, pydotplus, imblearn, xgboost, prince, numpy and adjustText, with no version pins. The dmba entry is the book's own companion package and several notebooks depend on it.

What licence is practical-statistics-for-data-scientists released under?

The repository is licensed GPL-3.0, and the LICENSE file sits at the top level. That covers the code in the repository; the book text itself is published by O'Reilly.

Official sources

  1. gedeck/practical-statistics-for-data-scientists on GitHub
  2. Issues
  3. License: GPL-3.0
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/gedeck-practical-statistics-for-data-scientists.svg)](https://hysenlabs.com/projects/gedeck-practical-statistics-for-data-scientists)