# Hands-On Data Analysis with Pandas, frozen at pandas 1.2.0

> Twelve chapter directories of Jupyter notebooks that accompany a Packt book, plus a solutions directory and three companion packages installed from GitHub at the 2nd_edition branch. Versions are exposed as git tags rather than releases, and the dependency set is pinned to pandas 1.2.0, numpy 1.19, and scikit-learn 0.23.2.

**stefmolin/Hands-On-Data-Analysis-with-Pandas-2nd-edition** — Materials for following along with Hands-On Data Analysis with Pandas – Second Edition

- Repository: https://github.com/stefmolin/Hands-On-Data-Analysis-with-Pandas-2nd-edition
- Website: https://www.amazon.com/Hands-Data-Analysis-Pandas-visualization/dp/1800563450
- Stars: 736 · Forks: 1,551
- Language: Jupyter Notebook
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/stefmolin-hands-on-data-analysis-with-pandas-2nd-edition

## Two tags carry the versions, and there are no releases at all

The repository does not version its notebooks the way a library would. There are no GitHub releases, and the versions sit in two git tags, 1st_edition and 2nd_edition, each holding a snapshot of the chapters as they stood at the time of publishing. Packt published the first edition on July 26, 2019 and the second on April 29, 2021, and the master branch that you land on holds the second edition content. The top level is one directory per chapter, ch_01 through ch_12, plus solutions, appendix, acknowledgements.md, environment.yml, requirements.txt, visual-aids, and an _img directory for the images the front page embeds. No topic tags are set on the repository. The last push on the default branch is 2026-04-11, which is later than either publication date, so the branch takes work without a second edition being re-published.

## The pinned stack stops at pandas 1.2.0 and numpy 1.19

requirements.txt is short enough to read in full, and it is the whole environment contract:

```text
graphviz==0.14.1
imbalanced-learn==0.7.0
ipympl==0.6.2
jupyterlab>=3.0.4,<=3.5.3
matplotlib==3.3.2
numpy>=1.19.2,<=1.19.5
pandas==1.2.0
requests==2.24.0
scikit-learn==0.23.2
scipy>=1.5.0,<=1.7.3
seaborn>=0.11.0,<=0.11.2
sqlalchemy>=1.3.19,<=1.3.20
statsmodels>=0.11.1,<=0.12.1
wheel
```

Seven entries are exact pins and six carry a lower bound with a ceiling, so the range entries have some room while pandas, matplotlib, scikit-learn, requests, graphviz, and imbalanced-learn have none. The front page explains the pandas number: the first edition was written against 0.23.4, which was pre 1.0, and the second edition moves the content to the 1.x line. The pin lands on 1.2.0, which sits early in that line rather than at its end. environment.yml sits next to requirements.txt at the root as the second format for the same environment.

## Three GitHub URLs and a relative path sit below the pinned packages

The last four lines of requirements.txt are not packages from an index. They are:

```text
git+https://github.com/stefmolin/login-attempt-simulator.git@2nd_edition
git+https://github.com/stefmolin/ml-utils.git@2nd_edition
git+https://github.com/stefmolin/stock-analysis.git@2nd_edition
./visual-aids
```

Three companion repositories are fetched at a branch named 2nd_edition rather than at a tag, so the content of an install is decided by whatever those branches point to. Two of them are named in the table of contents: login-attempt-simulator is the data generator behind chapter 8, and stock-analysis is the Python package chapter 7 builds. The third, ml-utils, is not referenced anywhere in the visible front page, so it arrives as an undeclared dependency of the setup. The fourth line installs from the repository's own visual-aids directory, which makes the install only partly remote. Chapter 3 is the one that shows how to explore an API to gather data, and requests is pinned at 2.24.0 for it.

## The front page offers three hosted launches and stops mid sentence on setup

You do not have to install anything to read a chapter. Three buttons at the top of the front page open the notebooks elsewhere: a Binder link that targets the master branch with a lab path, a Google Colab link, and an nbviewer link. That makes the repository readable as a reference with no environment work at all, which is also how most people first meet it. The written setup instructions are a separate matter. The section headed Notes on Environment Setup carries a link to an env-checks workflow run and then the sentence Environment setup instructions are in the chap, and the front page ends there, one word into the sentence that would point you at them. The two setup artifacts are both present at the root, environment.yml and requirements.txt, so what is missing is the walkthrough rather than the files.

## Chapters 7 and 8 hand their real code to other repositories

The first six chapters are where the pandas content lives. Chapter 2 introduces DataFrames, chapter 3 covers data manipulation with a walk through exploring an API to gather data and then cleaning and reshaping, chapter 4 teaches querying and merging plus rolling calculations and time series, chapter 5 builds visualizations first with matplotlib and then directly from pandas objects, and chapter 6 moves to seaborn for long form data and customization. Chapter 7 changes subject. Instead of teaching pandas it walks through the creation of a Python package for analyzing stocks and links to the separate stock-analysis repository. Chapter 8 covers simulating data by way of a separate login-attempt-simulator repository, applied to catching hackers attempting to authenticate to a website with rule-based strategies. Chapter 11 returns to that same login attempt data using machine learning techniques. So the back third of the book is a scikit-learn book, and pandas is the middle.

## The Python prerequisite is a notebook inside chapter 1

The front page states its own prerequisite before the learning objectives, and it is a file rather than a link to an external course. If you do not have basic knowledge of Python, or past experience with another language such as R, SAS, or MATLAB, it points at the ch_01/python_101.ipynb Jupyter notebook and calls it a Python crash-course refresher. That is the only language primer the repository carries, so a reader arriving from a statistical background has one notebook to work through before the DataFrames material makes sense. Chapter 1 itself is titled Introduction to Data Analysis and is described as covering the fundamentals, giving a foundation in statistics, and getting the environment set up for working with data in Python and using Jupyter Notebooks, which is where the setup write-up the front page starts pointing to would sit.

## The learning objectives promise scripts, modules, and packages

The bullet list of what you will learn reaches past analysis into packaging and algorithms. It asks the reader to build Python scripts, modules, and packages for reusable analysis code, to collect data from APIs, to write and run simulations, and to use computer science concepts and algorithms to write more efficient code for data analysis. It also lists combining, grouping, and aggregating data from multiple sources, creating visualizations with pandas, matplotlib, and seaborn, and applying machine learning algorithms with sklearn to identify patterns and make predictions. Three promises deserve a second look against the repository layout. scripts and packages is exactly what chapter 7 delivers through the stock-analysis repository rather than from a directory here, simulations is chapter 8's login-attempt-simulator, and the solutions directory at the root holds answers for the exercises that the visible front page never itemises.

## Conclusion

This repository is worth using as a worked reference rather than as a course you complete once, because the chapters build one dataset pipeline on top of another and the solutions directory lets you check a step without reading ahead. It fits an application built on pandas 1.x, since that is what the pins describe. Two things to check first. Read requirements.txt before you install anything, because the set includes three GitHub URLs pinned to a 2nd_edition branch plus a relative path into visual-aids/, so a clean pip install depends on state outside this repository. And decide what to do about the age of the stack yourself: pandas 1.2.0 and scikit-learn 0.23.2 are fixed rather than ranged, and a notebook written against those APIs will not simply run on a current environment. If you want pandas idioms from the 1.x era, this is exactly the right source. If you want current ones, read it for structure and rewrite the calls.

## FAQ

### How do I run the notebooks in Hands-On Data Analysis with Pandas without installing anything?

The front page carries three hosted launch buttons: a Binder link targeting the master branch with a lab path, a Google Colab link, and an nbviewer link. That opens a chapter in a hosted Jupyter environment without touching a local environment.

### What pandas version do the Hands-On Data Analysis with Pandas notebooks need?

requirements.txt pins pandas at 1.2.0 exactly, with no range. The front page explains that the first edition was written against 0.23.4, which was pre 1.0, and that this edition moves the content up to the 1.x line. NumPy is bounded to 1.19.2 through 1.19.5 and scikit-learn is pinned at 0.23.2.

### Is the code for the second edition of Hands-On Data Analysis with Pandas in this repository?

The chapter material is, as ch_01 through ch_12 alongside solutions/, appendix/, acknowledgements.md, visual-aids/, environment.yml, and requirements.txt. Three companion pieces are not: requirements.txt fetches login-attempt-simulator, ml-utils, and stock-analysis from GitHub at the 2nd_edition branch.

### What is the difference between the 1st_edition and 2nd_edition tags?

Each one holds the chapters as they stood at publication time. Packt published the first edition on July 26, 2019 and the second on April 29, 2021. The default branch master holds the second edition, and the repository publishes no GitHub releases.

### Does Hands-On Data Analysis with Pandas require Python experience to follow along?

The front page names a prerequisite and points at one file. Readers without basic Python knowledge, or with experience only in another language such as R, SAS, or MATLAB, are directed to the ch_01/python_101.ipynb notebook described as a Python crash course.

## Sources

- [Issues](https://github.com/stefmolin/Hands-On-Data-Analysis-with-Pandas-2nd-edition/issues)
- [License: MIT](https://github.com/stefmolin/Hands-On-Data-Analysis-with-Pandas-2nd-edition/blob/master/LICENSE)
- [Project website](https://www.amazon.com/Hands-Data-Analysis-Pandas-visualization/dp/1800563450)
- [README](https://github.com/stefmolin/Hands-On-Data-Analysis-with-Pandas-2nd-edition/blob/master/README.md)
- [stefmolin/Hands-On-Data-Analysis-with-Pandas-2nd-edition on GitHub](https://github.com/stefmolin/Hands-On-Data-Analysis-with-Pandas-2nd-edition)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/stefmolin-hands-on-data-analysis-with-pandas-2nd-edition
