# datawhalechina/statistical-learning-method-solutions-manual: a Jupyter-based answer manual for Li Hang's textbook

> The repository pairs worked solutions to Li Hang's Statistical Learning Methods and Machine Learning Methods with runnable Jupyter notebooks and Python code. It is an Alpha build, and the README says so.

**datawhalechina/statistical-learning-method-solutions-manual** — 机器学习方法习题解答，在线阅读地址：https://datawhalechina.github.io/statistical-learning-method-solutions-manual

- Repository: https://github.com/datawhalechina/statistical-learning-method-solutions-manual
- Website: https://datawhalechina.github.io/statistical-learning-method-solutions-manual
- Stars: 2,090 · Forks: 251
- Language: Jupyter Notebook
- License: NOASSERTION
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/datawhalechina-statistical-learning-method-solutions-manual

## What the answer manual actually covers, and who it is written for

This is a solutions manual, not a library. The README describes it as answers to the exercises in Li Hang's Statistical Learning Methods and Machine Learning Methods, organized into three parts: supervised learning, unsupervised learning, and deep learning. The chapter list runs from the perceptron and k-nearest neighbors through logistic regression, support vector machines, boosting, EM, hidden Markov models and conditional random fields, then clustering, SVD, PCA, latent semantic analysis, MCMC, LDA and PageRank, and finally feedforward, convolutional and recurrent networks, sequence-to-sequence models, pretrained language models and GANs. The README names three audiences: students working through the textbook, engineers who want Python implementations of the algorithms, and people revising for exams or interviews. That last group is the honest one. The material is written to be read alongside a textbook, and the README asks contributors to keep the mathematics at a level a beginner with calculus can follow. If you want a maintained implementation of a conditional random field, this is the wrong repository. If you want to see how the forward-backward recursion falls out of the HMM definitions in chapter 10, it is aimed squarely at you.

## How the notebooks, docs and codes directories fit together

The repository has three parallel trees, and understanding the split saves confusion. Solutions are authored as Jupyter notebooks under notebook/notes, then exported to Markdown and copied into docs/, which is the Vitepress site published at the homepage URL. Standalone exercise scripts live in codes/, one directory per chapter, with filenames that name the exercise: codes/ch02/perceptron.py for exercise 2.2, codes/ch07/svm_demo.py for exercise 7.2, codes/ch11/crf_matrix.py for exercise 11.4. The README's collaboration rules state the reason for the notebook-first workflow: exercises need program output, so the notebook is the source of truth and the Markdown is a rendering of it. That has a consequence worth naming. The docs site can lag the notebooks, because the copy step is manual rather than a build hook. The codes/ tree is separate again, holding plain scripts rather than notebooks, so a fix applied to a notebook in notebook/notes does not automatically reach the corresponding script in codes/. When you find a discrepancy, the notebook is the version the project treats as canonical.

## Installing the environment and running your first exercise

The README specifies Python 3.12 or later and Node 18.20.4 or later, and uses uv for dependency management. Start by installing uv and pointing it at the mirror the README configures:

```bash
pip install uv
set UV_INDEX=https://mirrors.aliyun.com/pypi/simple
```

The project itself pins Python 3.12 in pyproject.toml, and the README's sync command follows that:

```bash
uv sync --python 3.12 --all-extras
```

That installs the declared dependencies, which include numpy, pandas, scikit-learn, matplotlib, graphviz, notebook, h5py, tqdm and simple-elmo-pytorch. PyTorch is not in that list. The README tells you to install it separately from the official PyTorch site, with one concrete example pinned to CUDA 11.8 wheels:

```bash
uv pip install torch==2.7.1 torchvision==0.22.1 torchaudio torchviz --index https://download.pytorch.org/whl/cu118
```

Graphviz is also a separate install, and the README links to an external blog post for it rather than documenting the steps itself. That is a gap: the decision-tree chapters render trees through graphviz, so a missing binary shows up as a failed cell rather than a clear message. Once the environment is in place, start the notebook server:

```bash
jupyter notebook
```

The README also documents the docs site, which needs the Node side:

```bash
npm run docs:dev
```

Open a notebook from notebook/notes, run the cells top to bottom, and compare the printed output with the analysis in the surrounding Markdown. For a first pass, codes/ch02/perceptron.py is the smallest complete example in the tree.

## The Alpha label is a real constraint, not a formality

The README carries a caution block stating that this is an early internal build, incomplete, and possibly containing errors, and it asks readers to file issues. The progress table backs that up with specifics. Most chapters are marked complete, but chapter 26, sequence-to-sequence models, is marked in progress, and the table's chapter numbering skips from 11 to 14 and from 21 to 23, matching the structure of the book rather than a continuous sequence. Several chapters are also missing from the table entirely. The practical failure mode is not a crash; it is a derivation that reads plausibly and is wrong in a sign or an index. The README's own answer to that is the review column, which lists named reviewers per chapter, so the completed chapters have had at least one pass by someone other than the author. That is a meaningful signal, and it is also the limit of what the repository claims. There is no test suite in the top-level entries, no CI badge, and no statement that the code output has been verified against an independent implementation. Treat every result as something to check, particularly the self-programmed algorithms where the point of the exercise is to write the method from scratch rather than call a library.

## Where a library like scikit-learn is the better tool

Several exercises deliberately solve the same problem twice: once by calling a library and once by writing the algorithm by hand. Chapter 5 has codes/ch05/k_neighbors_classifier.py using scikit-learn's DecisionTreeClassifier alongside codes/ch05/my_decision_tree.py implementing C4.5 directly. Chapter 8 pairs an AdaBoostClassifier example with a hand-written AdaBoost. Chapter 9 pairs GaussianMixture with a from-scratch GMM. That pairing is the whole point, and it also tells you when to look elsewhere. If your goal is to train a decision tree or a Gaussian mixture on real data, scikit-learn is the maintained, tested, documented choice, and the repository's own code says so by using it in the comparison scripts. The manual's value is in the gap between the two versions: the hand-written file shows the update rule, the library call shows the convention the ecosystem settled on, and the difference between them is often where textbook notation and production APIs diverge. Read the pair, do not adopt the pair.

## Licence and the cost of keeping a fork current

The repository reports its licence as NOASSERTION, which means GitHub could not match the LICENSE file to a known identifier. The file exists at the top level, so the terms are stated somewhere, but you have to open LICENSE and read it yourself before you reuse the content. That matters more here than for a typical code repository, because the substantial work is prose and derivations rather than software, and the licence may treat those differently from the Python files. On maintenance: the last push was on 2026-05-08, so the repository is not archived but has not received a push in roughly four months. Upgrades are cheap in the sense that there is no runtime to keep patched, and expensive in the sense that the dependency set is pinned to recent versions. pyproject.toml requires Python 3.12 or later and lists numpy 2.4.2, pandas 3.0.0 and scikit-learn 1.8.0 as minimums, and the README pins torch 2.7.1 with torchvision 0.22.1. Those are aggressive floors. If your environment is on an older Python or an older numpy, uv sync will not resolve, and you will be editing pyproject.toml before you can run a single notebook. For a teaching repository, that is a defensible choice, since it keeps the examples on current APIs. It does mean the manual is not something you can drop into a legacy environment unchanged.

## Conclusion

Adopt it if you are working through Li Hang's exercises and want a second derivation plus runnable Python to compare against your own. Do not adopt it as a reference implementation to ship, and do not treat the Alpha notice as boilerplate: the README states the build is incomplete and may contain errors, and chapter 26 is still marked in progress. Before relying on any chapter, open its notebook under notebook/notes, run it against the pinned dependency set in pyproject.toml, and check the corresponding page under docs/, because the site is generated from the notebooks and the two can diverge.

## FAQ

### What are the key concepts in statistical learning?

The README frames the book as three parts: supervised learning methods such as the perceptron, k-nearest neighbors, naive Bayes, decision trees, logistic regression, support vector machines, boosting, EM, hidden Markov models and conditional random fields; unsupervised methods including clustering, SVD, PCA, latent semantic analysis, MCMC, LDA and PageRank; and deep learning methods from feedforward networks through GANs.

### What Python version does statistical-learning-method-solutions-manual require?

The README specifies Python 3.12 or later, and pyproject.toml sets requires-python to >=3.12. The README's setup command is uv sync --python 3.12 --all-extras, so the pinned interpreter matches that floor.

### Is statistical-learning-method-solutions-manual complete?

No. The README labels the build an Alpha internal test version that is incomplete and may contain errors, and its progress table marks chapter 26 on sequence-to-sequence models as still in progress. The project asks readers to file issues for problems or suggestions.

### What licence does statistical-learning-method-solutions-manual use?

The repository reports NOASSERTION, meaning the LICENSE file could not be matched to a standard identifier. The file is present at the top level, so the terms are stated there and need to be read directly before reuse.

## Sources

- [datawhalechina/statistical-learning-method-solutions-manual on GitHub](https://github.com/datawhalechina/statistical-learning-method-solutions-manual)
- [Issues](https://github.com/datawhalechina/statistical-learning-method-solutions-manual/issues)
- [Project website](https://datawhalechina.github.io/statistical-learning-method-solutions-manual)
- [README](https://github.com/datawhalechina/statistical-learning-method-solutions-manual/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/datawhalechina-statistical-learning-method-solutions-manual
