Open-source project
szcf-weiya/ESL-CN avatar
szcf-weiya/ESL-CN

ESL-CN: A Chinese Translation of The Elements of Statistical Learning, With Code and Exercise Solutions

The Elements of Statistical Learning (ESL)的中文翻译、代码实现及其习题解答。

2,787 stars618 forksJupyter NotebookGPL-3.0

At a glance

What is it?
ESL-CN is a community translation project for Hastie, Tibshirani and Friedman's textbook, bundled with code implementations in R, Julia, Python and C++ and per-chapter exercise solutions tracked as GitHub milestones. It is a study companion, not a library you import.
Who is it for?
Adopt ESL-CN if you are studying the ESL textbook in Chinese and want worked code next to the derivations, or if you want to contribute translations and solutions chapter by chapter. Do not adopt it if you need a maintained Python package, an English-language reference, or a complete solution set: the README's own progress list, kept as an HTML comment, shows several subsections still unchecked, and the milestone badges are the only signal of what is done.
Can I use it commercially?
Yes, with conditions. GPL-3.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
Is it still maintained?
Yes. The repository last received commits 9 days ago.
What is it written in?
Mainly Jupyter Notebook, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What ESL-CN Actually Is, and Who It Is For

The repository describes itself as the Chinese translation of The Elements of Statistical Learning, together with code implementations and exercise solutions. Those are three separate deliverables living in one tree, and they age at different rates.

The translation lives under docs/ and rmds/, rendered through mkdocs.yml to the site at esl.hohoweiya.xyz. The code lives under code/, with one directory per topic: EM, NaiveBayes, CART, boosting, MARS, rbm, Gibbs, SOM, nonParametrics, Resampling, nn and HighDim. The exercise solutions are not files at all. Each chapter has a GitHub milestone, and the README lists eighteen of them, chapter 2 through chapter 18, as badges.

That structure tells you who the project is for. It is for a reader working through ESL in Chinese who wants to see a method implemented rather than only derived. It is also for contributors, because a milestone is a task list: an open issue under chapter 7's milestone is a concrete piece of work someone can claim. It is not for someone who wants to pip install something. There is no package here.

How the Repository Is Organized and How the Site Gets Built

The top-level layout is a documentation project with a code appendix bolted on. mkdocs.yml drives the site. docs/ holds the rendered pages. rmds/ holds the R Markdown sources that produce some of them, which is why several README links point at rmd.hohoweiya.xyz rather than at the site itself. imgs/ and material/ are assets. gentag.py sits at the root, a small script whose name suggests tag generation for the docs.

A GitHub Actions workflow, gh-pages.yml, carries the build badge at the top of the README, so the published site is regenerated from the repository rather than committed by hand. The consequence for a reader is that the site and the repository can differ: a page you read at esl.hohoweiya.xyz reflects whatever the last successful workflow run produced.

The code side is deliberately polyglot. The topics list includes cpp, julia, python and r, and the README's code section names the language per entry: the EM simulation is em.R, AdaBoost is described as R and Julia, and the neural network entry links to Jupyter notebooks hosted on nbviewer. Two of those links point outside the repository entirely, at szcf-weiya/TFnotes and at an nbviewer URL. That is a real fragility: external notebook hosting is not covered by the GPL-3.0 licence of this repository, and a dead nbviewer link is not something a commit here can fix.

Installing Nothing: Reading the Site and Running a Code Example

There is no install step, because ESL-CN is not a package. The README points readers at the published site, and the code is meant to be cloned and run file by file.

Start by cloning the repository so the code and data directories are on disk:

bash
git clone https://github.com/szcf-weiya/ESL-CN.git
cd ESL-CN

The README's code section then names each entry by path. To work through resampling methods, which the README describes as covering cross validation and the bootstrap, go into that directory:

bash
cd code/Resampling
ls

What you see depends on which language that topic was written in. The EM entry is em.R, so that one needs an R interpreter; the AdaBoost entry is R and Julia; the neural network entry is a Jupyter notebook. There is no single environment file at the repository root that pins all of this, so expect to install interpreters per topic rather than once.

If you want the book text rather than the code, read it at esl.hohoweiya.xyz and treat the repository as the source of that site. Building the site yourself means invoking MkDocs against mkdocs.yml, which is the same configuration the gh-pages workflow uses.

The Exercise Solutions Are Milestones, Not Files

This is the design choice most likely to surprise a new reader. You cannot open a solutions directory and read chapter 7's answers. You open milestone 4, which the README labels as chapter 7, and read the issue list.

That has one clear advantage: partial work is visible and claimable. If exercise 7.4 is open, you know nobody has finished it, and you can volunteer. A static PDF of solutions cannot express that state.

The disadvantage is that completeness is hard to judge at a glance. The README's progress list is inside an HTML comment, so it does not render on the repository page, and it is stamped with a 2017 date. That list shows whole subsections unchecked: 3.7 through 3.9, 5.5 through 5.9, 6.7 through 6.9, 8.8 and 8.9, 10.7 through 10.14, 11.7 through 11.10, 12.3 through 12.7, 13.2 through 13.5, 14.4 through 14.10, 15.2 through 15.4, 16.1 through 16.3, and 18.1 through 18.8. The list is old and the repository has been pushed since, so treat it as a lower bound on what is missing rather than a current inventory. The milestones are the authoritative view.

There is also a literature statistics table in the README, counting how many references in each chapter come from AOS, JASA, JRSS and BKA, with a coverage ratio. Chapter 4 shows 0/7 and chapter 17 shows 0/12. It is an unusual thing to publish, and it is honest about where the annotation work has not reached.

Where ESL-CN Is the Wrong Tool

If you need a maintained library, this is not it. Nothing here is versioned, released, or importable. There are no releases in the repository metadata, so there is no changelog to read and no compatibility promise. Code examples are teaching artifacts, not APIs, and a change in an R package or a Julia version can break a script without anything in this repository changing.

The second mismatch is language. The project is a Chinese translation. If you are reading ESL in English, the translation adds a layer between you and the original text, and the original is the thing the exercises refer to. Use the code and the milestones, skip the prose.

The third is scope. ESL is a theory-heavy book. The code directory covers a selection of methods, not the book end to end. Chapter 4's literature table reading 0/7 is a hint that some chapters have had less attention than others, and the neural network material partly lives in a separate repository. If your interest is a chapter with a thin code directory and an open milestone, this project will not carry you.

Finally, the licence matters for reuse. GPL-3.0 is copyleft. If you want to lift a translation passage or a code file into a proprietary course pack, the licence is a constraint you need to read for yourself; this is a description of the licence, not legal advice.

How It Compares With an Introduction to Statistical Learning

The obvious alternative is An Introduction to Statistical Learning, usually abbreviated ISL, and its Chinese translation. The search data around this project shows readers moving between the two, and the comparison is worth stating plainly because the two books occupy different positions.

ISL is the lighter, applied companion, written for readers who want to fit models in R and interpret the output. ESL assumes more mathematics and spends its pages on the theory behind the same methods. ESL-CN inherits that split. Its code directory is there to illustrate a derivation, not to teach a workflow. If your goal is to run a lasso on a dataset this afternoon, the ESL-CN code for that topic is a demonstration you read, whereas ISL's material is organized around doing.

The other alternative is simply the English ESL with the official book site and the authors' own resources. That path gives you the authoritative text and no translation drift, at the cost of the Chinese-language annotation and the milestone-based solution tracking that ESL-CN adds. Neither is a substitute for the other; they answer different questions.

Maintenance, Contribution Cost and the Licence

The repository is not archived, and the last push was on 2026-09-22, which places it within the last week. The gh-pages workflow badge in the README is the visible sign that the site build still runs. There are no releases, which for a documentation project is normal rather than alarming: the site is the artifact.

The upgrade cost is low in one direction and unpredictable in the other. Because there is nothing to install, nothing breaks on your side when the repository changes. But the code examples depend on external interpreters and, in the neural network case, on notebooks hosted elsewhere. Pinning those is your problem, not the project's.

Contributing has a defined shape. Pick an open issue under a chapter milestone, or fill a gap the progress list identifies, and open a pull request. The GPL-3.0 licence covers the repository contents, which means derivative works that incorporate them carry the same licence. That is the practical implication for anyone thinking about republishing the translation or the code in another product; read the licence text and, if the stakes are high, a lawyer, because this article is not legal advice.

Editorial conclusion

Adopt ESL-CN if you are studying the ESL textbook in Chinese and want worked code next to the derivations, or if you want to contribute translations and solutions chapter by chapter. Do not adopt it if you need a maintained Python package, an English-language reference, or a complete solution set: the README's own progress list, kept as an HTML comment, shows several subsections still unchecked, and the milestone badges are the only signal of what is done. Before relying on a chapter, open that chapter's milestone and confirm the exercises you need are closed, then check whether the code directory for that topic has a runnable entry rather than a link to an external notebook host.

Frequently asked questions

What does ESL-CN stand for?

ESL stands for The Elements of Statistical Learning, the textbook by Hastie, Tibshirani and Friedman, and CN marks the Chinese translation. The repository's own description is the Chinese translation of ESL together with code implementations and exercise solutions.

Where can I get the ESL-CN PDF?

The README does not offer a PDF. It points to the published site at esl.hohoweiya.xyz, which is built from the docs directory via mkdocs.yml, and to R Markdown pages hosted at rmd.hohoweiya.xyz for some examples.

Is ESL-CN the same as An Introduction to Statistical Learning?

No. ESL-CN translates ESL, which is the more theoretical of the two books, and adds code implementations and milestone-tracked exercise solutions. ISL is a separate, more applied text, and the search results around this project show readers comparing the two.

What programming languages does ESL-CN use for its code?

The repository topics list cpp, julia, python and r, and the README names the language per entry: the EM simulation is em.R, AdaBoost is described as R and Julia, and the neural network material is Jupyter notebooks.

Official sources

  1. Issues
  2. License: GPL-3.0
  3. Project website
  4. README
  5. szcf-weiya/ESL-CN on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/szcf-weiya-esl-cn.svg)](https://hysenlabs.com/projects/szcf-weiya-esl-cn)