# Twenty-five chapter folders, a bug list, and no LICENSE file

> This is the companion code repository for a Chinese machine learning textbook: 25 chapter directories spelled charpter, a shared mlbook package, and a round of bug fixes dated April 2026 that the code printed in the book does not contain.

**luwill/Machine_Learning_Code_Implementation** — Mathematical derivation and pure Python code implementation of machine learning algorithms.

- Repository: https://github.com/luwill/Machine_Learning_Code_Implementation
- Stars: 1,556 · Forks: 585
- Language: Jupyter Notebook
- License: not declared
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/luwill-machine-learning-code-implementation

## Twenty-six algorithms, twenty-five folders, three library chapters

The project opens by promising derivations and code for 26 classic algorithms across four categories: supervised single models, supervised ensembles, unsupervised models and probabilistic models. The tree offers 25 chapter directories, `charpter1_ml_start` through `charpter25_MCMC`, and nothing in the page maps twenty-six algorithms onto them. The four categories are not reflected in the naming either, since every folder is one chapter rather than one category, so a reader has to work out which of the four buckets any given folder belongs to. Three folders are named for libraries that already exist, `charpter12_XGBoost`, `charpter13_LightGBM` and `charpter14_CatBoost`, and `charpter16_ensemble_compare` compares ensembles rather than deriving one. That leaves room for the count to be reached more than one way: a chapter covering two algorithms, or a claim that was never checked against the tree. The repository description says pure Python code implementation, while the primary language here is recorded as Jupyter Notebook, so what ships is notebooks.

## Every folder is spelled charpter, and the typo is load bearing

All 25 chapter directories carry the misspelling `charpter` in place of chapter, and the consistency is what makes it tolerable: nobody has to remember which half of the pair is right. It is not cosmetic. The revision note refers to chapters as Ch5, Ch7, Ch11, Ch15, Ch19, Ch23 and Ch25, the book's prose says 章, and the folders say `charpter`, so every reference between the text, the revision notes and the tree has to be translated by hand. Renaming would break the book's own cross-references and every external link shaped like the errata link the README shows, which is a blob URL on the master branch. The default branch being `master` rather than `main` adds a second thing to remember when you follow a path from the page. Three spellings of one idea, chapter, Ch and charpter, now have to stay in sync, and only the folders are load bearing at runtime.

## The shared library replaced the file paths the book teaches

The April 2026 review created `mlbook/` and collapsed six duplicated `utils.py` and `cart.py` files from Ch7, Ch11, Ch12 and Ch15 into two shared ones, and the page says plainly that the repository code is kept updated and iterated relative to the code printed in the book. So the refactor is a deliberate divergence, not a repair of a slip. The effect lands on imports: a notebook or listing that opens with a chapter-local path to `utils` or `cart` no longer matches anything in the tree, and the revision note is the only place the new location appears. The same review added a `tests/` directory with 25 cases covering the shared library and the key algorithms, and its closing line points readers at the commit history and that directory rather than at the errata table, which the README links separately for the book itself. Code changes and book changes are therefore tracked in two different places, and only one of them is a file you can read.

## The bug list is specific enough to check, and one fix lands on 1.0

The listed repairs are concrete: Ch5 LDA standardized wrongly inside `calc_cov`, which the README says moved accuracy from 0.85 to 1.0; Ch25 MCMC passed y=-1 into `p_xy` on every Gibbs step instead of the state transition value; Ch3 logistic regression carried an O(n^2) loop in `accuracy` and needed log clipping in cross-entropy to avoid NaN; Ch23 HMM hardcoded a state count of 4 in both the forward and Viterbi implementations; Ch19 SVD had a hardcoded Windows path replaced by `os.path.join`. The HMM entry is the quiet failure, because a fixed state count still produces plausible output for four states and wrong output for every other number. The LDA entry deserves the same suspicion in reverse: an accuracy of exactly 1.0 on whatever fixture was used is what a small separable dataset returns, and the page never names the dataset or the split. Spelling fixes in the same list, `missclassification` and `initilize_with_zeros`, mean a call written against the book can fail against this repository.

## pyproject.toml names a package it does not build

The project metadata is short: name `mlbook`, version 0.1.0, a description and `requires-python = ">=3.9"`. There is no build-system table and no dependencies array, so nothing here tells pip which backend to use or what to install, and the dependency list exists only in `requirements.txt`. What the file does configure is tooling. Ruff runs with line-length 100, target-version py39 and the E, F, W, I, N and UP rule sets, which is the rule list that catches the unused imports and the old typing spellings the review was cleaning up. Pytest is pointed at `testpaths = ["tests"]` with `python_files = ["test_*.py"]`. Two limits are worth naming. The lint target is Python modules while the primary language here is notebooks, so the files that dominate the tree are not what those rules were written for, and test discovery counts only files whose names start with `test_`.

## Dependency floors without ceilings, and one floor from another era

Eight dependencies are listed with lower bounds and no upper bounds: numpy>=1.24, pandas>=2.0, scikit-learn>=1.3, scipy>=1.11, matplotlib>=3.7, cvxopt>=1.3, pgmpy>=0.1 and jupyter>=1.0. Six of those floors sit at major version 2 or 3, while pgmpy is still written as 0.1, a numbering style the others left behind before reaching 1.0. The numpy floor explains part of the API repair: `np.float` and `np.matrix` are not available at 1.24, which is why the review replaced them with `np.float64` and standard arrays and why `sklearn.datasets.samples_generator` had to become `make_blobs`. With nothing pinned above, a fresh install resolves to the newest release of each package, newest pgmpy included, and the book has no mechanism for naming the versions its text and its numbers were written against. Anyone reproducing a chapter result is on their own there.

## The written material sits behind the purchase links

This is the code half of a published book. The page links a JD listing and a Dangdang listing, says the book content will be open-sourced later, and confirms that only the code is public now. The companion slides are not in the repository at all: readers who bought the paper book are told to contact the author through a WeChat public account. Video lectures are announced as still being produced, with a single chapter-one link. Book errata live in `Errata/Errata.md`, reached through the master branch. Licensing is stated in prose rather than in a file: the README names the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International license and links the CC BY-NC-SA 4.0 deed, so the code is non-commercial and share-alike, while the repository's own license field is empty and no LICENSE file sits at the top level next to `pyproject.toml` and `requirements.txt`.

## Conclusion

Use this repository as the corrected version of a book's code, not as a drop-in substitute for it. The April 2026 review is the reason to prefer it: six duplicated modules became one shared library, and bugs that reached print, including a Gibbs sampler that never updated its state and an HMM hardcoded to four states, are fixed here and not in the book. Before you rely on any chapter, check which folder your import should point at, since `mlbook/` replaced the chapter-local `utils.py` and `cart.py`, and pin your dependency versions yourself, because every requirement is a floor with no ceiling and `pgmpy` is still written as 0.1. The last commit on the default branch is dated April 29, 2026.

## FAQ

### How many algorithms are implemented in Machine Learning Code Implementation?

The README claims 26 classic algorithms in four categories: supervised single models, supervised ensembles, unsupervised models and probabilistic models. The repository carries 25 chapter directories, `charpter1_ml_start` through `charpter25_MCMC`, and the page does not say which chapter holds the twenty-sixth algorithm.

### What license does Machine Learning Code Implementation use?

The README states a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International license and links the CC BY-NC-SA 4.0 deed. The repository's license field is empty, and there is no LICENSE file at the top level.

### What did the April 2026 code review change?

It created the shared `mlbook/` library to replace six duplicated `utils.py` and `cart.py` files from Ch7, Ch11, Ch12 and Ch15, added `tests/` with 25 cases, fixed LDA covariance standardization, a Gibbs sampling bug in Ch25, an O(n^2) accuracy loop and NaN clipping in Ch3, and a hardcoded state count in Ch23 HMM.

### How do you install the dependencies for Machine Learning Code Implementation?

`requirements.txt` asks for numpy>=1.24, pandas>=2.0, scikit-learn>=1.3, scipy>=1.11, matplotlib>=3.7, cvxopt>=1.3, pgmpy>=0.1 and jupyter>=1.0, all with lower bounds only. `pyproject.toml` names the package mlbook, requires Python >=3.9, and declares no dependencies and no build backend.

### Is the book text included in Machine Learning Code Implementation?

No. The repository says only the code is open so far and that the book content will be open-sourced later. The companion slides go to readers who bought the physical book through a WeChat account, and the video lectures have a chapter-one link with the rest still in progress.

## Sources

- [Issues](https://github.com/luwill/Machine_Learning_Code_Implementation/issues)
- [luwill/Machine_Learning_Code_Implementation on GitHub](https://github.com/luwill/Machine_Learning_Code_Implementation)
- [README](https://github.com/luwill/Machine_Learning_Code_Implementation/blob/master/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/luwill-machine-learning-code-implementation
