ML-From-Scratch, a NumPy teaching library that skips the optimization
GitHub describes it as Machine Learning From Scratch. Bare bones NumPy implementations of machine learning models and algorithms with a focus on accessibility. Aims to cover everything from linear regression to deep learning.. The repository metadata lists Python as its primary language. The metadata lists the MIT license. This article stays within the project description and details documented in the GitHub repository README.
At a glance
- What is it?
- ML-From-Scratch is a MIT-licensed collection of machine learning models written in plain NumPy so you can read the math, not so you can ship them. It is built for reading the inner workings of an algorithm, which is exactly why the setup.py pins nothing and the examples print parameter tables instead of benchmarks.
- Who is it for?
- Adopt ML-From-Scratch if you want to read a working implementation of gradient descent, a CNN, a GAN or an RBM in one sitting, and you are learning rather than deploying. Do not adopt it for anything that has to run fast or on current libraries; it targets old pinned dependencies and its last push was 2023-10-15, with no GitHub releases to fall back on.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Probably not. The repository last received commits 36 months ago, on October 15, 2023.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The project optimizes for readability, not for speed
The stated purpose is not to produce algorithms as optimized and computationally efficient as possible, but to present their inner workings transparently. That single choice explains most of what you will notice. There is no GPU code, no compiled operators and no batching tricks. Every model is written in plain NumPy so you can follow what happens to an array at each step.
What you get spans the usual textbook spread, and the spread is wide. Polynomial regression, a convolutional network, DBSCAN, a generative adversarial network, a deep Q-network, a restricted Boltzmann machine, a neuroevolution example, a genetic algorithm and association analysis all ship, grouped under supervised, unsupervised, reinforcement and deep learning. Note what is absent: nothing here wraps a PyTorch or TensorFlow tensor, so there is no autodiff underneath to hide the derivative step you are trying to read. The design target is a reader who wants to see the update rule, not a user who wants the fastest result, and that gap shows up most in the deep learning examples, where every backward pass is written out by hand.
Installation is a clone plus setup.py install
The README gives three commands, and they are the whole install path:
git clone https://github.com/eriklindernoren/ML-From-Scratch
cd ML-From-Scratch
python setup.py installThe package name is `mlfromscratch` and the version in setup.py is `0.0.4`. Once it is installed, you run any model as a standalone example script. A supervised one looks like this:
python mlfromscratch/examples/polynomial_regression.pyThe other examples follow the same pattern, with some living under the library rather than the examples directory. Generating handwritten digits, for instance, is run with `python mlfromscratch/unsupervised_learning/generative_adversarial_network.py`, so the path tells you whether the file was filed as an example or as part of a topic module. That is the intended workflow: install the library once, then execute individual scripts to watch one algorithm train. Nothing bundles the examples into a single entry point.
requirements.txt pins no versions, so your Python decides what breaks
Every dependency is listed without a version constraint: matplotlib, numpy, sklearn, pandas, cvxopt, scipy, progressbar33, terminaltables and gym. Meanwhile setup.py asks for `numpy>=1.10` and `scipy>=0.17` under setup_requires, and its own reading of requirements.txt pulls the rest into install_requires.
Two consequences follow. You get whatever version of NumPy and SciPy your resolver finds, which on a fresh environment is a modern one, while the code was written against a much older API surface. And the set of extra packages is broad: if you only want to read the regression example, you are still pulling in a reinforcement learning library and an SVM solver. Check each example's imports before installing everything, and expect to pin numpy and scipy yourself if a given script errors on import.
The examples print layer tables and accuracy, not throughput
Instead of a benchmark, most examples dump a per-layer summary. The convolutional example on the digit dataset prints a table of layer type, parameter count and output shape, ending in a total of 538570 parameters, then a training bar and an accuracy figure. The generative adversarial network prints two such tables, one for a 1489936-parameter generator and one for a 533762-parameter discriminator.
Read those numbers as illustrations of architecture, not as a performance claim. The point of the table is that you can see a convolution expand 8x8 input channels to 16 and then 32, or watch a dense layer grow from 256 to 1024 units, while the code runs. The accompanying accuracies and timings exist to confirm the model learned, not to compare it against anything else. If you need speed, this library is telling you it is the wrong shape of tool.
No GitHub releases, and the last push was 2023-10-15
The repository is not archived, but it also has no GitHub releases, so there is no tagged version to pin and nothing to upgrade to. The last push was on 2023-10-15. The only version string in the project is the `0.0.4` hard-coded in setup.py, which describes the packaging, not a maintained release train.
For a reference library that is a defensible choice. These files are meant to be read, copied into your own project, and edited. But it means there is no upstream channel for bug fixes or compatibility shims against newer NumPy. If a script breaks because of a library API change, you fix it in your own copy. Treat the whole repository as a snapshot to fork from, not a dependency to track.
Editorial conclusion
Adopt ML-From-Scratch if you want to read a working implementation of gradient descent, a CNN, a GAN or an RBM in one sitting, and you are learning rather than deploying. Do not adopt it for anything that has to run fast or on current libraries; it targets old pinned dependencies and its last push was 2023-10-15, with no GitHub releases to fall back on. Before you rely on a single file from it, read the imports in that file and the unpinned requirements.txt, because the library assumes a NumPy and SciPy era that your environment may no longer match.
Frequently asked questions
How do I install ML-From-Scratch?
Clone the repository and run `python setup.py install` from the root, as the README's Installation section shows. The package installs under the name `mlfromscratch` and its dependencies come from requirements.txt, which pins no versions.
How do I run a single model in ML-From-Scratch?
Each example is a standalone script you execute directly, for example `python mlfromscratch/examples/polynomial_regression.py` or `python mlfromscratch/examples/convolutional_neural_network.py`. The library provides the implementations; the example scripts wire up the data and print the results.
Does ML-From-Scratch include deep reinforcement learning?
Yes. There is a deep reinforcement learning example, `python mlfromscratch/examples/deep_q_network.py`, which is presented as a Deep Q-Network solution to the CartPole-v1 environment in OpenAI gym. The requirements also list gym.
What is the license of ML-From-Scratch?
The repository is MIT licensed, and setup.py declares the same MIT license with author Erik Linder-Noren. The intent is that you can read the code and reuse it in your own projects.
Is ML-From-Scratch good for production use?
No. The project states its purpose is to present the inner workings of algorithms in a transparent way rather than to produce optimized, computationally efficient ones. It is written in plain NumPy with no speed optimizations, so it suits reading and learning, not deployment.