Self-hosted service
rushter/MLAlgorithms avatar
rushter/MLAlgorithms

rushter/MLAlgorithms: From-Scratch Machine Learning Implementations in Python

Minimal and clean examples of machine learning algorithms implementations

11,176 stars1,767 forksPythonMIT

At a glance

What is it?
MLAlgorithms is a Python package of minimal, readable implementations of classic machine learning algorithms, built with numpy, scipy and autograd. It is a study resource for engineers and students who want to understand how gradient boosting, SVMs or LSTMs work at the code level, not a library for training production models.
Who is it for?
Engineers who want to study how k-nearest neighbors, gradient boosting or an LSTM works in plain Python will find MLAlgorithms worth cloning. It is not the right tool for training real models on real data: the dependencies are fixed at old floor versions (numpy>=1.11.1, scipy>=0.18.0 per requirements.txt), there is no GPU path, and the package version remains 0.0.1.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 146 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 17, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Why MLAlgorithms Exists and Who It Targets

The rushter/MLAlgorithms repository is designed for one audience: engineers and students who want to read how a machine learning algorithm actually works in Python code. The project README states directly that the code is "much easier to follow than the optimized libraries and easier to play with." Production libraries like scikit-learn rely on C and Cython extensions that eliminate Python overhead at the cost of readability. MLAlgorithms makes the opposite choice: every algorithm is written in pure Python using numpy and scipy, so the data flow from input matrix to output prediction is visible at the source level.

The package is named mla and is registered in setup.py at version 0.0.1. The project targets two groups: people who want to understand the internals of algorithms they already use in production, and people who want to implement an algorithm from scratch as a learning exercise. It is not a general-purpose tool for building pipelines. Anyone looking for cross-validation utilities, preprocessing helpers or hyperparameter tuning should use a different library. The gap MLAlgorithms fills is genuine: most ML textbooks explain algorithms in pseudocode or mathematics, and most production libraries hide the key steps behind compiled code. A readable Python implementation sits between these two ends and serves a specific pedagogical purpose.

The mla/ Package Structure and Its numpy-Based Architecture

The repository is organized around a core Python package called mla, located in the mla/ directory. Each algorithm maps to its own file or a small subdirectory. Linear regression and logistic regression are in mla/linear_models.py. The gradient boosting implementation is in mla/ensemble/gbm.py (the README notes this covers GBDT, GBRT, GBM and XGBoost). The SVM with three kernel options (Linear, Poly, RBF) lives in mla/svm. Deep learning variants are grouped in mla/neuralnet, covering MLP, CNN, RNN and LSTM. Reinforcement learning via Deep Q-Learning is in mla/rl.

The requirements.txt lists all external dependencies: tqdm, matplotlib>=1.5.1, numpy>=1.11.1, scikit-learn>=0.18, scipy>=0.18.0, seaborn>=0.7.1, autograd>=1.1.7 and gym. The autograd library handles automatic differentiation, which is used in the neural network implementations instead of manual gradient derivations. The gym dependency supports the reinforcement learning environment for the Deep Q-Learning example. A separate examples/ directory holds a standalone runnable script for every algorithm, named predictably: examples/gbm.py, examples/svm.py, examples/nnet_mlp.py. The repository also includes a Dockerfile based on python:3 that installs scipy, numpy and then the mla package itself.

Installing MLAlgorithms and Running the First Example

The README offers three paths. The standard development install clones the repository and installs it with setup.py:

sh
git clone https://github.com/rushter/MLAlgorithms
cd MLAlgorithms
pip install scipy numpy
python setup.py develop

After setup.py develop completes, the mla package is importable from the current environment. To run any algorithm example without the full install, the README gives a direct module invocation:

sh
cd MLAlgorithms
python -m examples.linear_models

This runs the linear models example file directly without requiring setup.py. The Docker path is also documented for engineers who prefer an isolated environment:

sh
cd MLAlgorithms
docker build -t mlalgorithms .
docker run --rm -it mlalgorithms bash
python -m examples.linear_models

The Dockerfile builds from python:3, copies the project directory, installs scipy and numpy, and then installs the mla package via pip. Inside the container the same module invocation works. Each example script is self-contained: it imports the algorithm class from mla, generates or loads a dataset, fits the model and prints or plots the result. There is no configuration file to set up before running.

What the Codebase Covers: Fourteen Algorithm Families

The README lists fourteen implementation areas. Supervised learning covers deep learning (MLP, CNN, RNN, LSTM in mla/neuralnet), linear regression and logistic regression (mla/linear_models.py), random forests (mla/ensemble/random_forest.py), SVM with three kernels (mla/svm) and k-nearest neighbors (mla/knn.py). Probabilistic methods include naive Bayes (mla/naive_bayes.py) and Gaussian Mixture Models (mla/gaussian_mixture.py). Unsupervised methods include K-Means (mla/kmeans.py) and PCA (mla/pca.py). More specialized implementations appear as well: factorization machines (mla/fm.py), Restricted Boltzmann Machine (mla/rbm.py), t-SNE (mla/tsne.py), gradient boosting trees (mla/ensemble/gbm.py) and Deep Q-Learning (mla/rl).

The examples directory mirrors this list with a one-to-one mapping of example scripts, including examples for RNN text generation (examples/nnet_rnn_text_generation.py) and an MNIST convnet (examples/nnet_convnet_mnist.py). That breadth is notable for a single repository: few reference implementations cover everything from naive Bayes to reinforcement learning in one install. The trade-off is that each implementation is minimal. The gradient boosting implementation is a teaching version of the algorithm, not a competitive alternative to XGBoost or LightGBM.

Where MLAlgorithms Is the Wrong Tool

MLAlgorithms is a study aid, and several design decisions make it unsuitable outside that context. The version in setup.py is 0.0.1, and the repository has no GitHub releases. The dependency floor is old: numpy>=1.11.1 and scipy>=0.18.0 were current in 2016, and the autograd library has seen no major release since then either. There is no test suite visible in the repository layout (no tests/ directory appears in the top-level entries), so no automated check guards against regressions.

There is no GPU support. The neural network examples use numpy arrays directly and run on CPU only. For the MNIST convnet example or any meaningful deep learning experiment, this becomes a significant constraint. The implementations also do not expose the full API surface that production users expect: no partial_fit for incremental learning, no sample_weight for weighted datasets, no pipeline integration, and no serialize/deserialize for saving trained models. Finally, scikit-learn integration is listed in requirements.txt, but it is there for dataset utilities in the examples, not to make the mla algorithms interoperable with scikit-learn pipelines.

MLAlgorithms vs scikit-learn: A Difference in Purpose

scikit-learn is the standard Python machine learning library and covers most of the same algorithms: SVM, random forests, gradient boosting, k-nearest neighbors, naive Bayes, PCA, k-means and more. The key difference is that scikit-learn is built for users of algorithms, while MLAlgorithms is built for readers of algorithms. scikit-learn's gradient boosting estimator calls into Cython routines that are fast but opaque. MLAlgorithms' equivalent in mla/ensemble/gbm.py is entirely in Python and can be stepped through with a debugger line by line.

The performance difference is substantial and intentional. scikit-learn's implementations are optimized for production workloads with thousands of features and millions of rows. MLAlgorithms makes no performance claims at all. For someone trying to understand why gradient boosting uses residuals as the target for each successive tree, reading mla/ensemble/gbm.py alongside the theory is more useful than reading scikit-learn's Cython. For someone training a model for deployment, scikit-learn, XGBoost or LightGBM is the correct choice. The two libraries answer different questions and do not compete.

Last Push Date and MIT License

The last push to the repository was on 2026-05-07. There are no GitHub releases. The version number in setup.py is 0.0.1, where it has remained since the project was published. The repository has been open to contributions, and the README welcomes pull requests and issue-based proposals, but the maintainer cadence for merging contributions is not documented.

The license is MIT. This imposes no restrictions on commercial or private use, requires only that the license text and copyright notice be retained in any copy or substantial portion of the software. The author is listed as Artem Golubin in setup.py with contact email [email protected]. Because the package version is 0.0.1 and there are no releases, any downstream project that pins to this repository should expect no version-level API guarantees. The safest approach is to clone and develop against a specific commit rather than depending on an npm-style versioned release.

Editorial conclusion

Engineers who want to study how k-nearest neighbors, gradient boosting or an LSTM works in plain Python will find MLAlgorithms worth cloning. It is not the right tool for training real models on real data: the dependencies are fixed at old floor versions (numpy>=1.11.1, scipy>=0.18.0 per requirements.txt), there is no GPU path, and the package version remains 0.0.1. Clone the repository, run python -m examples.gbm to see gradient boosting in action, and read mla/ensemble/gbm.py against the corresponding scikit-learn source to understand exactly which steps the production library abstracts away.

Frequently asked questions

What algorithms does rushter/MLAlgorithms implement?

The repository implements fourteen algorithm families including MLP, CNN, RNN and LSTM neural networks, linear and logistic regression, random forests, SVM with three kernels, k-means, Gaussian mixture models, k-nearest neighbors, naive Bayes, PCA, factorization machines, RBM, t-SNE, gradient boosting trees and Deep Q-Learning. Each has a corresponding example script in the examples/ directory.

What Python dependencies does MLAlgorithms require?

The requirements.txt lists tqdm, matplotlib>=1.5.1, numpy>=1.11.1, scikit-learn>=0.18, scipy>=0.18.0, seaborn>=0.7.1, autograd>=1.1.7 and gym. The setup.py also enforces numpy>=1.10 and scipy>=0.17 as setup requirements. These version floors are old and the project has no declared upper bounds.

Can MLAlgorithms be run inside Docker?

Yes. The repository includes a Dockerfile based on python:3 that installs scipy, numpy and the mla package. After building with docker build -t mlalgorithms . and starting a container with docker run --rm -it mlalgorithms bash, examples can be run with the same python -m examples.linear_models command used in the standard install.

Official sources

  1. Issues
  2. License: MIT
  3. README
  4. rushter/MLAlgorithms on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/rushter-mlalgorithms.svg)](https://hysenlabs.com/projects/rushter-mlalgorithms)