Model or dataset
amitshekhariitbhu/build-your-own-x-machine-learning avatar
amitshekhariitbhu/build-your-own-x-machine-learning

Build Your Own X: Machine Learning is a from-scratch tutorial collection, not a library

Build your own X - Master machine learning by building everything from scratch. It aims to cover everything from linear regression to deep learning to large language models (LLMs).

707 stars110 forksPythonApache-2.0

At a glance

What is it?
Amit Shekhar's repository collects standalone Python implementations of algorithms from linear regression to neural networks. It is a reading and typing exercise, not a package you install into a production pipeline.
Who is it for?
Adopt this repository if you learn by writing gradient descent, KNN or a perceptron by hand and you want a curated index of single-file Python references to read alongside your own code. Do not adopt it if you need a maintained library with tests, versioned releases or support, because the repository is a set of tutorials and the README documents no such guarantees.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 88 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

What problem the repository actually solves

Most people meet machine learning through an estimator object. You call fit, you call predict, and the arithmetic happens inside a package you never open. That is the right way to ship software. It is a poor way to understand why a model fails on a particular dataset, because the failure lives in a line of calculus you have never written.

This repository addresses that gap. It is a collection of tutorials, each a single Python file, that reimplements one algorithm without a machine learning framework underneath. The README lists entries such as linear regression, logistic regression, K-Nearest Neighbors, Naive Bayes, decision trees, random forests, support vector machines, K-Means, PCA, the perceptron, gradient descent, gradient boosting, AdaBoost, LDA, Ridge and Lasso, polynomial regression, ElasticNet, Bayesian regression, mean-shift clustering, spectral clustering, independent component analysis and factor analysis, plus cost functions, activation functions and optimizers. The stated aim is to cover everything from linear regression to deep learning to large language models.

The intended reader is someone who already writes Python and wants the derivations in code form. It is not aimed at a team that needs a model in production this quarter.

How the repository is organised and how a tutorial runs

The top level holds LICENSE, README.md, assets/ and tutorials/. The README is the index, and every entry is a link to a file inside tutorials/, grouped into sections such as Core Machine Learning Algorithms, Neural Networks and Deep Learning, Recommendation Systems, Computer Vision Applications, Natural Language Processing, Time Series and Forecasting, Anomaly Detection, Sentiment and Text Analysis, and Miscellaneous Applications.

The data flow is deliberately flat. A tutorial file is a script. It defines the algorithm, often as a class with fit and predict methods, and then runs a small example at the bottom so you can execute it and watch numbers appear. There is no shared core package, no plugin interface and no configuration layer connecting one tutorial to another. Each file is meant to be read on its own.

That flatness is the point and also the cost. You can open linear_regression.py and follow the whole thing in one sitting, which is exactly what a learner needs. You cannot import a common base class from one tutorial into another, because the repository does not present one. The README does not describe a shared module layout, so treat each file as self-contained.

Installing the repository and running your first tutorial

There is no package to install. The repository is not published as a library, and the README gives no installation command. You get the code by cloning it, and you run the tutorial files directly with Python.

Start by cloning the default branch:

bash
git clone https://github.com/amitshekhariitbhu/build-your-own-x-machine-learning.git
cd build-your-own-x-machine-learning

The README links each tutorial by path. Linear regression lives under tutorials/core-machine-learning-algorithms/linear-regression/. Change into that directory and run the script:

bash
cd tutorials/core-machine-learning-algorithms/linear-regression
python linear_regression.py

The README does not list a requirements file or a pinned dependency set, so the imports inside each script are the only signal about what you need. If a file fails on a missing module, install that module yourself; the repository does not tell you which version it was written against. Expect a script to print its intermediate values, such as a loss that decreases across iterations or a set of learned coefficients, because that printed trace is the teaching device. If you see nothing, you are probably running a file that only defines classes and has no example block at the bottom.

Where the tutorial format breaks down

The most important limitation is that a tutorial is not a library. Nothing in the README claims these implementations are numerically stable, vectorised for speed, or validated against a reference. A from-scratch SVM written for readability will not match the convergence behaviour of a production solver, and the repository does not present it as a drop-in replacement for one.

A second limitation is scope in the other direction. The README states the aim of covering deep learning and large language models, but the visible index is dominated by classical algorithms and the linked files are single-script demonstrations. Anyone arriving for a from-scratch transformer with training infrastructure should check the Neural Networks and Deep Learning section of the README before assuming the material is there.

The third is maintenance surface. There are no releases, so there is no version to pin and no changelog to read. The last push was on 2026-07-04, and the README says new tutorials will keep being added. That means paths and file lists can change under you. If you copy a tutorial into your own project, copy the file rather than depending on the repository layout, because the layout is the part most likely to move.

How this differs from scikit-learn and from a course

scikit-learn is the obvious alternative, and the difference is not quality. scikit-learn is a library: it exposes a stable estimator API, it is tested, and it is built to be imported. This repository is a set of teaching scripts that you read and run. The trade-off is direct. Choosing scikit-learn gets you correct, fast implementations and no insight into the internals. Choosing this repository gets you the internals and nothing you should put behind a production endpoint.

A structured online course is the other comparison. A course gives you a sequence, exercises and a schedule. This repository gives you an index and a folder of files, and the README is explicit that it is maintained by Amit Shekhar, the founder of Outcome School, which also sells AI and machine learning and Android programs. The tutorials are free to read; the course is the paid path. If you need someone to tell you what to do next, a course fits better. If you already know what you want to understand and just need the code in front of you, the repository is faster.

Licence and the cost of keeping up

The repository is licensed under Apache-2.0, and the LICENSE file sits at the top level. That permits commercial use and modification, and it includes a patent grant, which permissive licences such as MIT do not. It also carries notice and attribution conditions, so if you copy a tutorial file into your own codebase, keep the licence and attribution intact. This is a description of the terms, not legal advice; read the LICENSE file before you redistribute anything.

Upgrade cost is low in one sense and unpredictable in another. There is nothing to upgrade, because there are no releases and no dependency manifest. But there is also no compatibility promise. If a tutorial imports a library that later changes its API, the repository has no stated process for catching that, and the README does not document one. Budget for reading the file and fixing the import yourself rather than waiting for a patch.

Editorial conclusion

Adopt this repository if you learn by writing gradient descent, KNN or a perceptron by hand and you want a curated index of single-file Python references to read alongside your own code. Do not adopt it if you need a maintained library with tests, versioned releases or support, because the repository is a set of tutorials and the README documents no such guarantees. Before you start, verify two things: that the specific tutorial file you plan to study exists at the path linked in the README, and that its imports match the Python and NumPy versions on your machine, since the README does not document a dependency pin or a supported version range.

Frequently asked questions

What is Build Your Own X - Machine Learning?

It is a repository of Python tutorial files that implement machine learning algorithms from scratch, from linear regression through to neural networks and beyond. The README describes it as a way to master machine learning by building everything from scratch, and the files live under the tutorials/ directory.

How do I build my own machine learning model with this repository?

Clone the repository, open the tutorial folder for the algorithm you want, and run the Python file directly, for example python linear_regression.py inside tutorials/core-machine-learning-algorithms/linear-regression. The README does not provide a package install, so the code is read and executed as standalone scripts.

Is there a PDF version of build-your-own-x-machine-learning?

The README does not mention a PDF. The material is published as Python files inside the tutorials/ directory of the repository, indexed by links in the README.

Can I use these implementations in a production system?

The repository presents them as tutorials, not as a library, and the README makes no claims about testing, numerical stability or performance. There are no releases to pin, so there is no version of these implementations to depend on.

What licence does build-your-own-x-machine-learning use?

It is licensed under Apache-2.0, with the LICENSE file at the top level of the repository. That permits commercial use and modification and includes a patent grant, subject to the notice and attribution conditions in the licence.

Official sources

  1. amitshekhariitbhu/build-your-own-x-machine-learning on GitHub
  2. Issues
  3. License: Apache-2.0
  4. Project website
  5. README
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/amitshekhariitbhu-build-your-own-x-machine-learning.svg)](https://hysenlabs.com/projects/amitshekhariitbhu-build-your-own-x-machine-learning)
Community notes

Community notes