ethen8181/machine-learning: A Deep Jupyter Notebook Reference for Practical ML
:earth_americas: machine learning tutorials (mainly in Python3)
At a glance
- What is it?
- ethen8181/machine-learning is a continuously updated collection of Jupyter notebooks that covers machine learning theory, from-scratch implementations in Python, and applied usage of scikit-learn, PyTorch, TensorFlow, and Hugging Face. It is structured as a personal learning log with enough depth to serve as a working reference.
- Who is it for?
- ethen8181/machine-learning suits data scientists and ML engineers who want to see both the mathematical underpinning and a working Python implementation of a given algorithm in one place. It is not an interactive course, a certification path, or a library; it is a reference collection of notebooks.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 83 days ago.
- What is it written in?
- Mainly HTML, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What This Repository Is and How It Is Organized
ethen8181/machine-learning is a personal machine learning study repository that the author describes as documenting a personal journey on learning data science and machine learning topics. The goal, as stated in the README, is to introduce machine learning content in Jupyter Notebook format while striking a balance between mathematical notation, from-scratch educational implementations in Python, and open-source library usage.
The repository is organized by topic, with one top-level directory per domain. These directories include deep_learning/, recsys/ (recommender systems), model_deployment/, operation_research/, reinforcement_learning/, ad/, search/, time_series/, model_selection/, dim_reduct/ (dimensionality reduction), trees/, clustering/, keras/, text_classification/, regularization/, networkx/, association_rule/, big_data/, data_science_is_software/, ga/ (genetic algorithms), unbalanced/ (imbalanced datasets), linear_regression/, and python/ for general Python programming topics.
Each directory contains Jupyter notebooks and in many cases HTML exports of those notebooks. The README provides links to both nbviewer (for interactive rendering) and hosted HTML (for fast static access) for every listed notebook.
The Balance of Theory, Implementation, and Library Usage
The distinguishing characteristic of these notebooks is the three-level treatment they give to most topics. A notebook on softmax regression, for example, covers the mathematical derivation, then builds the algorithm from scratch using numpy, then demonstrates the equivalent using a framework. This approach is more demanding than a pure tutorial but more useful for engineers who need to debug production implementations or adapt an algorithm to an unusual domain.
The libraries used across the collection are extensive. The requirements.txt lists numpy, numba, scipy, pandas, matplotlib, pyspark, scikit-learn, fasttext, huggingface Transformers, onnx, onnxruntime, xgboost, lightgbm, PyTorch, Keras, TensorFlow, gensim, h2o, ortools, and ray. The inclusion of ortools (Google's operations research solver) alongside deep learning libraries reflects the breadth of the repository; it is not exclusively a deep learning notebook collection.
The changelog.md and contributors.md files indicate that the repository has received contributions from others, making it more than a single author's notes.
Setting Up the Environment
The repository does not ship a setup.py or a pyproject.toml. The entry point for reproducibility is requirements.txt, which pins specific versions of every major dependency:
git clone https://github.com/ethen8181/machine-learning.git
cd machine-learning
pip install -r requirements.txtA selection of the pinned versions from requirements.txt gives a sense of the environment:
torch==1.13.1
numpy==1.23.2
pandas==1.5.0
scikit-learn==1.3.0
transformers==4.34.0
lightgbm==3.3.2
xgboost==1.6.1
onnx==1.11.0These pins ensure that the notebooks run with the library behavior they were written for, but they also mean users on newer Python versions or with other projects in the same environment may encounter version conflicts. Using a dedicated virtual environment or conda environment before installing is the safer approach. The convert_to_html.py script in the root generates the static HTML exports linked from the README.
Deep Learning Notebooks: From Softmax to Transformers
The deep_learning/ directory is the most extensive section of the repository. It starts with softmax regression implemented from scratch, then progressively covers feedforward networks in TensorFlow, CNNs for image classification, and recurrent architectures including vanilla RNN and LSTM in both TensorFlow and PyTorch.
Word representation notebooks cover Word2vec with skip-gram and negative sampling using Gensim. The Seq2Seq section includes a PyTorch implementation of sequence-to-sequence translation for German to English, with a separate notebook adding attention, and uses torchtext. A transformer notebook implements Attention Is All You Need from scratch in PyTorch with Hugging Face Datasets, and a separate notebook demonstrates fine-tuning mT5 for machine translation.
The subword tokenization section covers Byte Pair Encoding (BPE) implemented from scratch alongside a sentencepiece walkthrough. Multi-label text classification notebooks use fasttext and Hugging Face Tokenizers. Product Quantization for model compression is documented with a notebook on approximate nearest neighbor search using Navigable Small World (NSW) graphs.
Graph neural networks are covered through a DGL and GraphSAGE notebook on node classification. A question answering notebook fine-tunes a pre-trained encoder. The breadth here is unusual for a personal repository: most single-author ML notebook collections do not reach from softmax regression to GNNs and transformers with implementation detail at each step.
Beyond Deep Learning: Recommender Systems, Operations Research, and Deployment
The model_deployment/ section addresses a practical gap in many ML tutorial collections: how to move from a trained model to a serving environment. The README lists this section, though specific notebooks are not itemized there. The presence of onnx, onnxmltools, and onnxruntime in requirements.txt indicates that at least some deployment notebooks cover the ONNX model format for cross-platform serving.
The recsys/ directory covers recommendation systems, a domain that requires both collaborative filtering theory and engineering knowledge of sparse matrices and approximate nearest neighbor search. The inclusion of the NSW algorithm notebook in the deep learning section suggests the two areas are cross-linked.
The operation_research/ section includes the ortools library, covering scheduling and optimization problems. This is notably outside the machine learning mainstream; its inclusion reflects the repository's goal of covering practical data science broadly rather than only neural network techniques.
Reinforcement learning, time series forecasting, clustering (with both traditional and newer approaches in separate directories), dimensionality reduction, trees (decision trees and ensemble methods), and genetic algorithms all have their own directories. The big_data/ section uses pyspark, meaning some notebooks require a Spark environment to run.
What the Repository Leaves Out
The README makes no claim to being a complete machine learning curriculum. Several significant areas are not listed in the documentation. Large language model fine-tuning beyond the Hugging Face Transformers examples is not mentioned. Diffusion models, which became prominent after many of these notebooks were written, are absent. The repository does not cover model evaluation in production environments, A/B testing frameworks, or feature stores.
The notebooks assume Python familiarity and some mathematical background. The README's stated goal of balancing notation, from-scratch implementation, and library usage means that a reader with no linear algebra background will find some notebooks difficult. The repository is not structured as an introductory course.
The pinned library versions also create a practical limit. torch 1.13.1 is several major versions behind current PyTorch. Notebooks that use TorchText will encounter API changes. Readers who want to adapt these notebooks to current library versions should expect to update import paths, deprecated function calls, and occasionally data loading patterns.
ethen8181/machine-learning vs. Fast.ai Course Notebooks
Fast.ai publishes Jupyter notebooks as part of its Practical Deep Learning for Coders course, also freely available on GitHub. The two repositories take different approaches. Fast.ai notebooks are organized as a sequential course with a consistent narrative and a high-level library (fastai) that simplifies many operations. ethen8181/machine-learning is organized by algorithm or technique, with no prescribed learning order and a preference for showing both the low-level implementation and the library equivalent.
For a learner following a structured curriculum, Fast.ai's sequential course design is easier to follow. For a practitioner who already knows what technique they need and wants to see both the math and working code together, ethen8181/machine-learning provides depth that Fast.ai's abstraction-first approach does not. The two resources are complementary rather than competing for the same use case.
Fast.ai also maintains active course updates aligned with PyTorch releases. ethen8181/machine-learning's last push was on 2026-07-10, reflecting ongoing but less frequent updates. Contributors are credited in contributors.md, and a changelog.md documents major additions to the repository over time.
Editorial conclusion
ethen8181/machine-learning suits data scientists and ML engineers who want to see both the mathematical underpinning and a working Python implementation of a given algorithm in one place. It is not an interactive course, a certification path, or a library; it is a reference collection of notebooks. Users who need an official API reference for scikit-learn or PyTorch should consult those projects' own documentation. The last push was on 2026-07-10. Before cloning for heavy use, verify that the pinned library versions in requirements.txt (including torch 1.13.1 and scikit-learn 1.3.0) are compatible with your environment, since newer library versions may require notebook adjustments.
Frequently asked questions
How do I install the dependencies for ethen8181/machine-learning?
Clone the repository and run pip install -r requirements.txt. The file pins specific versions for torch, scikit-learn, pandas, transformers, and other libraries. Using a dedicated virtual environment is recommended to avoid version conflicts with other projects.
What topics do the machine learning notebooks cover?
The notebooks cover deep learning (from softmax regression to transformers), recommender systems, reinforcement learning, time series, operations research with ortools, model deployment with ONNX, NLP, clustering, dimensionality reduction, and more. Each topic has its own directory in the repository.
Are there HTML versions of the notebooks available without running Jupyter?
Yes. The README links to both nbviewer (interactive) and static HTML exports for each notebook. The convert_to_html.py script in the repository root generates the HTML files.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/ethen8181-machine-learning)