Model or dataset
stefan-jansen/machine-learning-for-trading avatar
stefan-jansen/machine-learning-for-trading

machine-learning-for-trading: What the 3rd Edition Repository Actually Ships

Code for Machine Learning for Trading, 3rd edition, from data sourcing to live execution.

21,078 stars5,648 forksJupyter NotebookMIT

At a glance

What is it?
Stefan Jansen's third-edition codebase runs nine case studies through one research pipeline, from Polars data handling to live execution. The install path is documented, the Python floor is 3.14, and one image is amd64-only.
Who is it for?
Adopt this repository if you already write Python and want one worked pipeline across nine markets rather than a pile of disconnected model demos. Skip it if you need a plug-and-play trading system, a maintained package you can import into production, or anything that runs on Python 3.13 or below.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 5 days ago.
What is it written in?
Mainly Jupyter Notebook, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What machine-learning-for-trading Is For

This is the code companion to Machine Learning for Trading, 3rd Edition, and it is organized around a single workflow rather than around techniques. The README describes the path as going from data infrastructure and strategy research, across what it calls an evidence boundary that separates tuning from evaluation, to deployment and monitoring, with a feedback loop that retrains, pauses, or retires a strategy as its edge decays. That framing matters for who should read it. The repository assumes you want to understand how a research idea becomes a sized, cost-aware, risk-managed portfolio, not that you want a library to call.

The audience is closer to a quantitative researcher or a graduate student than to a discretionary trader. Twenty-seven chapter directories sit at the top level, from 01_process_is_edge through 27_systematic_edge, and nine case studies run the same pipeline over ETFs, crypto perpetuals, NASDAQ-100 intraday data, S&P 500 equities with options, US firm characteristics, FX pairs, CME futures, S&P 500 options, and a US equities panel. The point of repeating one process nine times is to show where it holds up and where it does not.

Two things distinguish the third edition from earlier ones in the same repository lineage. Transaction costs and risk management are now full chapters of their own, which the README notes did not exist before. And the data layer moved to Polars for expression-based manipulation, with PyTorch, LightGBM, Optuna, and Plotly rounding out the stack. Generative AI material is new as well: retrieval-augmented generation grounded in SEC filings, knowledge graphs, and multi-agent research systems occupy chapters 22 through 24.

How the Repository Is Organized

The top-level layout is a map of the book. Numbered directories correspond to chapters, case_studies/ holds the nine end-to-end studies, and utils/, tests/, scripts/, envs/, and docs/ carry supporting code. There is a pyproject.toml, a uv.lock, a docker-compose.yml, and a matplotlibrc. The primary language is Jupyter Notebook, so most of what you read is a notebook rather than an importable module.

That has a practical consequence. You cannot pip install this repository as a trading framework. The pyproject.toml declares a project named ml4t at version 3.0.0, but its purpose is to pin the environment the notebooks expect, not to publish a stable API. Dependencies are pinned with reasons in comments, which is unusual and useful: scipy is held below 1.18 because version 1.18 removed a private scipy.cluster.hierarchy linkage attribute that PyPortfolioOpt 1.6.0 still references for HRP linkage validation in chapter 17. kaleido is held at 1.3 or above because 0.2.1 bundled its own Chromium and drove a deprecated plotly.io.kaleido.scope API.

There is a verification artifact worth knowing about: a .verified-notebooks.tsv file at the top level. Combined with the v3.0.0-artifacts release, which the release notes describe as pre-computed case study artifacts, the repository gives you a way to compare your own runs against committed outputs instead of trusting that a notebook executes cleanly on your machine.

Installing machine-learning-for-trading and Running a First Notebook

The README points to docs/installation.md as the entry point and says it walks a blank Linux, Windows or macOS machine to a running notebook, prerequisites included. The Docker route is the one the repository documents most concretely. Pre-built images live on Docker Hub under docker.io/ml4t/, and the compose file gives the two commands directly.

The pull is large. The compose comments put the ml4t image at roughly 12 GB on amd64 and roughly 3 GB on arm64, so budget disk and time before starting.

Platform Limits and the Python 3.14 Floor

The compose file is unusually explicit about what does not work everywhere, and it is worth reading before you build anything. Linux x86_64 supports all services with optional GPU. Windows x86_64 runs everything through Docker Desktop on WSL2, with GPU access via WSL2 and the nvidia toolkit. macOS Intel supports all services except GPU.

Apple Silicon is where it narrows. The ml4t and benchmark images are native arm64, but ml4t-py312 is amd64 only and therefore runs under Rosetta emulation. That image is not optional for everyone: the compose comments list chapter 5 notebooks 01, 03 and 07, chapter 9 notebooks 06 and 12, chapter 10 notebooks 01 through 03, chapter 12 notebook 10, chapter 14 notebook 06, and chapter 15 notebook 06 as requiring it, because those notebooks depend on signatory, esig, gensim, and tfcausalimpact. On Apple Silicon the documented options are reading the committed .ipynb outputs or running under Rosetta with DOCKER_DEFAULT_PLATFORM=linux/amd64.

The second constraint is the interpreter version. pyproject.toml sets requires-python to >=3.14,<3.15. That is a hard floor and a hard ceiling in one line. If your environment is pinned to 3.12 or 3.13 for other reasons, the declared dependency set will not resolve, and the ml4t-py312 image exists precisely because some libraries have not caught up. A GPU benchmark for gradient boosting in chapter 12 needs the rapids service, which the compose file says requires an NVIDIA GPU. None of this is hidden, but all of it is the kind of thing you discover after a 45-minute local build if you skip the comments.

Where This Repository Is the Wrong Tool

The most important limitation is structural: this is teaching code, not a trading system. Nothing in the repository describes a supported API surface, a release cadence for a library, or a compatibility promise. The version number 3.0.0 tracks the book edition, and the recent releases are labeled v3.0.0-alpha.2, v3.0.0-alpha.3, and v3.0.0-artifacts, which is a book's release naming rather than a library's.

If you want to run a strategy in a live account next week, this is the wrong starting point. Chapter 25 covers live trading systems for Interactive Brokers, Alpaca, and QuantConnect, but that is instructional material about wiring execution, not a broker abstraction you can depend on. Treating a notebook that places an order as production code is how people lose money.

The second limitation is cost of entry. A 12 GB image, a 45-minute local build, a Python 3.14 requirement, and an amd64-only fallback image for specific chapters add up. If your goal is to test one model on one CSV, this repository is a poor fit; the surrounding apparatus is most of the work.

The third is methodological. The README states that the book confronts multiple-testing and overfitting problems with tools like the Deflated Sharpe Ratio, the Rademacher Anti-Serum, and White's Reality Check. Those tools exist because naive backtests are misleading. A reader who lifts a single notebook's results without the walk-forward cross-validation and the evidence boundary the book describes will get exactly the false confidence the book is written to prevent.

How It Compares to Backtrader and Vectorbt

The obvious alternative for someone who wants to backtest is a dedicated framework such as Backtrader or Vectorbt. The difference is in what each one is. A backtesting framework gives you an event loop or a vectorized engine, a strategy class, a data feed interface, and a report. You supply the signal. The framework owns the execution model.

This repository inverts that. It owns the research process and treats execution as one stage among many. Chapters 16 through 20 move from strategy simulation to portfolio construction, transaction costs, risk management, and strategy synthesis, and the case studies carry each market through that whole sequence. You get labels, features, models, backtests, cost modeling, risk overlays, and a deployment assessment in one narrative. What you do not get is a stable Strategy base class to subclass.

The practical split: use Backtrader or Vectorbt when you already have a signal and need a tested engine to evaluate it against a broker's fill model. Use machine-learning-for-trading when the open question is how to define the learning task, which features survive, and whether a result is real. The overlap is the backtest, and the repository's backtests are illustrative rather than a reusable engine. If you adopt it expecting the second thing, you will spend a week discovering you have notebooks.

Maintenance, Licence, and What Upgrades Cost

The repository is not archived, and the last push was on 2026-07-24, which is under two months before the date of writing. The v3.0.0-artifacts release carries the same timestamp, and the two alpha releases before it landed on 2026-06-29 and 2026-06-22. That is a book release cycle: bursts of activity tied to edition milestones, not continuous library maintenance. Plan around that rather than expecting patches.

The licence is MIT, declared both in the repository root and in pyproject.toml. MIT is permissive, so you can reuse the code in your own work, including commercially, provided you keep the copyright notice. That is a statement about the licence text, not legal advice; if you are folding this code into a product, read the LICENSE file and get your own counsel.

The upgrade cost is real and it is concentrated in the pins. pyproject.toml holds scipy below 1.18 with a comment explaining that PyPortfolioOpt 1.6.0 still references a private scipy attribute, and holds kaleido at 1.3 or above because plotly 6.x deprecates the older API. Those comments tell you the pins are load-bearing. When PyPortfolioOpt ships a scipy 1.18 fix, the scipy constraint can lift; until then, upgrading scipy breaks chapter 17. The same pattern applies to requires-python. Moving the floor is not a one-line change when several dependencies have not shipped 3.14 support, which is why the ml4t-py312 image exists at all. Budget for reading the pin comments before any dependency bump.

Editorial conclusion

Adopt this repository if you already write Python and want one worked pipeline across nine markets rather than a pile of disconnected model demos. Skip it if you need a plug-and-play trading system, a maintained package you can import into production, or anything that runs on Python 3.13 or below. Before committing, verify that your platform matches the image you need: ml4t-py312 is amd64 only, so on Apple Silicon you either read the committed notebook outputs or run it under Rosetta with DOCKER_DEFAULT_PLATFORM=linux/amd64. Then check the pinned scipy<1.18 in pyproject.toml, which exists because PyPortfolioOpt 1.6.0 still references a private scipy hierarchy attribute.

Frequently asked questions

Can machine learning be used for trading?

The repository is built on the premise that it can, and its nine case studies carry signals from raw data through features, models, backtests, costs, and risk overlays to a deployment assessment. It also treats methodological rigor as a first-class topic, using walk-forward cross-validation and tools such as the Deflated Sharpe Ratio and White's Reality Check to confront the multiple-testing and overfitting problems that the README says quietly invalidate most backtests.

How do I use machine learning for trading with this repository?

Start at docs/installation.md, which the README says walks a blank Linux, Windows or macOS machine to a running notebook. The Docker route pulls a pre-built image and starts Jupyter Lab on port 8888, and the repository notes that the default settings in .env.example work as-is for the early chapters.

What is machine-learning-for-trading?

It is the code repository for Machine Learning for Trading, 3rd Edition by Stefan Jansen, organized around one end-to-end workflow from research idea to live execution. It contains 27 chapter directories, nine case studies across different asset classes, and reproducible Docker environments.

Which AI is best for learning trading with this codebase?

The repository does not rank tools. It ships a wider model toolkit spanning gradient boosting with XGBoost, LightGBM and CatBoost, deep time-series architectures such as PatchTST, iTransformer, TSMixer, TCN and Mamba, and newer tabular models including TabPFN and TabM. Which one suits a given case study is what the nine case studies are for.

Can I use AI to run my trading with this repository?

Chapter 25 covers live trading systems for Interactive Brokers, Alpaca, and QuantConnect, and chapters 22 through 24 cover retrieval-augmented generation, knowledge graphs, and autonomous agents. The README frames these as instructional material within a research workflow, and the repository does not document a supported production API for placing orders.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
  4. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/stefan-jansen-machine-learning-for-trading.svg)](https://hysenlabs.com/projects/stefan-jansen-machine-learning-for-trading)