AlphaPy: a machine learning pipeline for speculators, now a legacy repository
Python AutoML for Trading Systems and Sports Betting
At a glance
- What is it?
- Two ready made pipelines for market and sports prediction, kept alive under Apache-2.0 while the README steers new work toward a separate commercial edition with a different Python baseline.
- Who is it for?
- AlphaPy is worth reading as a worked example of how a full machine learning pipeline gets laid out for financial data, because the code shows the whole chain from data pull to blended ensemble to portfolio tear sheet in one place, and the docs at alphapy.readthedocs.io explain the parts the README skips. It is not worth installing into anything you intend to keep running.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 4 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 8, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Two pipelines, one shared dependency pile
AlphaPy is two applied workflows sharing a common feature engineering and model training core. MarketFlow handles market data, models and trading systems; SportFlow handles sporting event prediction. The README is blunt about who this is for, describing the package as a machine learning framework for speculators and data scientists, written mainly with scikit-learn and pandas plus a longer list of feature engineering and plotting libraries.
The model side is broad rather than deep. You can run scikit-learn estimators, Keras networks, and XGBoost, LightGBM and CatBoost, and the README advertises blended and stacked ensembles as a headline capability. That is a wider surface than most quant libraries offer, which lean on one or two model families, so the interesting comparison is with general AutoML tooling rather than with a trading-specific stack.
The repository description says Python AutoML for Trading Systems and Sports Betting. The topics list runs to twenty entries, covering backtesting, portfolio, keras, xgboost-era machine learning terms, sports and stocks. Licensed Apache-2.0, 1763 stars, 279 forks, 16 open issues.
Installing the legacy package and its optional extras
The install command in the README is deliberately plain:
pip install -U alphapyThree of the model libraries do not come along automatically. XGBoost will not install with pip on Mac or Windows, so the README sends you to the XGBoost build documentation rather than pretending there is a one-liner. LightGBM and CatBoost each get their own installation link for the same reason.
The one genuinely instructive install note is about pyfolio. AlphaPy pulls pyfolio in for portfolio tear sheets, and the README documents a specific failure you may hit:
*AttributeError: 'numpy.int64' object has no attribute 'to_pydatetime'*and prescribes installing from source instead:
pip install git+https://github.com/quantopian/pyfolioThat is a workaround for a version skew between pyfolio and a newer NumPy, and it tells you something about the package's age. A framework that documents an upstream library's incompatibility as a permanent install step is one whose dependency set has drifted.
The README also points out that you need pip and Python already present, and treats XGBoost, LightGBM and CatBoost as optional rather than required.
The pipeline diagrams come from a different repository
Every diagram in the README is served from a different GitHub account. The model pipeline, the market pipeline, the system pipeline and the sports pipeline are all referenced as raw files under github.com/Alpha314/AlphaPy, on the `master` branch. This repository is ScottfreeLLC/AlphaPy, and its default branch is `main`. The support link at the bottom of the README points at ScottfreeLLC/AlphaPy/issues, so the project is clearly aware of its own canonical location.
Both facts hold at once. The diagrams render, so the image URLs are live, and they come from a mirror or fork under a different owner. The consequence is that the visual documentation is not versioned alongside the code you are reading. When a diagram and the implementation disagree, there is no commit in this repository that explains which one is current, and the branch name mismatch makes it obvious that the images predate the move to `main`.
For a project whose selling point is a pipeline you are meant to follow end to end, having the architecture pictures live outside the repository is the single most surprising thing about the setup. Worth knowing before you spend an afternoon mapping the module tree yourself.
Version 2.5.0 since 2020, with pushes continuing in 2026
The release history has three entries and stops. 2.4.0 landed in February 2020 with feature names in the importance plots, updated categorical encoders, confusion matrices showing counts alongside percentages, and an upgrade to pandas 1.0 and scikit-learn 0.22. 2.4.3 followed in August 2020 as an ecosystem upgrade. 2.5.0, also August 2020, added LightGBM and CatBoost support. Nothing has shipped since.
Against that, `setup.py` still declares `VERSION = '2.5.0'`, and the repository was pushed on 2026-10-05. Both things are true: the packaged version is six years old and the repository is receiving commits. If you need to know whether a fix exists, the push history is the place to look, not the release page.
The classifier metadata makes the age concrete. The package declares support for Python 3.7 and 3.8 and `Development Status :: 4 - Beta`. Those are EOL interpreter versions, and the README's own answer to that problem is the AlphaPy Pro section, which advertises modern Python 3.12 or newer with UV package management, MetaLabeling support, NLP features and automated CI/CD through GitHub Actions. The open source version and the promoted version therefore disagree about what a supported Python looks like, by four minor releases in each direction.
Why active development is declared to have moved
The README does not bury this. After the feature list comes an announcement that AlphaPy Pro is publicly available, a separate repository at github.com/ScottfreeLLC/alphapy-pro, a documentation site hosted from GitHub Pages, and the install command `pip install alphapy-pro`. There is then an explicit note: active development has moved to AlphaPy Pro, and this repository remains available for users who rely on the original version.
That framing is what you should evaluate this repository through. The Pro edition's feature list is a de facto roadmap of what the legacy version will not gain, and it is revealing about priorities: MetaLabeling for financial modeling, NLP features for sentiment analysis, and CI/CD automation, alongside the Python version bump. Those are all real gaps in a 2020-era pipeline framework.
So the two signals pull in opposite directions and both are informative. A repository with a commit in the last day is one someone is still touching, whether for maintenance, for documentation, or for a dependency fix. A README that points at a paid successor and labels itself legacy is one you should not plan a dependency on. Treat the recent push as evidence of occasional care, not as evidence of a maintained release line.
What the dependency list reveals about running this today
`setup.py` lists eighteen install requirements with a mix of floors. Most are open-ended: `arrow>=0.13`, `bokeh>=1.3`, `category_encoders>=2.1`, `imbalanced-learn>=0.5`, `matplotlib>=3.0`, `numpy>=1.17`, `pandas>=1.0`, `scikit-learn>=0.23.1`, `seaborn>=0.9`, `tensorflow>=2.0`. Two deserve a second look.
The first is `scipy==1.10.0`, the only exact pin in the entire list. Everything else allows a newer release, and this one does not, which in practice caps your SciPy version no matter what else you upgrade. Pinning a single scientific library hard inside an otherwise floating set is a smell: it usually means someone worked around a specific break rather than fixing the code.
The second is `iexfinance>=0.4.3`, which is the market data path. AlphaPy also depends on `pyfolio>=0.9` and `pandas-datareader>=0.8` for portfolio analysis and data access, and the README's pyfolio workaround points at the Quantopian repository specifically. When a financial library's data providers and its tear sheet library both trace back to services whose corporate arrangements have changed since 2020, the honest move is to verify each one resolves before trusting any backtest output. The dependency names are all still listed, but a listed dependency and a working dependency are different claims, and only the second one matters for a model that trades.
Editorial conclusion
AlphaPy is worth reading as a worked example of how a full machine learning pipeline gets laid out for financial data, because the code shows the whole chain from data pull to blended ensemble to portfolio tear sheet in one place, and the docs at alphapy.readthedocs.io explain the parts the README skips. It is not worth installing into anything you intend to keep running. The version has sat at 2.5.0 since August 2020, the metadata still advertises Python 3.7 and 3.8, one dependency is pinned to an exact SciPy version while everything else floats, and two of the external services the package leans on are worth verifying before you trust a backtest. The README itself says active development has moved to AlphaPy Pro, which requires Python 3.12 or newer. Take the pipeline ideas, and check each data source and library status yourself.
Frequently asked questions
Is AlphaPy still maintained, or should I use AlphaPy Pro?
The README states that active development has moved to AlphaPy Pro and that this repository stays available for users relying on the original version. The open source line's last release was 2.5.0 in August 2020, even though the repository was pushed on 2026-10-05. AlphaPy Pro is a separate package requiring Python 3.12 or newer, so the choice is roughly legacy and free versus current and proprietary.
What is the difference between MarketFlow and SportFlow in AlphaPy?
They are two application pipelines built on the same feature engineering and model training core. MarketFlow covers market data, predictive models, trading system development and portfolio analysis, while SportFlow covers prediction of sporting event outcomes. AlphaPy Pro is described as having an enhanced MarketFlow, with MetaLabeling and NLP sentiment features as additions.
Which machine learning models can AlphaPy use?
scikit-learn estimators, Keras models, and the three gradient boosting libraries XGBoost, LightGBM and CatBoost. The README also advertises blended and stacked ensembles. LightGBM and CatBoost support arrived in release 2.5.0 in August 2020, and the three of them are optional extras that may need platform-specific installation rather than coming in with `pip install -U alphapy`.
Why does creating a pyfolio tear sheet fail with a numpy.int64 AttributeError?
It is a version skew between pyfolio and a newer NumPy, and the README documents it with the specific message `AttributeError: 'numpy.int64' object has no attribute 'to_pydatetime'`. Its prescribed fix is to install pyfolio directly from the Quantopian repository with `pip install git+https://github.com/quantopian/pyfolio` rather than the version pip resolves on its own.
What Python version does AlphaPy support?
The packaging metadata in `setup.py` declares Python 3.7 and 3.8, both of which are end of life, and marks the project as Beta. AlphaPy Pro is the edition that advertises Python 3.12 or newer. There is no `requires-python` ceiling in the legacy setup, so the effective constraint comes from the dependency floors rather than from an explicit Python bound.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/scottfreellc-alphapy)