Open-source project
edtechre/pybroker avatar
edtechre/pybroker

PyBroker: bootstrap metrics, walkforward training, and two truncated README snippets

Algorithmic Trading in Python with Machine Learning

3,556 stars456 forksPythonNOASSERTION

At a glance

What is it?
PyBroker is a Python backtesting framework for algorithmic trading strategies built around machine learning, published as lib-pybroker. Its honest parts are the resampled metrics and the walkforward windows; its rough edges are the quick examples that stop mid call and a licence field the metadata leaves unasserted.
Who is it for?
PyBroker earns a trial from anyone who wants backtest numbers they can defend, because bootstrap metrics and walkforward windows are cheap to adopt and awkward to retrofit later. Judge it on the guide notebooks rather than the feature list, since the quick examples are cut off mid expression and the repository ships both bundled agent skills and an asv benchmark harness without publishing any figures from it.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 5, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Bootstrap metrics replace the single backtest number

Every backtesting framework can print one return figure. PyBroker goes a step further and reports metrics derived from randomized bootstrapping, which is why evaluation sits in its own user guide notebook rather than in a paragraph at the end of the backtest one. A single run over one date range produces one number, and that number moves with the start date, so you can pick a range that flatters a rule without meaning to. Resampling the result set repeatedly produces a distribution instead of a point, and the distribution becomes the thing you compare strategies on. The practical cost is that the evaluation pass costs more than the backtest pass. PyBroker keeps the two apart, so your execution function stays free of metric code and the metric pass can be repeated against cached data. None of this rescues a rule written with look-ahead in it. It does make the summary statistics less fragile than a single printed figure, which is the part that survives contact with real money.

Walkforward analysis is the path a trained model takes

When a model is fitted on the same bars it is scored on, the backtest measures memorization. PyBroker's walkforward analysis splits the timeline into successive windows: fit on one stretch, score on the bars that follow, then roll forward and repeat. The project describes this as simulating how the strategy would have performed during actual trading, which is the point, because the model only ever sees data that already existed at decision time in that construction. The hook shape shows up in the model snippet:

python
   def train_fn(symbol, train_data, test_data):
      # Train the model using indicators stored in train_data.
      ...
      return trained_model

   # Register the model and its training function with PyBroker.
   my_model = pybroker.model('my_model', train_fn, indicators=[...])

   def exec_fn(ctx):
      preds = ctx.preds('my_model')
      if not ctx.long_pos() and preds[-1] > buy_threshold:
         ctx.buy_shares = 100
      elif ctx.long_pos() and preds[-1] < sell_threshold:
         ctx.sell_all_shares()

The split happens before your callback runs, so `train_fn` never holds the future bars and `ctx.preds('my_model')` is the only documented way to read model output inside an execution function. The rule-based path has no model and still crosses an explicit range, `Strategy(YFinance(), start_date='1/1/2025', end_date='8/1/2026')` with execution registered for `['AAPL', 'MSFT']`. Indicators are attached at `add_execution`, not inside the function, so one callback serves both symbols and the indicator cache is keyed per symbol.

Both quick example snippets stop mid expression

The two snippets under the quick example are trimmed. The rule-based block ends at `indicators=highest('high_10d', 'close'`, an unclosed call with no closing parenthesis, and the model-based block's final line is the bare word `alpaca`, which is where the data source was being assigned. Neither gap is cosmetic. The model snippet imports `Alpaca` at the top and never constructs it, so a reader copying it end to end has no data source at all and a bare `alpaca` raises a NameError in Python. `highest` is given a name and one field name with the trailing argument missing, so the rolling window length is left implicit. The rest of the surface is not guesswork: `ctx.long_pos()` reports whether a long position is open, `ctx.buy_shares` sets size, `ctx.hold_bars` sets how many bars to hold, and `ctx.stop_loss_pct` sets the stop. The complete versions of both examples live in the guide notebooks, and the applies-stops notebook is where the stop interaction with `hold_bars` is spelled out.

Three feeds, one custom data source contract, one cache

Historical data comes from Alpaca, Yahoo Finance, AKShare, or a provider you write yourself, and that last option decides whether the other three matter to you. The two examples already assume different paths, with the model snippet reaching for `Alpaca` while the rule-based snippet constructs `YFinance()`. If your data lives elsewhere, the custom data source notebook defines the interface your object has to satisfy, and until it returns bars in the shape the engine expects, none of the backtest notebooks run. Downloaded data, computed indicators and trained models all pass through a cache, so re-running the bootstrap evaluation does not repeat the download or the training. The dependency list explains the plumbing: `alpaca-py>=0.44.0`, `yfinance>=1.6.0`, `yahooquery>=2.4.1` and `diskcache>=5.6.3` all arrive as hard requirements rather than extras, and `AKShare` is named as a supported feed without appearing in that file.

Optuna search and joblib parallelism multiply the backtest count

Parameter optimization runs Optuna over the strategy's parameters to select the best set, with a dedicated notebook for it and separate notebooks around it for margin trading, modeling slippage, rebalancing and rotational trading. The parallelization notebook exists because of the arithmetic: a search over many parameter sets means many backtests, and each backtest already crosses a symbol list for a date range. `joblib>=1.5.3` is a hard requirement, so the parallelism is job or process based rather than an async event loop, which suits CPU-bound numeric work in NumPy. Two of the twenty-five dependencies carry an upper bound, `numpy>=2.5.2,<2.6` and `optuna>=4.9.0,<5`; every other entry is a floor only, including `numba>=0.67.0`, `pandas>=3.0.5` and `scikit-learn>=1.9.0`. The numpy cap sits directly beside numba, a compiled extension, so the resolver has to land on a compatible pair instead of taking the newest numpy and hoping.

The distribution name is lib-pybroker, not pybroker

The install command names something other than the project. The README says:

bash
   pip install -U lib-pybroker

The download-count badge in the header links to the same `lib-pybroker` name, so a dependency line written as `pybroker` resolves to something else or to nothing. The alternative is a plain clone:

bash
   git clone https://github.com/edtechre/pybroker

Cloning hands you `src/`, `tests/`, `docs/`, `benchmarks/` and `skills/` in one directory but leaves the install to you, which matters because the package metadata is assembled by setuptools from a `MANIFEST.in` and a `setup.cfg` that sit beside `pyproject.toml`. The documentation site is versioned and the badge points at `latest`, so a pinned dependency and a moving docs build can drift without either side saying so. The release list is short and recent: v1.2.14 on 2026-08-03, v2.0.0 on 2026-08-17 and v2.0.1 on 2026-08-28.

Ruff targets py312 while the install line promises 3.11

The lint block in `pyproject.toml` sets `target-version = "py312"` under a comment that says it assumes Python 3.12, while the installation section states support for Python 3.11 and up on Windows, Mac and Linux. Ruff's target steers rewrites and feature checks rather than the minimum interpreter, so a 3.11 user is not blocked, but the mismatch means the formatter can propose rewrites the running interpreter would refuse. Line length is pinned at 79, the same as Black, with four-space indentation, double quotes, spaces instead of tabs, and magic trailing commas respected rather than collapsed. The rule set is narrow, `select = ["E4", "E7", "E9", "F"]` with `ignore = ["E402"]`, leaving style warnings and complexity checks off, and docstring code formatting is enabled so examples inside docstrings match the module body. Per-file ignores widen that only where it matters: `src/pybroker/*.py` and `tests/*.py` drop E203 and E402, `src/pybroker/__init__.py` also drops F401 for its re-exports, and `tests/test_*` adds F403 and F405.

Agent skills and an asv harness ship inside the repository

Version 2 added skills for coding agents, and they live in the repository rather than in a separate package. The tree carries a `skills/` directory, a `.claude/` directory and a `CLAUDE.md` file at the root. Four are named in the README, Strategy Creator, Indicator Creator, Model Trainer and Parameter Optimization, and they map onto the four pieces you would otherwise write by hand: the execution function, an indicator, a `train_fn`, and an Optuna search space. A fifth link is cut off partway through its URL, so the bundled count is at least four and not provably four. Beside that sits `asv.conf.json`, a `benchmarks/` directory and a `.bench/` results directory, which is the airspeed velocity harness wired into the project itself. That is the file to check before repeating any speed claim: the project measures its own engine and stores the numbers, yet the README quotes none of them, so the assertion that the backtest engine is fast carries no figure anywhere you can read. The repository is not archived and the last push landed on 2026-09-28.

Editorial conclusion

PyBroker earns a trial from anyone who wants backtest numbers they can defend, because bootstrap metrics and walkforward windows are cheap to adopt and awkward to retrofit later. Judge it on the guide notebooks rather than the feature list, since the quick examples are cut off mid expression and the repository ships both bundled agent skills and an asv benchmark harness without publishing any figures from it. Before you depend on it, confirm the licence yourself: the repository metadata reports no assertion even though a LICENSE file sits at the root and the badge links to a licence page on the docs site. Pin the numpy upper bound rather than inheriting whatever your resolver picks.

Frequently asked questions

Which PyPI name actually installs PyBroker?

The install command is `pip install -U lib-pybroker`, so the distribution on PyPI is named lib-pybroker rather than pybroker. The header badge points at a download counter for the same name. Cloning https://github.com/edtechre/pybroker is the documented alternative, and that route leaves the install to you.

Does PyBroker need Python 3.12 or does 3.11 work?

The installation section states Python 3.11 and up on Windows, Mac and Linux. The Ruff configuration in pyproject.toml sets target-version to py312 with a comment about assuming 3.12, which steers lint rewrites rather than the minimum interpreter.

Which market data feeds does PyBroker read from?

Historical data comes from Alpaca, Yahoo Finance, AKShare, or a custom data source you implement yourself. Downloaded data, indicators and trained models go through a cache, and the dependency list carries alpaca-py, yfinance, yahooquery and diskcache as hard requirements.

How does PyBroker keep a trained model from cheating on the backtest?

Walkforward analysis trains on one window and scores the following bars, then rolls forward. The training callback receives train_data and test_data as separate arguments, and the execution function reads model output through ctx.preds('my_model').

What agent skills ship with PyBroker v2?

Four are named in the README: Strategy Creator, Indicator Creator, Model Trainer and Parameter Optimization. They live in a skills/ directory next to a .claude/ directory and CLAUDE.md at the repository root. A fifth link is cut off, so the full list is not visible there.

Official sources

  1. edtechre/pybroker on GitHub
  2. Issues
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/edtechre-pybroker.svg)](https://hysenlabs.com/projects/edtechre-pybroker)