Investing Algorithm Framework: one strategy definition, four ways to run it
Framework for quantitative trading. Complete framework for development, backtesting, and deploying automated trading algorithms and trading bots.
At a glance
- What is it?
- This is an Apache licensed Python framework for quantitative trading whose organising idea is that a strategy is written once and then executed by four different engines: a fast vectorised backtester for sweeping thousands of variants, an event-driven simulator for realism, a paper trader, and a live trader. The most defensible part of its feature list is not the trading but the six named backtest windows and a permutation test that asks how often your result could have happened by luck.
- Who is it for?
- This framework is worth evaluating if you write strategies and have ever discovered that the version you tested and the version you deployed were subtly different code paths, because keeping one definition and swapping the execution engine is the correct answer to that problem.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 6 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 5, 2026, and from our analysis. They are not legal advice.
Editorial analysis
One definition, four engines, and the promise that nothing gets rewritten
The tagline is the design brief: build a strategy, backtest it, compare results, and deploy the winner without rewriting the strategy.
That last clause is the whole point, and it is worth being precise about the problem it solves. In most trading codebases, a strategy lives in two places. There is a backtest version, which is written for speed and can assume it knows the future, and there is a live version, which is written for correctness and cannot. The two drift apart. A subtle difference in how a position is sized, or in how a missing bar is handled, or in the order of two operations that both look independent, means the strategy you validated is not the strategy you are running. That bug is invisible until the money is real, and it is the single most common way a backtest becomes fiction.
The framework's answer is that a strategy is defined once as entry and exit signals, and the same definition is handed to each of four execution engines. A vectorised backtester evaluates the signals across the whole dataset at once for speed. An event-driven simulator advances bar by bar, with realistic orders, fills, costs and portfolio management. A paper trader does the same against live data without sending orders. A live trader sends them.
That structure also implies a discipline about which engine answers which question. The vector engine is for asking many questions badly; the event engine is for asking one question well. Running a thousand variants through the fast engine and then re-running the survivor through the slow one is the correct sequence, and the framework makes it a sequence rather than a rewrite.
The deployment story follows the same logic. The documented targets are running locally, self-hosting, two serverless function platforms, and one commercial platform that also sponsors the project. Those are four ways to run the same loop on different infrastructure, not four codebases.
And there is an integration point that is becoming standard in this category: a built-in protocol server, so an assistant can drive the backtester rather than you writing glue code. For a framework whose interface is configuration objects rather than a user interface, that is a reasonable surface to expose.
Six named backtest windows, two of which exist to catch you
Most backtesting tools let you specify a start date and an end date. This one has six named window types, and the naming is the feature.
A study in this framework bundles the assets, the assumptions, and the evaluation periods, and a period can be rolling, anchored, holdout, walk-forward, time-based out of sample, or universe-based out of sample. Those are six distinct answers to the question of how to divide a dataset, and each one encodes a different belief about what a result means.
Rolling and anchored are the two familiar ones and they differ in a small but consequential way. A rolling window slides forward by a fixed step, so every result is measured over a recent period. An anchored window keeps the start fixed and extends the end, so the first result covers a short history and the last covers a long one. The anchored version is more stable and the rolling version is more current, and choosing wrongly between them is a decision you should make consciously.
The other four are the ones that matter. A holdout window is a period you promise not to look at while you are developing. Walk-forward is a sequence of in-sample periods each followed by its own out-of-sample period, so you are testing repeatedly rather than once. Time-based out of sample holds out a period of time. Universe-based out of sample holds out a set of instruments rather than a period, which catches a strategy that worked on the particular securities you happened to pick.
Making these first-class named objects rather than two date fields is a small design decision with a large effect. A date range is a string you can get wrong in a hundred ways. A named window type is a declaration of what kind of test you are running, and the framework can then report on window coverage, which the dashboard feature explicitly does. If your evaluation of a strategy rests on one in-sample period, that is visible in the report. If it rests on four windows of different kinds, that is visible too.
For anyone doing this work honestly, the holdout and the two out-of-sample variants are the features that separate a framework from a script. A strategy that only works in sample is the default outcome of a sweep, and the only defence is refusing to look at one part of the data, which is much easier to do when the framework has a name for it.
Results are columnar, compressed and indexed, which is why sweeps are fast
There is a claim in the feature list that sounds like a database detail and is actually the load-bearing engineering decision in the whole project.
It is that you can rank and filter more than ten thousand backtests through a relational database without decoding the full result bundles. Paired with that is a named, versioned, portable bundle format for storing complete vector and event results, and a storage demo sitting in the examples directory.
For that to work, three things have to be true, and the dependency list shows all three.
First, results must be compact. The framework depends on a binary serialisation format and a modern compression algorithm, which together mean a bundle holding thousands of bars of per-asset data does not occupy a database row per observation. Second, the contents must be columnar. A columnar library is a dependency, and a columnar format is what lets a query touch one field across a million rows without materialising the rest. Third, the index must be a separate, cheap thing. The claim is about ranking and filtering, which are metadata operations: which of ten thousand runs had a return above this, which had fewer than that many trades, which used this parameter value. None of those questions needs the price series.
So the architecture is a split: full fidelity results in compressed columnar bundles, and a small relational index of the summary statistics for the queries you run across runs. That is the design that makes a sweep of thousands of variants practical, and it is the design that makes the interactive comparison dashboard possible, since the dashboard needs to filter and rank before it needs to load anything.
The consequence for a user is worth stating plainly, because it is a real trade. The bundles are portable and versioned, which means a backtest you ran eighteen months ago can be reopened by a newer version of the framework, and it means you can move results between machines. The cost is that the index is a derived artefact. If the two drift apart, your filters will quietly return the wrong runs, and the framework will not necessarily notice. The portable format with an explicit version is the right mitigation, and it is the thing to check first if a sweep produces results you cannot reproduce.
Permutation testing asks the only question that matters about a good result
There is one advanced feature in this framework that deserves to be described without qualification, because it is the one most quant tools leave out and everybody should have.
The permutation test measures how often randomised market paths match or outperform the result a strategy actually produced.
Read that again. It is not a test that the strategy works. It is a test of whether the result was lucky. You take the observed result, generate many alternative price paths that have the same broad statistical character, and ask how many of them your strategy would have beaten. If your strategy beat the market by twenty percent and ten thousand random paths beat it by twenty percent, you have learned nothing about your strategy. If almost none of them did, the result is probably not an artefact.
The phrase in the documentation is careful. It asks how often randomised paths match or outperform, and it does not claim the answer is a pass or a fail. That restraint is correct, because a permutation test produces a number rather than a verdict, and the number has to be interpreted against how many paths you generated and what the distribution looks like.
Why this matters more than the eighty performance metrics in the same feature list is not hard to explain. A sweep over a thousand parameter combinations will find one that looks spectacular on the data you used to find it. Every metric you can compute on that result will look spectacular too: a high ratio, a shallow drawdown, a good recovery, a favourable comparison against a benchmark. The metrics are not wrong. They are measuring a run that was selected for being lucky, and there is no way to tell that from the metrics alone.
The permutation test is the only measurement in the list that is sensitive to the selection itself. That is why it belongs in the advanced section rather than the basics, and why a framework that includes it is taking the problem of overfitting seriously rather than only the problem of execution.
The other item in that section, cross-sectional pipelines, serves the same purpose from the other direction. Ranking, filtering and scoring a whole universe of instruments during each iteration means the strategy is evaluated on a population rather than on a handful of chosen names, which is the other way a result gets to look better than it is.
Confluence cards put a veto above a score, and that is the interesting part
One feature in this framework has nothing to do with trading and deserves attention anyway, because it is a decision-recording design that generalises.
A confluence scoring card is described as building explainable decisions from four ingredients: requirements, weighted evidence, vetoes, and score thresholds.
Weighted evidence with a threshold is the ordinary pattern, and it has a known weakness. Weights are positive, scores accumulate, and a high total wins. If one piece of evidence is disqualifying, a sufficient number of mediocre pieces will still add up to a passing score, because the arithmetic does not know the difference between important and disqualifying.
Vetoes are the addition. A veto is a condition that overrides the total regardless of what the rest of the evidence says. That is a categorically different decision structure from scoring, and it is the correct one for anything involving risk, where a single disqualifying fact should stop the decision rather than be outvoted.
The fourth ingredient is what makes the structure auditable. If a decision is a threshold, you can store the score and the threshold and reproduce the outcome. If a decision involves a veto, you have to store why the veto fired, which means the card is not a number with a comment attached. It is a structured record of the reasoning.
That is the honest value of the feature. In a trading framework, a decision record per signal or per parameter change is how you answer the question that matters six months later, which is why this particular configuration was chosen over the alternatives you tested. Without a card you have the configuration and no reasoning. With a card you have the requirement that was being satisfied, the evidence considered, the threshold, and any veto, and you can reconstruct not just the decision but the class of decision that was available at the time.
Whether the framework's implementation of this is any good is not something the feature list can tell you. That it is in the list at all, in a project about trading, suggests the author has thought about the same problem I have, which is that an automated system that makes decisions nobody can reconstruct is a system nobody will trust with money.
Three alpha tags in a day, an exact prerelease pin, and two lockfiles
Three process details in this repository are worth a reader's attention, and none of them is about trading.
The first is the release cadence. The three most recent tags are the twenty-first, twentieth and nineteenth alpha builds of the ninth major version, cut within about a day and a half of each other. That is a project shipping an alpha every few hours, and the readme says so directly: the alpha is available and you must install the prerelease explicitly.
Which is the unusual part. The documented install is not a version range and not a bare package name:
pip install investing-algorithm-framework==9.0.0a21It is a full version specifier with an exact equality, pinned to the specific alpha that was current when the readme was written. The project is telling you to pin, which is correct advice for a trading framework where a dependency change can silently alter a backtest, and it is also an admission that the alpha is not stable enough to trust with a range.
The practical consequence is that the readme's install command is already out of date relative to the next alpha, and will be within a day of you reading it. That is fine if you follow the pattern rather than the command: pin whatever the current alpha is, and record it, because a backtest result is only reproducible if you know which version produced it.
The second is two lockfiles. The repository contains both a lockfile for one dependency resolver and a lockfile for another. That usually means the project migrated to the second tool and kept the first file, and it is a real hazard: two resolvers can disagree about transitive versions, and a lockfile that is not used by your install command is decoration that looks like a guarantee. If you depend on this framework, use one resolver, commit the lockfile that resolver produced, and ignore the other.
The third is the dependency list itself, and there is one entry that will surprise anybody who is not trading cryptocurrency. A multi-exchange cryptocurrency trading library is a required dependency, listed without an optional marker and without an extra group, while every market data provider is optional behind an extra. So a user who wants daily equity bars from one provider, and nothing else, installs a library that speaks the order and market APIs of dozens of exchanges.
That is a consequence of how the project is built rather than an accident, and it is defensible if the framework's core abstraction is an exchange client. It is also thirty-odd transitive dependencies you did not ask for, and it is worth knowing before you decide the dependency list is unreasonably large.
Editorial conclusion
This framework is worth evaluating if you write strategies and have ever discovered that the version you tested and the version you deployed were subtly different code paths, because keeping one definition and swapping the execution engine is the correct answer to that problem. It is a poor fit if you want a finished product with sensible defaults, since a ninth major is still at alpha and the documented install pins an exact prerelease, and it is a poor fit for equity-only work unless you are willing to install a multi-exchange cryptocurrency library to get there. Read the window documentation before your first sweep and use the holdout and out-of-sample windows from the start, because a sweep over a thousand variants on a single window will find an overfit result and the framework will not stop you.
Frequently asked questions
What is the investing-algorithm-framework and what does it do?
It is a Python framework for the quantitative trading workflow in which a strategy's entry and exit signals are defined once and then run through a fast vectorised backtester, an event-driven simulator, paper trading, and live trading without rewriting the strategy. Results can be inspected in an interactive dashboard and exported as a self-contained report.
Why does investing-algorithm-framework have both vector and event-driven backtesting?
They answer different questions. The vector engine evaluates signals across the whole dataset at once so thousands of strategy variants can be swept quickly, and the event-driven engine advances bar by bar with realistic orders, fills, costs and portfolio management to validate a small number of candidates properly.
What backtest window types does investing-algorithm-framework support?
Six: rolling, anchored, holdout, walk-forward, time-based out of sample, and universe-based out of sample. They are named objects rather than a pair of dates, and the framework reports on window coverage, so a result evaluated on a single in-sample period is visible as such.
How does the investing-algorithm-framework store and index backtest results?
Complete results go into portable, versioned bundles, and a relational index holds the summary statistics so that ranking and filtering more than ten thousand backtests does not require decoding the bundles. The dependency list includes a binary serialisation format, a modern compression algorithm and a columnar library, which is what makes partial queries cheap.
What does Monte Carlo permutation testing do in investing-algorithm-framework?
It measures how often randomised market paths match or outperform the result a strategy actually produced, which is a check on whether an observed result could have been luck. It is the one measurement in the feature list that is sensitive to having selected the best run out of many, which is the failure mode of a parameter sweep.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/coding-kitties-investing-algorithm-framework)