Model or dataset
Miasyster/QuantGPT avatar
Miasyster/QuantGPT

QuantGPT gives the agent the decision rights, and keeps a folder of paths that failed

Agent-driven alpha factory — LLM autonomously designs, backtests, and submits factors to WorldQuant BRAIN

475 stars125 forksPythonMIT

At a glance

What is it?
QuantGPT is a factor research engine that hands a Claude agent fifteen MCP tools for designing, backtesting, scoring and rejecting alpha expressions, with a persistent knowledge base split into rules, findings and failures. The parts worth reading are the ones about restraint: a second model reviews every conclusion, a fitness gate skips the second market, and an enforced discipline list caps nesting depth and requires failed experiments to be written down.
Who is it for?
QuantGPT fits a researcher who already knows what a Sharpe ratio means, wants the tedious loop of generate, backtest and reject automated, and is willing to review the output rather than rubber stamp it. It does not fit someone who wants a turnkey strategy, since the deliverable is a factor expression and a backtest, and it does not fit anyone who will let an agent decide when to submit.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 135 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The agent is the researcher, and the human is the reviewer

The project's central claim is about decision rights rather than about automation, and it draws the contrast explicitly.

The traditional pattern, which it groups with a chat assistant pointed at a backtest library, is a human invents the factor, a tool runs the backtest, and a human reads the result. In that arrangement the agent is the executor and the human is the decision maker.

QuantGPT's pattern inverts both roles: a human defines the goal, the agent researches on its own, and a human reviews what comes out. The agent is the researcher and the human is the reviewer. The README is careful that this is not an interface difference, not natural language against code, but a difference in who decides.

The list of decisions the agent takes on its own is the part to read closely, because it is the actual scope: which direction to explore, what expression to generate, which metrics to evaluate, when to iterate, when to abandon a direction, and when to submit. That last one is the reason the submission path is described as optional and configurable rather than automatic.

The framing is also what makes the rest of the design coherent. A system that gives an agent the right to abandon 200 hypotheses needs a way to remember the 200 that failed, or it will rediscover them, and it needs a second opinion before it commits, or it will be confident. Both of those mechanisms exist, and the next sections are about them.

Fifteen tools, and three of them submit to another platform

The agent's toolbox is delivered over MCP from either Claude Code or Claude Desktop, and the README says there are fifteen tools of which it names ten.

The research tools are the bulk of it. `run_backtest` runs a grouped backtest across the whole market, `score_factor` produces a composite score from zero to one hundred, and `diagnose_factor` looks at why a factor failed rather than only that it failed. Three tools are about not fooling yourself: `run_anti_overfit` runs four anti overfitting tests, `run_rolling_validation` does walk forward validation, and `validate_expression` checks the syntax before anything expensive happens.

Two are for orientation. `list_operators` returns documentation for more than sixty expression operators, described as including non linear, ternary and technical indicator forms, and `list_universes` returns the stock pools and benchmarks available, which is the one that connects the expression to the market it will be tested on.

Then the three WorldQuant BRAIN tools, and they are the ones to think about twice. `wq_brain_submit` submits a single factor, `wq_brain_batch_submit` submits a parameter sweep in batch, and `wq_brain_submit_by_ids` submits a chosen set. Those are the tools that turn a research tool into something that acts on a third party's platform, and the credentials for them are an optional block in the environment file, described as what enables real submission.

So the submission path exists, is opt in, and is deliberately the smallest part of the surface.

A fitness gate below 0.1 skips the second market

Batch evaluation is where the compute saving lives, and the mechanism is a single threshold.

A round submits ten to twenty factor expressions at once, runs the backtests concurrently, and retries in three waves. Results come back sorted by fitness in descending order. Then the gate: when the fitness on one index is below 0.1, the run skips validation on the second index entirely, and the stated reason is to save compute.

The call is a single function, and the concurrency is a parameter rather than a config file:

python
from scripts.factor_miner import batch_evaluate
results = batch_evaluate(
    server, expressions, params,
    max_concurrent=10
)

Two things follow from that design. A factor that does badly on the first market is not carried to the second, which is a defensible filter and also a strong prior, since it bets that a factor failing one universe will not work in another. And the wave based retry means a batch that hits a rate limit is retried rather than lost, which matters when the batch is twenty expressions at a time.

The README's own track record puts a round of eight candidate factors at about fifteen minutes, against a cumulative count of more than 370 backtest tasks. Those are the project's figures, reported without a breakdown of what the compute cost was.

Every conclusion goes to a second model before it counts

The cross review is the mechanism aimed at confirmation bias, which the README names as the core problem of single model factor research.

The rule is that every conclusive judgement, and it names three kinds, adopt, do not adopt, and close the direction, has to be independently reviewed by a second LLM. What gets sent is not just the result: the factual data and the first model's reasoning chain go to DeepSeek Reasoner, which is asked to assess independently whether the reasoning is sound and whether any angle was missed.

Then the two outcomes are handled differently. Agreement outputs directly. Disagreement presents the evidence from both sides and adopts the more conservative conclusion.

That asymmetry is the design, and it is worth being clear about what it does and does not buy. It removes a class of failure where a model commits to a direction because it framed the problem in a way that made the direction look obvious. It does not remove a class of failure where both models share the same framing, or where the first model's reasoning chain anchors the reviewer, since that chain is sent along with the data and is the thing being reviewed.

The cost is a second model call per decision, on a loop that makes many decisions, which is one more reason the batch path and the fitness gate exist. The stack table puts DeepSeek in the AI row for both factor generation and cross review, so the two roles are the same model family unless configured otherwise through the OpenAI compatible endpoint.

Three folders in the knowledge base, and one of them is off limits

The knowledge base is the reason a tenth research session is cheaper than a first, and its layout is three directories.

`rules/` holds validated stable rules that must be followed. `findings/` holds empirical findings, which are for reference rather than binding. And `failures/` holds falsified paths, with the instruction that they must not be repeated.

The README makes the claim this is not chat history but structured research assets, and the operational version of that is worth reading: a tenth session can use the findings of the previous nine, avoid repeating their experiments, follow the rules that have already been validated, and steer around the paths that have already been falsified. The failure folder is the one doing the real work, because a direction that has been killed once will be proposed again by a model that does not remember being killed.

Next to it sits a research discipline list that the README insists is not advice but hard rules, and it reads like the lessons from exactly those killed directions. Change one variable per experiment. Check whether the experiment has already been run before running it, against both the notes and the knowledge base. Label analytical conclusions as hypotheses only. Record the reason for a failed experiment, not just its failure. Anything nested more than four levels deep needs extra justification. And simple and clear beats complex and clever.

The nesting limit is the one that would change your results. A four level cap on an expression is a hard constraint on the search space, and it is imposed by the tool rather than discovered by the model.

Three factors, three expressions, and three claimed Sharpe ratios

The README reports three factors that passed the WorldQuant BRAIN in sample tests and were submitted, with a table of expressions and results that is worth reading as claims rather than as evidence.

Debt Momentum Composite is written as a negative rank of the average difference of close over ten days plus a rank of debt to enterprise value, with a claimed Sharpe of 1.77, a fitness of 1.26 and returns of 20.18 percent. VWAP Decay Reversal is a negative rank of the linear decay of close over vwap over ten days, at 1.69, 1.07 and 18.63 percent. Returns Volume Momentum is a negative rank of the linear decay of returns times volume over a twenty day average, at 1.60, 1.03 and 24.15 percent. All three are marked as passing every in sample test and as submitted.

The README attributes the three to different sources of return, which is the part that shows the intent: the first combines momentum reversal with a fundamental ratio and is industry neutralized, the second captures mean reversion of the deviation from volume weighted average price and is market neutralized, and the third captures decay momentum of returns against relative volume and is also market neutralized.

What the page does not contain is anything outside the project. There is no independent verification of those three results in the visible text, no out of sample period named for them, and the numbers sit in the same table as a claim that A grade factors go to a hosted cloud service for independent validation. If the Sharpe figures matter to a decision, they are the project's own backtest output, and the walk forward and anti overfitting tools are what you would run yourself to find out whether they hold.

The container ships with authentication switched off

The deployment files are the place where a research tool becomes a service, and they are worth reading before you run one.

The compose file builds the image, maps one port, mounts two named volumes for data and reports plus the SQLite file, loads the environment file, and then sets two things explicitly. Authentication is disabled, and the database URL is a Postgres connection string with the password spelled as the literal word password, pointing at a host port outside the container. The restart policy is unless stopped. Anyone copying that file as a starting point is starting with no login and a placeholder credential, and the environment example file is more careful than the compose file, carrying a warning never to enable it in production.

The Dockerfile is two stages, a Node 20 alpine build of the React frontend and a Python 3.12 slim runtime that installs the package in editable mode along with the Postgres extra, copies the frontend bundle in, exposes the same port, and disables authentication again as an environment variable, along with a process task backend and two worker processes. The entry command runs the package over HTTP transport.

The environment file is the most useful document in the repository, because it is ordered by what you need rather than alphabetically. The LLM block is DeepSeek over an OpenAI compatible base URL, and the comment says the key is optional because leaving it empty gives you expression only mode with no natural language input. Then the database, which is empty by default so SQLite is used. Then JWT settings including a generated secret and separate access and refresh expiries, SMTP settings for email verification codes, an optional feedback webhook, an optional RiceQuant credential block described as the primary A share data source, an admin password that literally reads change this password, task tuning, a rate limit, the allowed origins, the task backend choice, a Rust engine switch that falls back to pure Python, and finally the WorldQuant BRAIN credentials, marked optional and described as what enables real submission.

The project version is 2.8.0 in the packaging file, there are no tagged releases, and the last push to main is dated 20 May 2026.

Editorial conclusion

QuantGPT fits a researcher who already knows what a Sharpe ratio means, wants the tedious loop of generate, backtest and reject automated, and is willing to review the output rather than rubber stamp it. It does not fit someone who wants a turnkey strategy, since the deliverable is a factor expression and a backtest, and it does not fit anyone who will let an agent decide when to submit. Before running it, decide whether you want the WorldQuant BRAIN submission path at all, keep the free market data sources as the default rather than a paid one, and read the access settings in the compose file, which ship with authentication switched off.

Frequently asked questions

What is QuantGPT and how is it different from an AI backtest tool?

It is an agent-driven factor research engine, and the README says the difference is decision rights rather than interface. The traditional pattern has a human inventing factors and an agent executing backtests; QuantGPT has a human setting the goal and the agent choosing what to explore, what to generate, when to iterate, when to abandon a direction and when to submit, with a human reviewing the output.

What tools does the QuantGPT agent have?

Fifteen MCP tools, of which the README names ten: run_backtest for a grouped market backtest, score_factor for a zero to one hundred score, diagnose_factor for failure analysis, run_anti_overfit for four overfitting tests, run_rolling_validation for walk forward checks, validate_expression, list_operators for more than sixty operators, list_universes, and three WorldQuant BRAIN submission tools for single submission, batch sweeps and selected ids.

How does QuantGPT avoid fooling itself about a factor?

Every conclusive judgement, whether to adopt, reject or close a direction, goes to a second LLM together with the data and the first model's reasoning chain, which is asked to assess the reasoning independently. Agreement outputs directly, and disagreement presents both sides and takes the more conservative conclusion. There is also a knowledge base with a failures folder of paths that must not be repeated.

What data sources and databases does QuantGPT use?

Market data comes from baostock and akshare, both described as free, cached to Parquet, with an optional RiceQuant credential block described as the primary A share data source. The database is SQLite by default with no configuration, or PostgreSQL as an optional extra, and the LLM is DeepSeek over an OpenAI compatible base URL, where an empty key gives expression only mode with no natural language input.

Does QuantGPT submit factors to WorldQuant BRAIN automatically?

The submission path is described as optional. Three of the fifteen MCP tools handle single submission, batch parameter sweeps and submission by ids, and the environment file marks the WorldQuant credentials as optional and as what enables real submission. Separately, factors that pass local validation in the A grade band are described as being uploaded to the project's own cloud service for independent verification and out of sample tracking.

Official sources

  1. Issues
  2. License: MIT
  3. Miasyster/QuantGPT on GitHub
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/miasyster-quantgpt.svg)](https://hysenlabs.com/projects/miasyster-quantgpt)