QuantaAlpha: an LLM-driven factor mining framework built on self-evolving trajectories
QuantaAlpha transforms how you discover quantitative alpha factors by combining LLM intelligence with evolutionary strategies. Just describe your research direction, and watch as factors are automatically mined, evolved, and validated through self-evolving trajectories.
At a glance
- What is it?
- QuantaAlpha turns a research direction written in plain language into mined, evolved and backtested alpha factors, using an LLM inside an evolutionary loop. The repository is small, the data requirements are not, and the licence file is missing even though the README shows an MIT badge.
- Who is it for?
- Adopt QuantaAlpha if you already run Qlib data locally and want to test whether trajectory-level evolution produces factors that survive a walk-forward backtest on CSI 300. Do not adopt it if you have no A-share market data, no LLM API budget, or no appetite for a project still classified as Development Status 3 - Alpha.
- Can I use it commercially?
- Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
- Is it still maintained?
- Yes. The repository last received commits 95 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem QuantaAlpha addresses: factor research as a search problem
Classical alpha factor research is a manual loop. An analyst forms a hypothesis about a market inefficiency, writes an expression over price and volume fields, backtests it, looks at the information coefficient, and revises. The bottleneck is not compute, it is the rate at which a human can generate and discard hypotheses. QuantaAlpha targets that bottleneck directly. The README describes the input as a research direction and the output as verified alpha factors, with the intermediate steps being diversified planning, trajectory evolution and validation. The intended user is a quantitative researcher who already has Qlib-format market data and wants a machine to propose the candidate expressions rather than writing them by hand. The project is classified in pyproject.toml as Development Status 3 - Alpha, and the last push to the repository was on 2026-06-29, so this is a research codebase rather than a production system. The paper linked from the README badge is arXiv 2602.07085, which is where the evaluation methodology lives; the README itself only summarises the headline numbers.
How the trajectory-based self-evolution actually works
The pipeline shown in the README is a four-stage flow: research direction, diversified planning, trajectory evolution, verified alpha factors. The mechanism that distinguishes it from a single-shot prompt is the middle stage. Instead of asking a model once for a factor expression, QuantaAlpha maintains trajectories, which the README describes as being evolved at the trajectory level and constrained by structured hypothesis-code pairs. In practice that means a hypothesis is expressed in a structured form and paired with the code that implements it, so the evolutionary step operates on something checkable rather than on free text. The README also mentions diverse planning initialisation, which is the mechanism that keeps the initial population from collapsing onto one idea. Validation runs against Qlib, and the README reports the CSI 300 test period as 2022 to 2025 with an information coefficient of 0.0472, a Rank IC of 0.0459, annualised return of 4.68 percent, an information ratio of 0.6453 and a maximum drawdown of 11.80 percent under the best configuration, described as QuantaAlpha with GPT-5.2. Those are the authors' reported results, not independently reproduced here. The zero-shot transfer claim is the more interesting one: factors mined on CSI 300 were transferred without retraining to CSI 500 and S&P 500, with the README reporting cumulative excess returns of roughly 40.3 percent and 19.1 percent respectively by the end of the test period.
Installing QuantaAlpha and running a first mining job
The README gives a conda-based install. The SETUPTOOLS_SCM_PRETEND_VERSION variable is needed because the version is derived from git metadata by setuptools-scm, and the README pins it to 0.1.0 for the editable install.
git clone https://github.com/QuantaAlpha/QuantaAlpha.git
cd QuantaAlpha
conda create -n quantaalpha python=3.10
conda activate quantaalpha
SETUPTOOLS_SCM_PRETEND_VERSION=0.1.0 pip install -e .
pip install -r requirements.txtConfiguration is loaded from a .env file, created by copying the example. The README marks the data paths and the LLM endpoint as required, and shows deepseek-v3 as an example model name for both the chat and reasoning roles.
cp configs/.env.example .envQLIB_DATA_DIR=/path/to/your/qlib/cn_data
DATA_RESULTS_DIR=/path/to/your/results
OPENAI_API_KEY=your-api-key
OPENAI_BASE_URL=https://your-llm-provider/v1
CHAT_MODEL=deepseek-v3
REASONING_MODEL=deepseek-v3Data comes from a HuggingFace dataset the README names as QuantaAlpha/qlib_csi300. It ships cn_data.zip (A-share Qlib data for 2016 to 2025), daily_pv.h5 and daily_pv_debug.h5. The README explains the HDF5 files exist because generating them from Qlib on first run is slow.
pip install huggingface_hub
huggingface-cli download QuantaAlpha/qlib_csi300 --repo-type dataset --local-dir ./hf_data
unzip hf_data/cn_data.zip -d ./data/qlib
mkdir -p git_ignore_folder/factor_implementation_source_data
mkdir -p git_ignore_folder/factor_implementation_source_data_debug
cp hf_data/daily_pv.h5 git_ignore_folder/factor_implementation_source_data/daily_pv.h5
cp hf_data/daily_pv_debug.h5 git_ignore_folder/factor_implementation_source_data_debug/daily_pv.h5The two destination directories are not cosmetic. The debug copy is renamed to daily_pv.h5 inside its own folder, so a debug run and a full run read from separate paths. After this, the README points to a local UI for the end-to-end flow from research direction through mining to backtest, and to docs/user_guide.md for the full documentation. The repository also exposes a console entry point, quantaalpha, mapped to quantaalpha.cli:app in pyproject.toml, though the README does not document its subcommands.
Where QuantaAlpha breaks down
The data requirement is the first real constraint. QuantaAlpha does not fetch market data. It expects a Qlib cn_data directory and precomputed HDF5 files, and the only dataset the README points to is A-share CSI 300. If you trade US equities or futures, you are on your own for building both the Qlib bundle and the price-volume HDF5 files, and the README does not document that conversion path beyond noting it is time-consuming. The second constraint is the LLM dependency. Both CHAT_MODEL and REASONING_MODEL must resolve through the configured OPENAI_BASE_URL, so every mining run costs tokens against a provider you supply. There is no documented offline or local-model mode, and no cost estimate. Third, the reported numbers are modest in absolute terms: an IC of 0.0472 and annualised return of 4.68 percent are credible for a factor library but not spectacular, and they come from a specific configuration the README identifies as GPT-5.2, which means swapping in a weaker model is likely to move the results. Fourth, the project is labelled Alpha, there are no retrieved releases, and the README's own roadmap points to a successor project, QuantaAlpha-claw, described as the next-generation research harness. That is a signal about where maintenance attention is likely to go. Finally, the README shows an MIT badge but the repository has no LICENSE file at the top level, which is a gap you should resolve before depending on the code.
QuantaAlpha compared with AlphaAgent and single-agent factor miners
The closest comparison in the search data around this project is AlphaAgent, and the difference is architectural rather than cosmetic. A single-agent factor miner typically runs one generate-evaluate-refine loop per candidate, where the model sees the last result and produces the next expression. QuantaAlpha instead keeps trajectories and evolves at that level, with structured hypothesis-code pairs as the constraint. The practical consequence is that diversity is maintained across the population rather than only within one chain of revisions, which matters when the search space of expressions is large and a single greedy chain converges early. That said, the trajectory approach costs more LLM calls per accepted factor, and the README gives no comparison of token consumption against a single-agent baseline. If your goal is a quick sanity check on whether an LLM can write plausible Qlib expressions at all, a single-agent loop is cheaper and easier to debug. QuantaAlpha earns its complexity only when you are running a sustained search and care about the population-level behaviour.
Maintenance, licence and upgrade cost
The last push to the default branch was on 2026-06-29, roughly three months before the time of writing, so the repository is not abandoned but it is not moving fast either. There are no retrieved releases, which means installation is from source at a moving commit and there is no versioned artefact to pin. The install itself depends on git metadata through setuptools-scm, which is why the README sets SETUPTOOLS_SCM_PRETEND_VERSION; if you build from a tarball without a .git directory, you should expect to set that variable yourself. Dependency pinning is partial: numpy is constrained to >=1.24,<2.0 and pandas to >=1.5,<3.0, but langchain-community, openai, pymupdf and the rest are unpinned, so a fresh install months from now may resolve differently than it did for the authors. On licensing, the README header carries an MIT badge and pyproject.toml declares the MIT classifier, but no LICENSE file appears in the top-level repository listing. The declared intent is MIT, which is permissive, but a missing licence file is a real ambiguity for anyone embedding this in a commercial stack. That is a factual gap, not legal advice; if the licence matters to your organisation, ask the maintainers to add the file.
Editorial conclusion
Adopt QuantaAlpha if you already run Qlib data locally and want to test whether trajectory-level evolution produces factors that survive a walk-forward backtest on CSI 300. Do not adopt it if you have no A-share market data, no LLM API budget, or no appetite for a project still classified as Development Status 3 - Alpha. Before anything else, verify three things: that your Qlib cn_data directory matches the layout the README expects, that the repository actually ships a LICENSE file despite the MIT badge in the README header, and that your chosen CHAT_MODEL and REASONING_MODEL are reachable through the OPENAI_BASE_URL you configure, because the framework will not run without a working LLM endpoint.
Frequently asked questions
What does QuantaAlpha do?
It combines a large language model with evolutionary strategies to mine, evolve and validate quantitative alpha factors, taking a research direction as input and producing verified factors as output. The README describes the flow as research direction, diversified planning, trajectory evolution, verified alpha factors.
What data does QuantaAlpha need before it can run?
Two things: Qlib market data for backtesting and precomputed price-volume HDF5 files for factor mining. The README points to the HuggingFace dataset QuantaAlpha/qlib_csi300, which ships cn_data.zip, daily_pv.h5 and daily_pv_debug.h5.
Is QuantaAlpha licensed under MIT?
The README shows an MIT badge and pyproject.toml carries the MIT classifier, but the top-level repository listing contains no LICENSE file. Treat the licence as declared but not shipped until a licence file appears.
Which LLM does QuantaAlpha use?
You supply it. The .env file requires OPENAI_API_KEY, OPENAI_BASE_URL, CHAT_MODEL and REASONING_MODEL, and the README shows deepseek-v3 as the example value for both model variables. The best reported results in the README are for a configuration it identifies as GPT-5.2.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/quantaalpha-quantaalpha)