# QuantaAlpha: the reported alpha numbers are in text, the baseline comparison is a picture

> An LLM plus evolutionary strategy framework for mining quantitative factors, where the headline figures are readable but the baseline table is an image, installation requires faking a version number because there are no tags, the debug dataset has to be renamed to the production filename, and the licence lives only in a classifier.

**QuantaAlpha/QuantaAlpha** — QuantaAlpha transforms how you discover quantitative alpha factors by combining LLM intelligence with evolutionary strategies. Just describe your research direction, and watch as factors are automatically mined, evolved, and validated through self-evolving trajectories.

- Repository: https://github.com/QuantaAlpha/QuantaAlpha
- Stars: 1,577 · Forks: 306
- Language: Python
- License: not declared
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/quantaalpha-quantaalpha

## The framework's own numbers are text; the baseline table is a picture

The results section states a small set of figures in prose and puts everything else in images. What is readable: for the best configuration, described as the framework combined with one named model, on one Chinese index over a 2022 to 2025 test period, an information coefficient of 0.0472, a rank IC of 0.0459, an annualised return of 4.68%, an information ratio of 0.6453 and a maximum drawdown of 11.80%.

The transfer claim is also in text: factors mined on one index are said to transfer zero-shot to two others, with cumulative excess return of roughly 40.3% on one and 19.1% on the other by the end of the test period.

The comparison that would make those numbers meaningful is not. The caption says the table is a complete comparison against traditional machine learning, deep learning, factor library and language model agent baselines, and the table itself is an image. So the framework's own figures can be read, quoted and checked, while every competing figure requires looking at a picture.

The robustness section follows the same pattern, describing year by year coefficients through a style rotation in 2023 and the coefficient curve across the first five mining iterations, both as images.

## Installation requires you to invent a version number

The install sequence is four commands and one of them exists only to satisfy the build system:

```bash
git clone https://github.com/QuantaAlpha/QuantaAlpha.git
cd QuantaAlpha
conda create -n quantaalpha python=3.10
conda activate quantaalpha
# 以开发模式安装包
SETUPTOOLS_SCM_PRETEND_VERSION=0.1.0 pip install -e .

# 安装额外依赖
pip install -r requirements.txt
```

The reason is the build configuration. The version is derived from git tags by a setuptools plugin, with a scheme that guesses the next development version, and the repository has no tags and no releases. So the pretend-version environment variable is not optional polish: without it the editable install has no version to report.

The second command is unusual in a different way. The project declares its dependencies as dynamic, read straight out of that requirements file during the editable install, and then the instructions tell you to install the same file again for the extra dependencies. The file's own header says it lists only third-party packages actually imported by the project code, which is a good discipline, but the two-step install means the pinned set is whatever the file says at that moment rather than a resolved lock.

## The debug dataset has to be renamed to the production filename

Data preparation is the longest section, and it contains two details that will cost you an hour if you miss them.

The first is the rename. Two directories are created under a folder whose name is literally about being git-ignored, one for the full price and volume file and one for the debug subset. The full file is copied in under its own name, and the debug file must be copied into the second directory renamed to the production filename. The note says so explicitly. So the code finds the debug data by filename convention rather than by a setting, and a copy that keeps the debug name silently finds nothing.

The second is the environment variables that point at those directories. They carry a prefix in mixed case that names a different project, one for the main data folder and one for the debug folder. Anyone writing a wrapper script has to reproduce a spelling that has nothing to do with this project's name.

The reason the files are shipped precomputed is also stated: the system can generate the full file from the underlying market data on a first run, but that process is very slow, so downloading it saves the wait.

## The published dataset covers one market; the transfer claim names two others

The data is published on a model hub as one dataset repository, with three files and a stated purpose for each: a compressed archive of raw market data for one Chinese index covering 2016 to 2025, required for initialising the backtester; a full precomputed price and volume file, required for mining; and a smaller precomputed debug subset, required for the debug and validation path.

Two download routes are given. One uses the hub's command line client with a repository type flag and a local directory, and the other uses wget against the dataset's resolve URLs for the same three files.

Here is the tension. The main experiment and the published data are both about one Chinese index. The transfer claim is about factors mined there moving to two other indices, one of them a United States benchmark. Nothing in the data section supplies those two markets, so reproducing the headline transfer number needs a second data source the project does not host.

The backtest configuration is also the one place that names a specific baseline set: a combined mode that adds a named set of twenty factor-library features to the custom factors.

## Mining is one script with an optional second argument

The mining entry point is a shell script taking the research direction as its first argument, with a second optional argument that sets a suffix for the factor library. Two examples are given in the documentation: one asking for price and volume factor mining, and one asking for microstructure factors with a short library suffix.

Everything discovered is written to a library file whose name carries a wildcard suffix, and that file is the input to the second stage, which is a separate Python module:

```bash
python -m quantaalpha.backtest.run_backtest \
  -c configs/backtest.yaml \
  --factor-source custom \
  --factor-json all_factors_library.json
```

The source flag has two values. One loads only the custom factors, and the other combines them with the named baseline feature set. There is also a dry-run flag with verbose output whose documented purpose is to load the factors without running a backtest, which is a factor-loading smoke test rather than a preview of results.

Where results land is not a command line argument. The documentation says results are saved to the directory named in the experiment output setting inside the backtest configuration file, so the output path lives in YAML rather than in the command.

## The web interface is a second application with its own start script

The repository ships a separate front-end directory and a separate way to start it:

```bash
conda activate quantaalpha
cd frontend-v2
bash start.sh
# 访问 http://localhost:3000
```

Port 3000, its own script, its own directory. So the project is two applications that share one environment: a Python package driven by the root shell script, and a web application that runs on its own. The web side is described as able to do the whole workflow without the command line.

Its four areas map onto the pipeline. System settings cover the model API, the data paths and the experiment parameters. Factor mining takes natural language input and shows progress live. The factor library lets you browse, search and filter what was found, with quality grading. And standalone backtesting picks from the library, runs a full-period backtest and shows the results visually.

The Python side also exposes a console command that points at a command line interface module, which is the entry point the two documented install paths do not mention.

## The licence is a classifier, and the project URLs use a different repository name

The packaging metadata is where the loose ends are. The classifier list claims an MIT licence, no licence is named in the repository metadata, and there is no licence file in the tree. Three sources, and the only one that names a licence is a string in a metadata field, which means anyone redistributing this needs to resolve the question themselves.

The project URLs point at a repository path written in lowercase while the repository itself is capitalised, which usually resolves through a redirect on the forge and will not resolve at all if the project moves.

The package discovery directive has no include filter and searches the working directory, so anything that looks like a package at the top level is fair game for the build. And the dependency list is almost entirely unpinned: three entries carry version ranges for compatibility with downstream packages, and the rest carry no version at all, including the file-locking library that prevents concurrent runs, the CLI framework, the fuzzy matching library and the model client.

The project also classifies itself as alpha status software, targets Python 3.10 and 3.11, and documents itself in Chinese with an English version linked from the header. Two of the four badge links in that header are empty anchors, and the next-generation research harness is announced as a separate repository rather than as a branch here.

## Conclusion

QuantaAlpha is interesting for the shape of the loop rather than for the results. The unit of work is a trajectory: you give a research direction, the system plans several routes, evolves candidates across iterations, validates them, and writes what survives into a factor library you can then backtest separately. That separation between mining and backtesting is the part worth copying, because it lets you re-run the evaluation against different data without paying for another mining pass.

Four things to check before trusting or reproducing any of it. The baseline comparison table is an image, so the competing numbers are not in the text you can read; the framework's own figures are. The reported transfer result moves from one Chinese index to two others while the published dataset covers only the first, so reproducing the transfer needs data you have to source. Installation requires setting a pretend version because the project derives its version from git tags and has none. And the licence exists only as a package classifier, with no licence file in the repository and nothing recorded in the metadata.

The honest summary is that this is an alpha-status research prototype that publishes its data and its commands in detail, which makes it reproducible in principle and unverified in practice until you run the backtest yourself.

## FAQ

### What results does the QuantaAlpha README report?

For one index over a 2022 to 2025 test period it reports an information coefficient of 0.0472, a rank IC of 0.0459, an annualised return of 4.68%, an information ratio of 0.6453 and a maximum drawdown of 11.80%. The comparison against baseline methods is presented as an image rather than as text.

### Why does the QuantaAlpha install command set a pretend version?

Because the version is derived from git tags and the repository has no tags or releases. The editable install therefore needs a pretend version supplied in the environment before the package will build.

### What data does QuantaAlpha need for factor mining?

Market data for the backtester plus a precomputed price and volume file for mining. The debug subset must be copied into its directory renamed to the production filename, because the code looks it up by that name.

### Does the QuantaAlpha repository ship a licence file?

No. The package classifier claims an MIT licence and no licence is named in the metadata, but there is no licence file at the root, so the terms are not stated anywhere in the files themselves.

## Sources

- [Issues](https://github.com/QuantaAlpha/QuantaAlpha/issues)
- [QuantaAlpha/QuantaAlpha on GitHub](https://github.com/QuantaAlpha/QuantaAlpha)
- [README](https://github.com/QuantaAlpha/QuantaAlpha/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/quantaalpha-quantaalpha
