Model or dataset
ZhuLinsen/alphasift avatar
ZhuLinsen/alphasift

AlphaSift: a YAML-driven stock screening engine with an optional LLM ranking layer

AI-native stock screening engine with full-market discovery, LLM ranking, risk-aware scoring, and auditable evaluation. AI选股

366 stars212 forksPythonApache-2.0

At a glance

What is it?
AlphaSift scans a full A-share snapshot, filters it with auditable YAML strategies, and optionally asks an LLM to rank the survivors. It is built for engineers who want the screening logic readable and the results replayable, not for anyone expecting trading signals.
Who is it for?
Adopt AlphaSift if you want the screening rules in version-controlled YAML and you are willing to supply your own market data source and read the fallback metadata on every hotspot payload. Do not adopt it if you need intraday quotes, a hosted service, or a component that tells you what to buy.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 89 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem AlphaSift addresses: turning a screening idea into something you can re-run and audit

Most retail screening tools give you a form with fixed fields and a result table you cannot inspect. AlphaSift takes the opposite route. A strategy is a YAML document, the engine applies it to a full market snapshot, and every run can be saved and evaluated against a later snapshot. The README describes the pipeline as three layers: L1 deterministic screening with hard filters and factor scoring, L2 optional LLM ranking with theses and risk buckets, and L3 pluggable post-analysis that defaults to a local scorecard.

The target user is an engineer or quant-adjacent researcher working on A-share data who wants the ranking rationale to be a file in git rather than a black box. The README is explicit that this is for learning, research, and engineering experiments, and that outputs can be delayed, incomplete, or wrong. The recorded examples were run on April 12, 2026 against the April 10, 2026 close, with LLM ranking disabled. That date-stamping is the tell: the project treats a screening result as a reproducible artifact, not a live recommendation.

How the three layers connect, and where the LLM actually sits

The data flow starts with a market snapshot. The README states that the default provider priority depends on whether a Tushare token is present: without one it is efinance, akshare_em, em_datacenter, and with one it becomes tushare, efinance, akshare_em, em_datacenter. It also notes that Tushare data is the most recent trading day's daily bars and daily_basic, not a live order book. That single sentence sets the expectation for the whole system: this is an end-of-day screener.

L1 reduces the universe. The README's recorded example shows 5190 stocks filtered down to 337 for the dual_low strategy, and 5190 down to 126 for volume breakout. The survivors are scored by deterministic factors and ranked.

L2 is opt-in. When enabled, AlphaSift sends a candidate set to an LLM for cross-candidate reasoning. Several environment variables govern this: LLM_CANDIDATE_MULTIPLIER controls how many candidates are pulled in relative to the final output, LLM_MAX_CANDIDATES caps the prompt size, LLM_MIN_COVERAGE sets the fraction of candidates the model must actually address, and LLM_RANK_WEIGHT (default 0.40) decides how much the LLM score moves the final ranking. That last variable is the honest part of the design: the LLM is a weighted input, not an oracle, and you can dial it to zero by passing --no-llm.

L3 runs after ranking. The default is a local scorecard, and the README shows --explain to print its reasoning. DSA is available as an optional analyzer via --post-analyzer dsa, which requires DSA_API_URL to be set. You can also disable the layer entirely with --no-post-analysis.

Installing AlphaSift and running a first screen without an API key

The package requires Python 3.10 or newer, per pyproject.toml. The README's quick start installs it in editable mode and copies the environment template. Nothing here needs an LLM key; the key only matters if you want L2 ranking.

bash
pip install -e .
cp .env.example .env

Before screening, list the strategies that ship with the package so you know which name to pass. The README shows this command returning the built-in strategy list.

bash
alphasift strategies

The no-key demo runs an end-to-end pass without any provider configuration.

bash
alphasift quickstart

For a real screen with the LLM layer switched off, pass a strategy name and --no-llm. The README's example output for dual_low shows a first line of the form "Universe 5190 -> filtered 337 -> output Top 5", followed by a table with rank, code, name, score, price, change, PE and PB columns. If your first line reports a universe far below the market you expect, the snapshot provider chain did not resolve.

bash
alphasift screen dual_low --no-llm

Saving a run is what makes later evaluation possible. The README pairs --save-run with the runs and report commands.

bash
alphasift screen dual_low --no-llm --save-run
alphasift runs --json
alphasift report <run_id> --output data/reports/dual_low.md

The LLM layer is optional, and that is a deliberate constraint

It is worth being blunt about the L2 boundary. AlphaSift does not require an LLM to produce a ranked list, and the README's own headline examples were generated with --no-llm. The LLM path adds theses, catalysts, risks, confidence, and portfolio risk buckets, but it also adds a provider dependency, a prompt-size budget, and a coverage threshold that can leave part of the candidate set unaddressed.

The configuration surface reflects that tension. LLM_MIN_COVERAGE defaults to 0.60, meaning the model is expected to cover at least 60 percent of candidates; LLM_MAX_RETRIES defaults to 1, so a malformed response is retried once and then, presumably, falls back. LLM_JSON_MODE defaults to true. None of this makes the LLM deterministic, and the README does not claim it does. The rank weight of 0.40 means a model that disagrees with the factor score can move a stock several places.

If your use case requires an explainable ranking with no external calls, run with --no-llm and treat the L3 scorecard as your explanation layer. That combination is fully local.

Hotspot discovery and the fallback metadata you have to read

The hotspot subsystem is separate from screening. It ranks topics and sectors by heat, resolves a topic into a detail payload, and writes cache and history sidecars. The README shows a discovery command that writes both a JSON cache and a JSONL history file, and a detail command that can read from a fallback cache file.

bash
alphasift hotspots --provider akshare --top 12 --output data/hotspots.json --history data/hotspot.history.jsonl --explain
alphasift hotspot "AI compute" --top-stocks 10 --timeline --fallback-cache data/hotspots.json --explain

The cache schema is versioned at 2 and carries generated_at, a metadata block with provider, row count, source errors, and stale or fallback state, plus normalized hotspot rows. Detail payloads keep a raw timeline for auditability and a compact route list grouped by day, newest first, trimmed for display.

The design choice worth noting is that fallbacks are labeled rather than hidden. When live constituent APIs fail and AlphaSift substitutes cached leaders, the returned stocks carry source="last_good_cache.leader_stocks", a source_confidence value, and fallback_used=true. That is a better failure mode than silently serving stale data, but it also means any downstream consumer has to check those fields. A dashboard that ignores fallback_used will present cached leaders as if they were live. The README does not document a way to make the engine fail hard instead of falling back.

A real alternative: a general-purpose backtesting framework

The closest thing to a drop-in alternative is a general-purpose backtesting library such as backtrader or vectorbt, paired with your own data loader. The difference in approach is where the work sits. A backtester assumes you already have a signal and asks how it would have performed over history; AlphaSift assumes you have a universe and asks which names survive your filters today, with an evaluation loop that compares saved runs to newer snapshots.

That evaluation loop is narrower than a backtest. The README describes saving runs, evaluating later using newer snapshots, deducting transaction cost, tagging follow-through and failed-breakout outcomes, reviewing failure samples, and optionally fetching price paths for max drawdown and max favorable excursion. It does not describe position sizing, portfolio construction, or multi-strategy capital allocation. If you need those, AlphaSift is a candidate generator that feeds into something else, not a replacement for the backtester.

The second difference is the strategy format. A YAML strategy file under strategies/ is readable by someone who does not write Python, which matters if the person defining the screening criteria is not the person maintaining the code. A backtesting script is not.

Licence, maintenance, and the cost of keeping a data pipeline alive

AlphaSift is Apache-2.0, with the licence file listed at the repository root and license-files declared in pyproject.toml. That permits commercial use and modification, and it includes an explicit patent grant. It does not remove the disclaimer's force: the project states it is not investment advice and that users are responsible for compliance checks, transaction costs, liquidity risks, and announcement timing. Nothing about the licence changes who owns a trading decision.

The repository is not archived and the last push was on 2026-07-03. There are no retrieved releases, so the version to track is the one in pyproject.toml, currently 0.2.0, meaning installation from source is the normal path rather than a published wheel.

The ongoing cost is the data layer, not the code. The dependency list includes five market data libraries: efinance, akshare, baostock, tushare, and yfinance. Each has its own upstream endpoints and its own breakage pattern. The fallback chain in SNAPSHOT_SOURCE_PRIORITY exists precisely because any one of them can fail, and ALPHASIFT_FALLBACK_SNAPSHOT_PATH points at a last-good snapshot file. Budget for keeping that chain working; the screening logic itself is unlikely to be what breaks.

Editorial conclusion

Adopt AlphaSift if you want the screening rules in version-controlled YAML and you are willing to supply your own market data source and read the fallback metadata on every hotspot payload. Do not adopt it if you need intraday quotes, a hosted service, or a component that tells you what to buy. Before committing, run alphasift screen dual_low --no-llm on your own machine and confirm the universe count in the first line matches the snapshot you expected; if it does not, the provider chain in SNAPSHOT_SOURCE_PRIORITY is not resolving the way you assumed.

Frequently asked questions

Does AlphaSift require an LLM API key to screen stocks?

No. The README's quick start includes a no-key demo via alphasift quickstart, and the recorded screening examples were all run with --no-llm. A key is only needed if you want the L2 LLM ranking layer, which reads GEMINI_API_KEY, OPENAI_API_KEY, DEEPSEEK_API_KEY, or the LiteLLM variables.

What data sources does AlphaSift use for its market snapshot?

The default provider priority is efinance, akshare_em, em_datacenter when no Tushare token is set, and tushare first when TUSHARE_TOKEN is present. The README notes that Tushare data is the most recent trading day's daily bars and daily_basic, not a live order book.

How do I install AlphaSift?

The README's quick start installs it in editable mode with pip install -e . and then copies the configuration template with cp .env.example .env. The package requires Python 3.10 or newer according to pyproject.toml.

Can I tell whether AlphaSift returned cached hotspot data instead of live data?

Yes. When live constituent APIs fail and cached leaders are substituted, the returned stocks carry fields including source="last_good_cache.leader_stocks", a source_confidence value, and fallback_used=true. The hotspot cache metadata block also records stale or fallback state.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. README
  4. ZhuLinsen/alphasift on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/zhulinsen-alphasift.svg)](https://hysenlabs.com/projects/zhulinsen-alphasift)