Model or dataset
DR-lin-eng/stock-scanner avatar
DR-lin-eng/stock-scanner

stock-scanner, and the six version folders the README still points at

开源A股量化分析(并且配合llm模型,进行高级分析)

1,146 stars438 forksPythonMIT

At a glance

What is it?
This is a Chinese A-share analysis tool that combines 25 financial indicators, technical signals and news sentiment with a large language model layer, and the most useful thing in it is structural: each release is a separate top-level directory rather than a branch, so the quick start tells you to run the desktop app from the 2.0 folder and the web app from a directory whose own name says it is a test version, while the changelog says 3.0 is what you should deploy.
Who is it for?
Read this project as a worked example of how to combine quantitative indicators with a language model, and take from it the weighting scheme, the streaming event design and the rules fallback, all of which are well specified here. Do not treat its output as investment advice, because the README advertises explicit buy and sell recommendations, which is a claim no indicator set can support.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Activity is slowing. The repository last received commits 7 months ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Six directories where a branch would normally be

The quick start that precedes all of this is three commands, and the second one does not match the first:

bash
git clone https://github.com/DR-lin-eng/stock-scanner.git
cd stock-analysis-system
pip install -r requirements.txt

Cloning a repository creates a directory named after the repository, and this one is called stock-scanner, so the change-directory line names something that does not exist. It is a small thing, and it is the kind that survives in a README nobody has followed end to end recently, which is worth knowing before treating the rest of the instructions as tested.

The top-level listing is then the most informative thing about this repository, and it is unusual enough to be worth reading slowly.

There is a directory for each major version: 1.0, 2.0 win app, 2.5 webapp, 2.6 webapp, 3.0 webapp, and 3.1 webapp. Six of them, in the repository root, at the same time. A project with this history would normally have used branches, and the presence of all six as parallel trees means every version is a full copy that lives in the default branch forever.

The consequences are practical rather than aesthetic. A change has to be made in one directory, which is a copy, rather than in a shared source tree, so anything the author fixed after 2.0 has to be fixed six times or only once and the other five keep the bug. A pull request touches one directory, so a reader has to work out which version it targets from the path. And the repository grows at the rate of the whole application rather than the rate of a diff.

The directory names also carry the author's own state of mind. The 3.0 directory is named for Hong Kong and US stock support. The 2.6 one carries a Chinese suffix meaning streaming transmission test version, which is an explicit label that this build is not the production one.

Here is where that collides with the instructions. The quick start tells you to change into the 2.0 directory and run the desktop GUI, and to change into the 2.6 directory to run the web server. The development log at the top of the README says the latest version is 3.0, that it can go into production, that it supports Hong Kong and US stocks and news compression, and separately that version 3.1 had a lot of bugs and was rolled back. So the documented path to a working application is the 2.6 folder, which by its own name is a test build, and the production version is a directory the instructions never mention.

The 3.1 directory is still in the tree, which is consistent with a rollback: the code was written, the bugs were found, the author retreated to 3.0, and the abandoned attempt stayed where it was. That is a reasonable thing to do and a confusing thing to ship.

Two smaller artefacts at the root are worth noting for the same reason. There is a PDF whose name says, in Chinese, that it is only a website example download and that it looks ugly, and a loose Python file for analysing all stocks at the recommendation stage, sitting next to the versioned applications.

Twenty-five indicators, five categories, and what a composite score rests on

The financial indicator layer is the part of this project that would survive without the language model, and it is specified concretely enough to check.

Twenty-five indicators across five categories. Profitability with net margin, return on equity and return on total assets. Solvency with the current ratio, the debt-to-asset ratio and the interest coverage ratio. Operating efficiency with total asset turnover and inventory turnover. Growth with revenue growth rate and net profit growth rate. And market performance with the price-earnings ratio, price-to-book and the PEG ratio.

That is a conventional set, and its conventionalness is a point in its favour. These are the ratios a financial analyst would reach for first, they are computable from published statements, and the calculation is not the interesting part. The interesting part is what happens to them next.

They feed a technical layer and a sentiment layer, and the three are combined into a composite score on a scale of zero to one hundred with a stated weighting: technical 0.4, fundamental 0.4, sentiment 0.2. The sentiment weight being half the technical weight is the design decision, and it is a defensible one for a tool that aggregates news. It is also a number someone chose, and the README gives no argument for it.

The technical layer is equally conventional: multi-period moving averages and MACD crossovers as trend, RSI overbought and oversold plus Bollinger band position as oscillators, and price-volume relationship with a volume ratio. The technical period is configurable, defaulting to 365 days.

The sentiment layer reads news, announcements and research reports, with a configurable maximum of 100 news items. Reading research reports specifically is an interesting choice, since those are written by people with a financial interest in the outcome, and a language model summarising a bullish analyst note will report bullishness as a fact about the market rather than as a position taken by someone paid to take it.

So the composite score is a weighted sum of three differently-trusted signals, and the weight on the least reliable one is the smallest, which is the right instinct. What a reader cannot check is whether the underlying computation is right, and for a project like this the honest test is to take ten companies whose fundamentals you know and see whether the ranking matches.

A rules fallback underneath the model, which is the right pairing

The AI layer is three model providers, a failover mechanism, streaming output and a fallback, and the fallback is the part worth reading first.

The README says that when the AI is unavailable the system automatically degrades to advanced rule-based analysis. That single sentence is what makes this design defensible. A tool that produces investment commentary has a failure mode that is worse than a crash: it produces a confident, well-formatted, wrong answer. With a fallback, the worst case is a rule-based output the reader can trace to the twenty-five indicators and the weighting, which is a much better failure than a fabricated paragraph.

The three providers are OpenAI GPT-4, Claude 3 and Zhipu's ChatGLM, with a model preference in the configuration and automatic switching between a primary and a backup API for availability. The example configuration names specific model identifiers, a small variant for OpenAI and a specific dated Claude model, which is a detail that dates the project and also tells you the author tested with cheap models rather than the expensive ones, which is the right choice for a per-request cost.

The Zhipu option is worth noting for a specific reason. It is a domestic Chinese model provider, and a Chinese A-share tool that has a domestic fallback is making a statement about availability. The demo log in the README says the demo site now runs on a DeepSeek-derived open model, so the author's own deployment settled on a self-hostable open model rather than a hosted one, which is both a cost decision and a data-residency one.

What the AI layer is asked to do is listed as four things: interpretation of the twenty-five indicators, an investment strategy with explicit buy and sell recommendations, risk and opportunity identification, and comparison against peers in the same industry. Three of those are reasonable. The fourth, the explicit recommendation, is the one to think about carefully, and the next section returns to it.

The streaming implementation is described as real-time AI analysis display over Server-Sent Events, which means the model output is not shown at the end but as it is produced. That is a presentation choice with a substantive consequence: a reader watching the analysis arrive can stop reading halfway if it is going somewhere unhelpful, which a wall of finished text does not let them do.

Six event types and an Nginx block that tells you what streaming costs

The web application is Flask with Server-Sent Events, and the interface is described well enough to be a specification rather than a feature list.

There are five documented endpoints. Two are GET, a status check and a system information endpoint. One is a GET event stream. Two are POST, single-stock analysis and batch analysis, both streaming. There is no separate non-streaming analysis endpoint in the list, which tells you the streaming path is the primary one rather than an add-on.

The event types are six, and the naming is the interesting part because it maps the lifecycle of an analysis onto the wire. Connected confirms the connection. Log carries progress messages. Progress carries progress updates. Scores update carries scoring results. Final result is the complete answer. And there is a separate event for the AI stream content.

Splitting the AI stream from the other five is the design decision that makes this work. Progress and score updates are facts the system knows, computed from the indicators, and they arrive as soon as they are computed. The AI interpretation arrives as the model produces it, which is slow. Merging them into one event type would mean the whole response is gated on the slowest part, and a user would stare at nothing for the model's first token. Separating them means the quantitative half of the result appears immediately and the language model half streams in behind it.

The Nginx configuration in the README is the second half of the story, and it is the part that tells you what streaming actually requires of your infrastructure. The event stream location sets the connection header to empty, forces HTTP version 1.1, turns buffering off, turns caching off, and disables chunked transfer encoding. Every one of those is there because a reverse proxy will otherwise break a long-lived event stream: buffering holds the response until it completes, which for an infinite stream is never, and the user sees nothing. Caching is the same failure with a different name.

The production run command is Gunicorn with four workers bound to all interfaces on port five thousand, and the same event stream caveats apply to it, which is why the Nginx block is in the documentation at all. A Flask development server is fine for the streaming demonstration the directory is named for and wrong for anything else.

The config file, the troubleshooting notes, and where the data comes from

The configuration is a single JSON file, and reading it tells you the security posture as clearly as any documentation would.

There are three sections. API keys for the three model providers, with placeholder values. An AI section with a model preference and per-provider model identifiers. And web authentication, which is a flag, a password and a session timeout in seconds. So the web application is password protected with a session, and the password lives in the same file as the API keys, in plaintext.

That is a reasonable choice for a tool one person runs on their own machine, and an inconvenient one for a shared deployment. The README's security section lists password authentication, session management with configurable timeout, cross-site request forgery protection, strict input validation, locally encrypted API key storage, error handling that does not leak sensitive information, and session-based access control. That is a competent list, and the one item that cannot be verified from outside is the encrypted key storage, since the configuration file itself shows the keys as readable strings.

The troubleshooting section is more revealing than the feature list, because it shows where the data actually comes from. When a data fetch fails, the first check is pinging a quote host, and the second is upgrading a package by name. So the market data comes from an open-source Chinese financial data library rather than a paid feed, which means it is free, it is as fresh as the library's scraping, and it is subject to whatever the upstream service does about automated access.

That last point is a genuine operational risk for anyone outside China, and the README's own note about market coverage is a symptom of it. The introduction says the system currently supports Chinese stocks only, with Hong Kong and US markets under optimisation because news and information retrieval there is limited and slow. A data source that is slow and limited for one market is likely to be unavailable from outside the region that hosts it, and the project does not document a proxy or a mirror. Anyone outside mainland China should test connectivity before investing an evening in the setup.

The rest of the troubleshooting is competent and specific. For an AI failure, it prints the configured keys with a one-line Python command and then tests the provider directly with a bearer token request. For a web access failure, it checks what is listening on port five thousand. That is what a maintainer writes who has actually debugged this, and it is more useful than a log of error messages would be.

The buy and sell recommendation, and the honest reading of what this is

The README says the AI layer provides explicit investment strategies and buy and sell recommendations, and it is worth sitting with that claim rather than repeating it.

Nothing in the underlying analysis can support it. Twenty-five accounting ratios, a handful of oscillators and a sentiment score from a hundred news items are inputs to a valuation view. A view is not a recommendation. The distance between them is a position size, a time horizon, a risk tolerance, a portfolio context and a judgement about what the analyst does not know, and none of those are in the data. A buy and sell recommendation derived from those inputs is a formatting decision.

The composite score makes the same point from the other direction. A number from zero to one hundred is a ranking, and a ranking is only as meaningful as the weights behind it. Technical 0.4, fundamental 0.4, sentiment 0.2 was chosen by one developer, and a different weighting would reorder the results. The tool has no way to express that uncertainty to the person reading the number, and a number formatted as a score invites exactly the confidence the inputs do not support.

There is also a structural problem with the sentiment layer that the weighting only partly addresses. The AI reads news, announcements and research reports. A research report is written by someone whose employer or client benefits from a particular view, and a model asked to summarise it will report the conclusion rather than the interest. The same applies to news coverage, where the volume of coverage tracks how much a company is being discussed rather than what is happening to it.

None of this makes the project bad. It makes it a data aggregator with a language model in front, and that is a legitimate and useful thing. The problem is the framing in the feature list, where professional investment advice appears as a feature of an open-source tool with a donation link. This is not investment advice, and a reader who takes a single output from it as a decision has been misinformed by the interface rather than by the analysis.

The genuinely reusable parts are elsewhere. The event stream design that separates computed facts from model output is well made and directly applicable to any application that combines computation with generation. The rules fallback is the right instinct. The weighted composition of three differently-trusted signals, with the weakest weighted least, is a defensible default even if these particular numbers are arbitrary. And the three-provider failover with a self-hosted open model as the demo's choice is a practical answer to availability and cost that most projects in this space have not thought about.

Editorial conclusion

Read this project as a worked example of how to combine quantitative indicators with a language model, and take from it the weighting scheme, the streaming event design and the rules fallback, all of which are well specified here. Do not treat its output as investment advice, because the README advertises explicit buy and sell recommendations, which is a claim no indicator set can support. Verify first by running one stock end to end from the directory the changelog names rather than the one the quick start names, checking that the data source is reachable in your region, and reading the weights in the config before trusting a composite score, because a 0 to 100 rating is only as meaningful as the 0.4 and 0.4 and 0.2 split behind it.

Frequently asked questions

What does the stock-scanner project do?

It is an AI-enhanced A-share analysis system combining 25 financial indicators across five categories, technical indicators including moving averages, RSI, MACD and Bollinger Bands, news sentiment drawn from over 100 items, and a large language model interpretation layer. It ships as a PyQt6 desktop application and a Flask web application with streaming output.

Which stock markets does stock-scanner support?

The introduction says only Chinese stocks are supported, with Hong Kong and US markets under optimisation because news retrieval there is limited and slow. The development log and the 3.0 directory name say Hong Kong and US support was added, so the coverage depends on which directory you run, and the version history section stops before that.

Which AI models does stock-scanner use, and what happens if they fail?

OpenAI GPT-4, Claude 3 and Zhipu ChatGLM, with a model preference in the config and automatic switching between a primary and a backup API. When the AI is unavailable the system degrades to rule-based analysis over the indicators, and output is streamed over Server-Sent Events as it is produced.

How do I run stock-scanner?

Install the requirements with pip, then run either the desktop GUI or the Flask web server. The README's paths point at the 2.0 desktop directory and the 2.6 web directory whose name marks it as a test build, while the development log states 3.0 is the production version and that 3.1 was rolled back because of bugs.

Where does stock-scanner get its market data?

From an open-source Chinese financial data library rather than a paid feed, according to the troubleshooting section, which pings a quote host and tells you to upgrade that library. That means the data is free and as fresh as the library's own retrieval, and the README's note about slow news access for other markets suggests accessibility varies by region.

What are the stock-scanner web API endpoints?

Two GET endpoints for status and system information, one GET for the event stream, and two POST endpoints for single-stock and batch streaming analysis. The stream emits six event types: connected, log, progress, scores_update, final_result and ai_stream, with the AI stream separated so computed results appear before the model output.

Official sources

  1. DR-lin-eng/stock-scanner on GitHub
  2. License: MIT
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/dr-lin-eng-stock-scanner.svg)](https://hysenlabs.com/projects/dr-lin-eng-stock-scanner)