mom-index publishes the formula, not the strategy behind it
👩👧 宝妈指数 — 追踪小白/宝妈投资情绪的反向指标。当菜市场大妈都在讨论股票时,就是你该离场的时候。
At a glance
- What is it?
- A behavioural finance framework that scores retail investor chatter on Chinese social platforms, released as a readable baseline with four fixed weights, synthetic data and collector interfaces. The maintainer's private calibration is deliberately absent.
- Who is it for?
- This fits someone who wants to see how a sentiment indicator is assembled rather than take one on trust, since the weights, the signal categories and the data flow are all in the repository. It does not fit anyone looking for a validated signal, and the project says so: the sample index has not been empirically validated and the formula is a research starting point.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 33 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The published part is the framework, not the strategy
The repository states plainly what it is not publishing.
The public version shows the complete data flow and the extension interfaces. It does not include the strategies the maintainer currently uses, the model prompts, the production calibration logic, or the raw social platform data.
That split explains most of what looks odd in the code. The analysis layer uses early keyword rules and fixed weights, described as chosen so they can be understood and modified. The data files are desensitised synthetic samples that correspond to no real user. The collectors are interfaces with example implementations, and confirming platform terms and data authorisation is left to whoever runs it.
The index formula is described as a research starting point that does not represent the current private version and does not guarantee investment effectiveness.
So the correct reading is that this is a teaching implementation. If you want the signal the author actually trades on, it is not here, and the file names rather than the prose are the honest summary of that.
Four weights that add up to a readable baseline
The composite index is four terms, and the weights are printed in full.
宝妈指数 = 新手讨论占比 × 40%
+ 新手信号强度 × 25%
+ 情绪极端度 × 20%
+ 信号纯度 × 15%In order: the share of discussion coming from newcomers carries the most weight, then the strength of the newcomer signal, then how extreme the sentiment is, and finally the purity of the signal.
Two design choices are visible in those numbers. The first term is a share rather than a count, so a busy day and a quiet day produce comparable readings. And the last term is a discount rather than a driver, so text that matches the rules loosely counts for less.
The documentation is direct about their status: these weights are only a readable baseline for the first version, and users are advised to add a semantic model, confidence evaluation, de-biasing, time decay and historical backtesting.
That list is effectively the to-do list for anyone treating the number seriously, and none of those five are in the repository.
The fourth signal category is the one pointing the other way
The text analysis is organised as four categories, and the fourth is the interesting one.
The first captures identity and knowledge-seeking signals, which is where someone asking how to do something first appears. The second covers chasing, panic and signals of depending on someone else's decision. The third is intent, separating buying, selling and waiting.
The fourth is a set of counter-signals: professional terminology, risk awareness, and a long-term perspective.
That category is what stops the index from being a simple measure of how many people are excited. Talking about drawdowns and time horizons lowers the reading even when the discussion volume is high, which is the mechanism that makes the number mean something other than attention.
The component that implements this is named for a history rather than for what it does. The text analyser file keeps an older name in its filename even though the first version is a rule-based analyser rather than a model call, and the file comment says so.
That mismatch between filename and behaviour is the clearest signal in the repository that the private version swaps in something else at exactly that point.
Collectors are interfaces, and one file is labelled historical
Three files sit in the collectors directory and they are not equivalent.
One is a public web page collection example. One is an optional data source interface for a different platform. And one is described in the structure listing as historical experimental code.
That last label matters. Anti-detection work is the kind of thing that arrives with a stability warning, and rather than omitting it the structure listing marks it as experimental history. Anyone reading the tree can see which files are examples and which are not maintained.
The interface contract for extending it is one line: a new collector should emit unified post fields. That is what lets the analysis layer stay unaware of where a post came from.
The example environment file holds two optional values for the optional source, both empty, under a comment saying never to commit real ones. There is no credential handling in the analysis layer at all, which is consistent with the dependency list containing nothing that would suggest otherwise.
Two dependencies, and neither of them analyses anything
The requirements file is short enough to read in full, and it contains two packages.
requests>=2.31,<3
playwright>=1.40,<2Both are collection libraries. One makes HTTP requests and one drives a browser, and both are capped below their next major version.
There is no model client, no numerical library and no machine learning framework in that list, which matches the description of the analysis as keyword rules with fixed weights. Everything after collection is arithmetic on counts, so the absence is consistent rather than an omission.
It also means the analytical half of the project is the part you are expected to replace. The extension notes point at exactly three files: a new collector, the analyser, and the index calculator where calibration, decay or different weights go.
The frontend has no Python dependency at all, because the charting library is vendored into the directory rather than fetched. The dependency file mentions a local copy of the chart library specifically to avoid variation from a CDN, which is the kind of decision that matters when a dashboard has to keep rendering.
The quickstart is written for PowerShell
The getting-started sequence is short, and the shell it assumes is visible in the first two lines.
py -3 -m venv .venv
.\.venv\Scripts\Activate.ps1
pip install -r requirements.txt
# 运行示例流程
py -3 pipeline.py
# 启动看板
py -3 frontend/server.py 8765Creating a virtual environment, activating it through the Scripts path and PowerShell activation is Windows-specific as written. On another platform the equivalent steps differ, and the repository does not give a second version.
Two entry points follow. The pipeline file runs the sample flow, described in the structure listing as collect, analyse, index, then write data out. The frontend server is then started on a port, and the dashboard is opened from that port in a browser.
The server is described as a local one with caching disabled, which suits a development dashboard where you want a change to appear on reload.
One structural detail the write-up does not mention: the repository root also holds a data synchronisation script and a tests directory, neither of which appears in the project structure listing. So the test suite exists, and how it is run is not documented here.
The data safety list is a do-not-commit list
The section on data safety is written as a list of things not to put in the repository, which is more specific than most projects manage.
Cookies, login state and browser profiles are first. Then API keys, access tokens and environment files. Then unauthorised raw posts, comments and user identifiers. Then screenshots, debug pages and logs containing real personal information.
The first and third items are the ones that would matter most for this project in particular, since the collectors exist to gather posts from accounts that are logged in. A browser profile committed by accident carries session state for every site that profile ever authenticated to, not just the one being collected.
The section opens by stating that the posts, time series and statistics in the repository are all desensitised demonstration data, which is consistent with the description of the sample files earlier in the document.
There is also a companion project, a Dad Index aimed at discussion among experienced investors, which suggests the same pipeline is meant to be pointed at the opposite population, since experienced-investor talk is the counter-signal this one treats as a discount.
Editorial conclusion
This fits someone who wants to see how a sentiment indicator is assembled rather than take one on trust, since the weights, the signal categories and the data flow are all in the repository. It does not fit anyone looking for a validated signal, and the project says so: the sample index has not been empirically validated and the formula is a research starting point. Two things to do before treating any reading as meaningful: add time decay and backtesting, both of which the documentation recommends and none of which the first version ships, and confirm that your data source permits collection, which is left as the user's responsibility.
Frequently asked questions
What is the Mom Index measuring?
Discussion heat, sentiment extremity and buy or sell tendency among retail investors and newcomers on Chinese social platforms. A higher reading means the discussion is more concentrated and more extreme, and it is described as a research observation rather than a trading instruction.
How is the Mom Index formula weighted?
Four terms: the share of discussion from newcomers at 40 percent, newcomer signal strength at 25 percent, sentiment extremity at 20 percent, and signal purity at 15 percent. The weights are described as a readable baseline for the first version.
Does mom-index ship real social media data?
No. The public version ships desensitised synthetic demonstration data and explicitly excludes the maintainer's private strategies, model prompts, production calibration logic and raw platform data.
What dependencies does mom-index require?
Two: requests and playwright, each capped below its next major version. Both exist for collection, and the example environment file holds optional collector credentials that must not be committed.
How do I extend the Mom Index framework?
Add a collector that emits unified post fields, replace the rule-based analyser with your own rules or a model, add calibration, decay or different weights in the index calculator, build offline backtests from historical samples, and extend the dashboard.
Is the Mom Index investment advice?
No. It is stated to be for behavioural finance research, software engineering demonstration and popular science, and the sample index is described as not sufficiently empirically validated.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/mihang123-mom-index)