BettaFish runs a seven-slot LLM panel and a forum loop over social media
微舆:人人可用的多Agent舆情分析助手,打破信息茧房,还原舆情原貌,预测未来走向,辅助决策!从0实现,不依赖任何框架。
At a glance
- What is it?
- BettaFish is a from-scratch Python multi-agent system that reads public posts and comments across Chinese and international social platforms, debates findings in a moderator-led forum loop, and renders an HTML report. It needs seven LLM endpoints, a PostgreSQL database and a Playwright browser, and its copyleft licence is the first thing to settle.
- Who is it for?
- Adopt BettaFish if you want an opinion report you can inspect as HTML in your repository, if you can supply one LLM endpoint per agent role, and if you are willing to own the crawling side yourself, since a 24 hour crawler against Weibo, Xiaohongshu, Douyin and Kuaishou breaks the moment those platforms change.
- Can I use it commercially?
- Yes, with conditions. GPL-2.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
- Is it still maintained?
- Yes. The repository last received commits 14 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Seven LLM slots, and every agent is meant to run a different model
The .env.example file is the real configuration surface. It opens with host settings and a database block, then gives every agent its own key, base URL and model name:
# Insight Agent
INSIGHT_ENGINE_API_KEY=
INSIGHT_ENGINE_BASE_URL=
INSIGHT_ENGINE_MODEL_NAME=
# Media Agent
MEDIA_ENGINE_API_KEY=
MEDIA_ENGINE_BASE_URL=
MEDIA_ENGINE_MODEL_NAME=The same triple repeats for the Query Agent, the Report Agent, MindSpider, the forum moderator and a SQL keyword optimizer, so seven separate endpoints in all, while the project describes itself as built around five categories of agent. Every call goes through the OpenAI request format, so any compatible service works, and a comment recommends each model by name, from kimi-k2 for the Insight Agent to gemini-2.5-pro for the Report Agent and qwen-plus for the moderator.
The comment above the Report Agent slot is the useful warning: that agent needs a strong model, and blank charts or broken paragraphs in the finished report are the symptom of an undersized one. Budget for it accordingly, because the report stage is where a weak model shows up rather than the search stage.
The forum loop is unbounded in the visible documentation
A run starts at a Flask app, launches three agents in parallel for a first pass, lets each draft a chunked research plan, and then enters a loop labelled 5-N. Inside each round the agents do deep research, the ForumEngine watches what they say and an LLM moderator produces guidance, and the agents adjust direction using a `forum_reader` tool.
The stated reason for the loop is to avoid the homogeneity a single model produces when it works alone. The cost of it is arithmetic: every round is a full extra pass of searching and reading, across three agents, plus a moderator call, and the rounds are described as multi-round without an upper bound. Nothing in the visible documentation sets a maximum iteration count, a wall-clock cap, or a token budget for a single analysis. You control the ceiling yourself through the models you configure and whatever you wrap the run in.
That is the single most important operational fact about this project, and it is the one the documentation leaves open.
Reports are bound into an IR document before a template sees them
The last three steps of the flow are the interesting engineering. The Report Agent collects everything the other agents and the forum produced. It then chooses a template and a style, generates metadata across several passes, and binds the result into an IR intermediate representation. Only after that is the report rendered, in chunks, each chunk quality-checked, into an interactive HTML file.
The IR is a buffer between research and presentation, and it buys two things. Template selection becomes a runtime decision rather than a fixed page, and a generation failure can be retried at the chunk level without re-running the crawl. It also means the report format is a separate concern from the agents, which is why the repository carries `templates/`, `static/` and a `regenerate_latest_html.py` script at the root.
The output is a file on disk with a timestamp in its name, such as `final_report__20250827_131630.html` in `final_reports/`, which the compose file mounts out to the host. Reports are reviewable artifacts in your repository, not a dashboard you have to keep running.
Compose publishes the database on 5444 while the config says 5432
The quick start is two steps. Copy `.env.example` to `.env`, fill it in, then start everything:
docker compose up -dThe compose file defines two services. The application runs from `ghcr.io/666ghj/bettafish:latest` and publishes 5000 for Flask plus 8501, 8502 and 8503. The database is `postgres:15`, named `bettfish-db`, and this is where the trap is:
ports:
- "${POSTGRES_PORT:-5444}:5432"So the database is reachable from the host on 5444, while `.env.example` sets `DB_PORT=5432` and the configuration table calls 5432 the default PostgreSQL port. Inside the container the app reaches the database through the service name `db` on 5432 and is unaffected. Run `app.py` on your host instead, and a `.env` left at 5432 finds nothing. Set `DB_PORT=5444` or set `POSTGRES_PORT` yourself.
The compose file also notes that image pulls are slow and ships a commented mirror, `ghcr.nju.edu.cn/666ghj/bettafish:latest`, ready to be uncommented.
Crawling needs Playwright chromium, and the platforms decide the outcome
Data collection runs through a MindSpider crawler built on Playwright. The source setup has a dedicated step for it:
# 安装浏览器驱动(用于爬虫功能)
playwright install chromiumThe dependency is pinned at `playwright==1.45.0`, and the Dockerfile sets `PLAYWRIGHT_BROWSERS_PATH=/ms-playwright` while installing the browser binaries and a long apt list that includes `ffmpeg` and the GTK and Pango libraries, so short video is part of the intended input rather than an accident.
That is where the fragility lives. The project describes a crawler cluster running around the clock across more than ten platforms including Weibo, Xiaohongshu, Douyin and Kuaishou, and reads not only posts but the mass of user comments underneath them. A headless chromium from 2024 is what stands between you and that data, and every one of those sites is free to change its page structure, its login wall or its rate limiting. Nothing in the visible documentation describes a selector-fallback strategy or a rate-limit backoff, so when collection returns empty you are debugging a browser session, not a query.
The machine learning block is optional, the PDF path is not
requirements.txt groups its entries and marks two of them as removable. The machine learning block pulls `torch>=2.0.0` on the CPU build, plus `transformers`, `sentence-transformers`, `scikit-learn` and `xgboost`, and the note says you can comment out that section if you do not want the local sentiment model, since its compute cost is described as small. A separate comment gives the GPU command, using the cu126 wheel index.
The PDF path is harder to skip. WeasyPrint needs system libraries that pip cannot install, which is why an optional system-dependency step exists with its own README, and why the Dockerfile installs `libpango`, `libcairo2`, `libgdk-pixbuf` and the rest before Python is set up. Skipping that step is the documented cause of a WeasyPrint install failure and a PDF feature that does not work. `export_pdf.py` sits at the repository root next to `regenerate_latest_html.py` and `regenerate_latest_md.py`, so HTML, Markdown and PDF are three parallel outputs of the same report.
One detail worth copying: the database is not initialised by hand. Running `app.py` detects and creates it, with `DB_CHARSET=utf8mb4` set for emoji and `DB_DIALECT` switching between postgresql and mysql.
Four listening ports for one analysis, three of them Streamlit
Port 5000 is the Flask application. Ports 8501, 8502 and 8503 are Streamlit, one per engine, and the compose file mounts a matching host directory for each: `insight_engine_streamlit_reports/`, `media_engine_streamlit_reports/` and `query_engine_streamlit_reports/`. The dependency list pins `streamlit==1.28.1` alongside `flask==2.3.3`, `flask-socketio==5.3.6` and `eventlet==0.33.3`, so this is a socket.io app with three separate Streamlit surfaces bolted beside it.
The compose file sets `STREAMLIT_SERVER_ENABLE_FILE_WATCHER=false` and `PYTHONUNBUFFERED=1`. The first stops Streamlit's file watcher from reacting to the report files the run itself writes into those mounted directories, and the second keeps container logs unbuffered so you can follow a long analysis in `docker logs`.
The repository also ships `SingleEngineApp/` and `report_engine_only.py`, which means you can run one agent or only the report stage instead of the whole panel. That is the practical way to iterate on a report template without paying for another full crawl.
Three releases inside one month, then pushes with no release
The repository is not archived, and the last push was on 2026-09-16. The releases tell a different story: v2.0.0 on 2025-11-28, v2.1.0 on 2025-12-09, and v3.0.0 on 2025-12-23. Three releases in twenty-six days, then nothing since December while the branch kept moving. So `main` is not v3.0.0, and if you need a known state you have to pin the tag yourself and accept that it predates most of the recent work.
Licensing is the part to settle before anything else. The project is GPL-2.0, which is copyleft, and a `THIRD_PARTY_NOTICES.md` sits next to the `LICENSE` file at the root. What that means for you depends entirely on whether you distribute the software or run it as a service, so treat the licence text as a question for your own counsel rather than something this repository can answer.
The scope claim is also worth reading precisely. The project describes itself as built from zero without depending on any framework, yet requirements.txt pins a large set of libraries including `tavily-python>=0.3.0` for search, `jieba==0.42.1` for Chinese segmentation, `plotly` and `wordcloud` for output, and `pyexecjs` and `xhshow` for crawling work. The no-framework claim means no agent orchestration framework, not no dependencies.
Editorial conclusion
Adopt BettaFish if you want an opinion report you can inspect as HTML in your repository, if you can supply one LLM endpoint per agent role, and if you are willing to own the crawling side yourself, since a 24 hour crawler against Weibo, Xiaohongshu, Douyin and Kuaishou breaks the moment those platforms change. Do not adopt it to embed in a closed-source product without reading GPL-2.0 first, and do not adopt it expecting a bounded run, because the forum loop has no round limit stated anywhere in the visible documentation. Verify three things in order: that `docker compose up -d` leaves port 5444 free for the database when your app runs on the host, that the Report Agent's model is strong enough to avoid blank charts in the output, and that the HTML file in final_reports/ that you are shown as the example is one your own models would have produced.
Frequently asked questions
what is bettafish github
666ghj/BettaFish is a Python multi-agent system for public opinion analysis, published under the name 微舆 and described as built from zero without depending on any agent framework. It crawls posts and comments across more than ten social platforms, has its agents debate in a forum loop moderated by an LLM, and renders an interactive HTML report.
How many LLM API keys does BettaFish need?
The .env.example file defines seven separate API key, base URL and model name triples: the Insight, Media, Query, Report, MindSpider, forum moderator and SQL keyword optimizer. Every call uses the OpenAI request format, so any compatible service can be substituted.
Can BettaFish run without Docker?
Yes. The source path needs Python 3.9 or higher, Conda or uv, PostgreSQL or MySQL, and about 2GB of memory, and it runs on Windows, Linux and MacOS. Dependencies come from `pip install -r requirements.txt` and the crawler needs `playwright install chromium`.
What licence does BettaFish use?
The repository states GPL-2.0 and points at a LICENSE file at the root, with a THIRD_PARTY_NOTICES.md alongside it. Because GPL-2.0 is copyleft, check how it applies to your own distribution before embedding the code.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/666ghj-bettafish)