# BettaFish runs a seven-slot LLM panel and a forum loop over social media

> BettaFish is a from-scratch Python multi-agent system that reads public posts and comments across Chinese and international social platforms, debates findings in a moderator-led forum loop, and renders an HTML report. It needs seven LLM endpoints, a PostgreSQL database and a Playwright browser, and its copyleft licence is the first thing to settle.

**666ghj/BettaFish** — 微舆：人人可用的多Agent舆情分析助手，打破信息茧房，还原舆情原貌，预测未来走向，辅助决策！从0实现，不依赖任何框架。

- Repository: https://github.com/666ghj/BettaFish
- Website: https://deepwiki.com/666ghj/BettaFish
- Stars: 42,324 · Forks: 7,618
- Language: Python
- License: GPL-2.0
- Published: 2026-08-24 · Updated: 2026-08-24 · Language: en
- Canonical page: https://hysenlabs.com/projects/666ghj-bettafish

## Seven LLM slots, and every agent is meant to run a different model

The .env.example file is the real configuration surface. It opens with host settings and a database block, then gives every agent its own key, base URL and model name:

```
# Insight Agent
INSIGHT_ENGINE_API_KEY=
INSIGHT_ENGINE_BASE_URL=
INSIGHT_ENGINE_MODEL_NAME=

# Media Agent
MEDIA_ENGINE_API_KEY=
MEDIA_ENGINE_BASE_URL=
MEDIA_ENGINE_MODEL_NAME=
```

The same triple repeats for the Query Agent, the Report Agent, MindSpider, the forum moderator and a SQL keyword optimizer, so seven separate endpoints in all, while the project describes itself as built around five categories of agent. Every call goes through the OpenAI request format, so any compatible service works, and a comment recommends each model by name, from kimi-k2 for the Insight Agent to gemini-2.5-pro for the Report Agent and qwen-plus for the moderator.

The comment above the Report Agent slot is the useful warning: that agent needs a strong model, and blank charts or broken paragraphs in the finished report are the symptom of an undersized one. Budget for it accordingly, because the report stage is where a weak model shows up rather than the search stage.

## The forum loop is unbounded in the visible documentation

A run starts at a Flask app, launches three agents in parallel for a first pass, lets each draft a chunked research plan, and then enters a loop labelled 5-N. Inside each round the agents do deep research, the ForumEngine watches what they say and an LLM moderator produces guidance, and the agents adjust direction using a `forum_reader` tool.

The stated reason for the loop is to avoid the homogeneity a single model produces when it works alone. The cost of it is arithmetic: every round is a full extra pass of searching and reading, across three agents, plus a moderator call, and the rounds are described as multi-round without an upper bound. Nothing in the visible documentation sets a maximum iteration count, a wall-clock cap, or a token budget for a single analysis. You control the ceiling yourself through the models you configure and whatever you wrap the run in.

That is the single most important operational fact about this project, and it is the one the documentation leaves open.

## Reports are bound into an IR document before a template sees them

The last three steps of the flow are the interesting engineering. The Report Agent collects everything the other agents and the forum produced. It then chooses a template and a style, generates metadata across several passes, and binds the result into an IR intermediate representation. Only after that is the report rendered, in chunks, each chunk quality-checked, into an interactive HTML file.

The IR is a buffer between research and presentation, and it buys two things. Template selection becomes a runtime decision rather than a fixed page, and a generation failure can be retried at the chunk level without re-running the crawl. It also means the report format is a separate concern from the agents, which is why the repository carries `templates/`, `static/` and a `regenerate_latest_html.py` script at the root.

The output is a file on disk with a timestamp in its name, such as `final_report__20250827_131630.html` in `final_reports/`, which the compose file mounts out to the host. Reports are reviewable artifacts in your repository, not a dashboard you have to keep running.

## Compose publishes the database on 5444 while the config says 5432

The quick start is two steps. Copy `.env.example` to `.env`, fill it in, then start everything:

```bash
docker compose up -d
```

The compose file defines two services. The application runs from `ghcr.io/666ghj/bettafish:latest` and publishes 5000 for Flask plus 8501, 8502 and 8503. The database is `postgres:15`, named `bettfish-db`, and this is where the trap is:

```yaml
    ports:
      - "${POSTGRES_PORT:-5444}:5432"
```

So the database is reachable from the host on 5444, while `.env.example` sets `DB_PORT=5432` and the configuration table calls 5432 the default PostgreSQL port. Inside the container the app reaches the database through the service name `db` on 5432 and is unaffected. Run `app.py` on your host instead, and a `.env` left at 5432 finds nothing. Set `DB_PORT=5444` or set `POSTGRES_PORT` yourself.

The compose file also notes that image pulls are slow and ships a commented mirror, `ghcr.nju.edu.cn/666ghj/bettafish:latest`, ready to be uncommented.

## Crawling needs Playwright chromium, and the platforms decide the outcome

Data collection runs through a MindSpider crawler built on Playwright. The source setup has a dedicated step for it:

```bash
# 安装浏览器驱动（用于爬虫功能）
playwright install chromium
```

The dependency is pinned at `playwright==1.45.0`, and the Dockerfile sets `PLAYWRIGHT_BROWSERS_PATH=/ms-playwright` while installing the browser binaries and a long apt list that includes `ffmpeg` and the GTK and Pango libraries, so short video is part of the intended input rather than an accident.

That is where the fragility lives. The project describes a crawler cluster running around the clock across more than ten platforms including Weibo, Xiaohongshu, Douyin and Kuaishou, and reads not only posts but the mass of user comments underneath them. A headless chromium from 2024 is what stands between you and that data, and every one of those sites is free to change its page structure, its login wall or its rate limiting. Nothing in the visible documentation describes a selector-fallback strategy or a rate-limit backoff, so when collection returns empty you are debugging a browser session, not a query.

## The machine learning block is optional, the PDF path is not

requirements.txt groups its entries and marks two of them as removable. The machine learning block pulls `torch>=2.0.0` on the CPU build, plus `transformers`, `sentence-transformers`, `scikit-learn` and `xgboost`, and the note says you can comment out that section if you do not want the local sentiment model, since its compute cost is described as small. A separate comment gives the GPU command, using the cu126 wheel index.

The PDF path is harder to skip. WeasyPrint needs system libraries that pip cannot install, which is why an optional system-dependency step exists with its own README, and why the Dockerfile installs `libpango`, `libcairo2`, `libgdk-pixbuf` and the rest before Python is set up. Skipping that step is the documented cause of a WeasyPrint install failure and a PDF feature that does not work. `export_pdf.py` sits at the repository root next to `regenerate_latest_html.py` and `regenerate_latest_md.py`, so HTML, Markdown and PDF are three parallel outputs of the same report.

One detail worth copying: the database is not initialised by hand. Running `app.py` detects and creates it, with `DB_CHARSET=utf8mb4` set for emoji and `DB_DIALECT` switching between postgresql and mysql.

## Four listening ports for one analysis, three of them Streamlit

Port 5000 is the Flask application. Ports 8501, 8502 and 8503 are Streamlit, one per engine, and the compose file mounts a matching host directory for each: `insight_engine_streamlit_reports/`, `media_engine_streamlit_reports/` and `query_engine_streamlit_reports/`. The dependency list pins `streamlit==1.28.1` alongside `flask==2.3.3`, `flask-socketio==5.3.6` and `eventlet==0.33.3`, so this is a socket.io app with three separate Streamlit surfaces bolted beside it.

The compose file sets `STREAMLIT_SERVER_ENABLE_FILE_WATCHER=false` and `PYTHONUNBUFFERED=1`. The first stops Streamlit's file watcher from reacting to the report files the run itself writes into those mounted directories, and the second keeps container logs unbuffered so you can follow a long analysis in `docker logs`.

The repository also ships `SingleEngineApp/` and `report_engine_only.py`, which means you can run one agent or only the report stage instead of the whole panel. That is the practical way to iterate on a report template without paying for another full crawl.

## Three releases inside one month, then pushes with no release

The repository is not archived, and the last push was on 2026-09-16. The releases tell a different story: v2.0.0 on 2025-11-28, v2.1.0 on 2025-12-09, and v3.0.0 on 2025-12-23. Three releases in twenty-six days, then nothing since December while the branch kept moving. So `main` is not v3.0.0, and if you need a known state you have to pin the tag yourself and accept that it predates most of the recent work.

Licensing is the part to settle before anything else. The project is GPL-2.0, which is copyleft, and a `THIRD_PARTY_NOTICES.md` sits next to the `LICENSE` file at the root. What that means for you depends entirely on whether you distribute the software or run it as a service, so treat the licence text as a question for your own counsel rather than something this repository can answer.

The scope claim is also worth reading precisely. The project describes itself as built from zero without depending on any framework, yet requirements.txt pins a large set of libraries including `tavily-python>=0.3.0` for search, `jieba==0.42.1` for Chinese segmentation, `plotly` and `wordcloud` for output, and `pyexecjs` and `xhshow` for crawling work. The no-framework claim means no agent orchestration framework, not no dependencies.

## Conclusion

Adopt BettaFish if you want an opinion report you can inspect as HTML in your repository, if you can supply one LLM endpoint per agent role, and if you are willing to own the crawling side yourself, since a 24 hour crawler against Weibo, Xiaohongshu, Douyin and Kuaishou breaks the moment those platforms change. Do not adopt it to embed in a closed-source product without reading GPL-2.0 first, and do not adopt it expecting a bounded run, because the forum loop has no round limit stated anywhere in the visible documentation. Verify three things in order: that `docker compose up -d` leaves port 5444 free for the database when your app runs on the host, that the Report Agent's model is strong enough to avoid blank charts in the output, and that the HTML file in final_reports/ that you are shown as the example is one your own models would have produced.

## FAQ

### what is bettafish github

666ghj/BettaFish is a Python multi-agent system for public opinion analysis, published under the name 微舆 and described as built from zero without depending on any agent framework. It crawls posts and comments across more than ten social platforms, has its agents debate in a forum loop moderated by an LLM, and renders an interactive HTML report.

### How many LLM API keys does BettaFish need?

The .env.example file defines seven separate API key, base URL and model name triples: the Insight, Media, Query, Report, MindSpider, forum moderator and SQL keyword optimizer. Every call uses the OpenAI request format, so any compatible service can be substituted.

### Can BettaFish run without Docker?

Yes. The source path needs Python 3.9 or higher, Conda or uv, PostgreSQL or MySQL, and about 2GB of memory, and it runs on Windows, Linux and MacOS. Dependencies come from `pip install -r requirements.txt` and the crawler needs `playwright install chromium`.

### What licence does BettaFish use?

The repository states GPL-2.0 and points at a LICENSE file at the root, with a THIRD_PARTY_NOTICES.md alongside it. Because GPL-2.0 is copyleft, check how it applies to your own distribution before embedding the code.

## Sources

- [Official documentation](https://deepwiki.com/666ghj/BettaFish)
- [Official README](https://github.com/666ghj/BettaFish#readme)
- [Project repository](https://github.com/666ghj/BettaFish)
- [Release notes](https://github.com/666ghj/BettaFish/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/666ghj-bettafish
