Usagi-org/ai-goofish-monitor: AI-driven Goofish monitoring with a web UI
基于 Playwright 和AI实现的闲鱼多任务实时/定时监控与智能分析系统,配备了功能完善的后台管理UI。帮助用户从闲鱼海量商品中,找到心仪产品。
At a glance
- What is it?
- A Python and FastAPI system that drives Chromium through Playwright to watch Xianyu listings, filters them with a vision-capable model, and pushes hits to ntfy, Bark, WeCom or Telegram. Here is how it installs, where it breaks, and who should skip it.
- Who is it for?
- Adopt it if you already run a Docker host, hold a vision-capable OpenAI-compatible endpoint, and want scheduled Xianyu searches with AI filtering and push notifications rather than a spreadsheet of links. Skip it if you need a supported commercial product, if you cannot supply a working Xianyu login state, or if you object to driving a marketplace through a browser session.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 135 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 26, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What ai-goofish-monitor actually automates
Searching Xianyu by hand means reopening the same keyword, scrolling past the same stale listings, and reading every description to work out whether a seller is real. The project replaces that loop with scheduled jobs. You define a task with a keyword, a price band, a publish-time window, an optional province/city/district filter, and either a keyword rule or an AI prompt. The system then runs the search on a timer, pulls new items, downloads their images, asks a multimodal model whether each one matches your stated requirement, and sends a notification only for the ones the model recommends.
The intended user is a single operator or a small team tracking a niche: a specific camera body, a mechanical keyboard switch, a discontinued part. The web UI is the control surface for tasks, accounts, prompts, logs and results, and the README describes it as covering task management, account management, AI standard editing, run logs and result browsing. It is not a general marketplace analytics platform. There is no dashboard of aggregate price trends, and the price history it stores exists to serve the monitoring loop rather than to be queried as a dataset.
How the monitoring pipeline moves from search to notification
The README's workflow diagram is the clearest description of the architecture. A task starts, picks an account and proxy configuration, searches the marketplace, and asks whether a new item appeared. If not, it pages or waits and searches again. If a new item appears, it fetches the item detail and seller information, downloads the product images, and calls the AI for analysis. A positive AI verdict triggers a notification; either verdict writes a record to SQLite. A risk-control or exception branch rotates the account or proxy and retries.
The storage split matters more than it first appears. SQLite at data/app.sqlite3 is the live primary store for tasks, results and price history, and the application creates the schema on startup. Login cookies, prompts, logs and images stay on the filesystem in state/, prompts/, logs/ and images/. Product images land in images/task_images_<task_name>/ and the README says they are cleaned up by default when the task ends. The task creation API reflects the two decision paths: POST /api/tasks/generate returns 202 with a job when decision_mode is ai, and returns the created task directly when decision_mode is keyword. Progress for the AI path is polled through GET /api/tasks/generate-jobs/{job_id}.
One consequence is easy to miss. Because results are read from SQLite rather than scanned from jsonl files, the old config.json, jsonl/ and price_history/ directories exist only as a one-time import source. The README states the app attempts that import once at startup. If you delete them before the first successful boot, you lose the history they held.
Installing ai-goofish-monitor with Docker and creating a first task
Docker is the recommended path and the image ships with Chromium, so the host does not need a browser installed. Clone the repository, copy the environment template, fill in the AI credentials, and bring the stack up.
git clone https://github.com/Usagi-org/ai-goofish-monitor && cd ai-goofish-monitor
cp .env.example .env
vim .env
docker compose up -d
docker compose logs -f appThe minimum configuration the README lists is three AI variables plus the web credentials. OPENAI_API_KEY, OPENAI_BASE_URL and OPENAI_MODEL_NAME are all required, and the model must accept image input, because the analysis step sends downloaded product images. WEB_USERNAME and WEB_PASSWORD default to admin/admin123.
After the container is up, the web UI answers on http://127.0.0.1:8000. Log in, open the Xianyu account page, and paste the login state JSON exported by the Chrome extension the README links to. The state file is written to state/acc_1.json. Then go to task management, create a task, bind the account, and run it. The README notes that if you change SERVER_PORT in .env you must also update the port mapping in docker-compose.yaml, and that the compose file mounts ./data to /app/data for the SQLite database.
If the image pull is slow, the README offers a mirror route: pull ghcr.nju.edu.cn/usagi-org/ai-goofish:latest, retag it as ghcr.io/usagi-org/ai-goofish:latest, and run compose again. Updates go through docker compose pull followed by docker compose up -d.
Running from source, and the checks start.sh performs
The source path is heavier and the README documents it as a developer route. It requires Python 3.10 or newer, Node.js and npm, and Playwright with Chromium. The README suggests installing both before the first run.
python3 -m pip install playwright && python3 -m playwright install chromium
chmod +x start.sh
./start.shstart.sh checks for the Playwright CLI and a browser first, then installs dependencies, builds the frontend, copies the build output, and starts the backend. The frontend build produces web-ui/dist/, which start.sh copies to the repository root dist/, and FastAPI serves dist/index.html and dist/assets/. If you open the page and get a message that the frontend build output is missing, the root dist/ directory is absent; running ./start.sh again, or running npm run build inside web-ui/ and confirming the copy, is the fix the README gives.
Manual startup splits the two halves. The backend runs as python -m src.app or as uvicorn src.app:app --host 0.0.0.0 --port 8000 --reload, and the frontend development server runs from web-ui with npm install and npm run dev. Vite proxies /api, /auth and /ws to http://127.0.0.1:8000. Tests run with PYTEST_DISABLE_PLUGIN_AUTOLOAD=1 pytest, and pyproject.toml registers live and live_slow markers for smoke tests that need real credentials and external services. That marker split is a useful signal: the maintainers separate tests that hit the real marketplace from tests that do not, which implies the offline suite does not cover the scraping path.
Where the browser-session design becomes a liability
The whole system depends on a captured login state. If that cookie expires or the account is challenged, tasks stop returning results, and the failure surfaces as a login error in the log page rather than as an exception you can catch upstream. The README points at the log page as the place to diagnose login-state failures, risk control and AI call problems, which is honest but also means unattended operation needs someone reading logs.
The README's own notice says to respect Xianyu's terms of service and robots.txt, not to request too frequently, and that the project is for learning and technical research. Treat that as a design constraint rather than boilerplate. There are failure guards in the configuration, including TASK_FAILURE_THRESHOLD, TASK_FAILURE_PAUSE_SECONDS and TASK_FAILURE_GUARD_PATH, plus proxy rotation settings such as PROXY_ROTATION_ENABLED, PROXY_ROTATION_MODE, PROXY_POOL and PROXY_BLACKLIST_TTL. Those exist because the failure mode is real. A cheap proxy pool will not make an aggressive polling interval safe.
The AI layer is the other soft spot. Every new item costs an image-bearing model call, so the running cost scales with how many listings your keywords surface, not with how many you care about. Broad keywords plus a short schedule is the expensive combination, and the README does not publish any guidance on expected token volume per task. The regional filter is the cheapest lever available: the README states it significantly narrows the result set and recommends leaving it empty only when you are surveying the whole market.
How it differs from a plain scraper such as xianyu_spider
The README credits superboyyy/xianyu_spider as a reference, and the difference in approach is worth stating plainly. A conventional Xianyu scraper returns a list of matching items and leaves the judgement to you; you read titles and prices and decide. ai-goofish-monitor inserts a model between discovery and delivery, so the notification you receive has already been through a prompt you wrote. That is the entire value proposition, and it is also the entire risk: a badly worded prompt produces confident wrong recommendations, and the README's answer is a dedicated AI standard editor in the web UI plus a background job that generates the analysis standard from a natural-language requirement.
Compared with a general-purpose browser automation framework, the trade is the opposite. Playwright alone gives you full control and no opinion about Xianyu; this project gives you task scheduling, account rotation, notification channels, a result store and a UI, at the cost of accepting its data model and its assumptions about how listings are fetched. If you need to monitor a different marketplace, the task layer is not portable. If you need a one-off scrape, the container and the AI key are overhead you do not need.
Licence, maintenance and the cost of upgrading
The project is MIT licensed and the README states it is provided as is, without warranty, with the authors disclaiming liability for damages arising from use. MIT is permissive: you can modify and redistribute it, including commercially, provided the licence and copyright notice travel with the code. That is a statement about the licence text, not legal advice, and it says nothing about whether your use of the marketplace complies with its terms.
The last push to the default branch was on 2026-05-18, and the most recent release listed is 2.4 from 2026-04-27. The repository is not archived. Those dates are the only maintenance signal available here; the README does not publish a support policy or a deprecation schedule.
Upgrade cost concentrates in the storage migration. Moving from the pre-SQLite layout means the app imports config.json, jsonl/ and price_history/ once at startup, and the README advises keeping those mounts until you have confirmed the data in data/app.sqlite3, then deciding whether to keep them. Custom prompts live in prompts/ and login states in state/, both outside SQLite, so a container replacement preserves them only if those directories are mounted. The compose file mounts both by default. If you write your own compose file, carry those volumes over or you will re-authenticate every account after each upgrade.
Editorial conclusion
Adopt it if you already run a Docker host, hold a vision-capable OpenAI-compatible endpoint, and want scheduled Xianyu searches with AI filtering and push notifications rather than a spreadsheet of links. Skip it if you need a supported commercial product, if you cannot supply a working Xianyu login state, or if you object to driving a marketplace through a browser session. Before trusting it, verify three things on your own machine: that state/acc_1.json appears after you paste the extension JSON, that data/app.sqlite3 is created on first boot, and that one keyword-mode task returns a result row without the log page showing a login failure.
Frequently asked questions
What does ai-goofish-monitor need before it can run a task?
The README lists OPENAI_API_KEY, OPENAI_BASE_URL and OPENAI_MODEL_NAME as required, and the model must support image input because product images are sent for analysis. You also need a Xianyu login state, exported with the Chrome extension and pasted into the account page, which is saved to state/acc_1.json.
How do I deploy ai-goofish-monitor with Docker?
Clone the repository, copy .env.example to .env and fill in the configuration, then run docker compose up -d. The image includes Chromium, the web UI listens on http://127.0.0.1:8000, and the compose file mounts ./data to /app/data for the SQLite database.
Why did my ai-goofish-monitor task stop returning results?
The most common cause the README points to is an expired login state or marketplace risk control, and it directs you to the log page to diagnose login-state failures, risk control and AI call problems. The configuration also exposes failure guards such as TASK_FAILURE_THRESHOLD and TASK_FAILURE_PAUSE_SECONDS, plus proxy rotation settings.
Does ai-goofish-monitor store results in SQLite?
Yes. The README states that SQLite at data/app.sqlite3 is the live primary store for tasks, results and price history, and that the schema is created on startup. Login cookies, prompts, logs and images remain filesystem directories under state/, prompts/, logs/ and images/.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/usagi-org-ai-goofish-monitor)