onWatch: a local daemon that tracks AI API quotas across ten providers
Track AI API quotas across Synthetic, Z.ai, Anthropic (Claude Code), Codex, GitHub Copilot & Antigravity in real time. Lightweight background daemon (<50MB RAM), SQLite storage, Material Design 3 dashboard. Zero telemetry.
At a glance
- What is it?
- onWatch is a Go daemon that polls your Synthetic, Z.ai, Anthropic, Codex, GitHub Copilot and Antigravity keys, stores the history in SQLite and serves a dashboard on port 9211. It is useful when you need per-cycle history rather than a current snapshot, and it is the wrong tool if you want a hosted service that watches keys on machines you do not own.
- Who is it for?
- Adopt onWatch if you run several AI coding tools on one machine and keep hitting throttling without knowing which key caused it. Skip it if you need a hosted, multi-tenant view of keys on machines you do not control, because the design assumes a local daemon reading local credential files.
- Can I use it commercially?
- Yes, with conditions. GPL-3.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly Go, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The gap onWatch fills between a usage snapshot and a billing surprise
Provider dashboards answer one question: how much have I used right now. They rarely answer the question that matters before throttling, which is how this cycle compares with the last one, and which of the four tools I have open is responsible. onWatch is built for that second question. It runs as a background agent, polls each configured provider, and keeps the results in SQLite so you can look at history per cycle rather than a live number that resets.
The audience is narrow and specific: developers who hold keys for Synthetic, Z.ai, Anthropic (Claude Code), Codex, GitHub Copilot, MiniMax, Gemini CLI, Cursor, Grok, Kimi Code or Antigravity, often several at once, and who work through tools such as Cline, Roo Code, Kilo Code, Claude Code, Codex CLI or Cursor. If you use exactly one provider and never approach its limit, the daemon adds a process for no benefit. The README states the memory target as under 50 MB with all providers polling in parallel, which is the number to hold it to when you decide whether the daemon is worth a permanent slot on your machine.
How the daemon, SQLite store and port 9211 dashboard fit together
The architecture is a single Go binary with three moving parts. A poller reads credentials and queries each provider's usage endpoint. A SQLite database, written through modernc.org/sqlite, holds the history. An HTTP server serves a Material Design 3 dashboard with dark and light modes. Nothing leaves the machine, and the README states zero telemetry. That claim is structural rather than a policy promise: the binary has no analytics dependency in go.mod, so there is no client library present to send anything.
Provider credentials come from different places, and this is where the design gets interesting. Anthropic tokens are auto-detected from Claude Code credentials, read from the macOS Keychain, a Linux keyring, or ~/.claude/.credentials.json. Codex tokens can be re-read from ~/.codex/auth.json while the daemon runs, but the README recommends setting CODEX_TOKEN in .env so a Codex-only startup works reliably. OpenCode Go is tracked differently again: it requires OPENCODE_GO_WORKSPACE_ID and a browser cookie named auth from opencode.ai, which means that provider is a dashboard scrape rather than an API call. Scraping a cookie-authenticated page is the most fragile of the integrations, and the README itself calls the GitHub Copilot support Beta.
The repository also carries a Prometheus client in its dependencies, so metrics export exists alongside the dashboard. The README does not document a metrics endpoint path, so treat that as something to confirm from the source before building alerting on it.
Installing onWatch and confirming it tracks a real key
The fastest path on macOS and Linux is the one-line installer, which the README says downloads the binary to ~/.onwatch/, creates a .env config, sets up a systemd service on Linux or a launchd agent on macOS, and adds onwatch to your PATH.
curl -fsSL https://raw.githubusercontent.com/onllm-dev/onwatch/main/install.sh | bashIf you prefer a package manager, Homebrew is supported and the setup wizard walks through API keys and config.
brew install onllm-dev/tap/onwatch
onwatch setupWindows has a PowerShell equivalent, and the README points at docs/WINDOWS_SETUP.md for manual setup and troubleshooting.
irm https://raw.githubusercontent.com/onllm-dev/onwatch/main/install.ps1 | iexAfter installation, edit ~/.onwatch/.env and set at least one provider key. The README gives these keys, and at least one is required.
SYNTHETIC_API_KEY=syn_your_key_here
ZAI_API_KEY=your_zai_key_here
ANTHROPIC_TOKEN=your_token_here
CODEX_TOKEN=your_token_here
COPILOT_TOKEN=ghp_your_token_here
ONWATCH_ADMIN_USER=admin
ONWATCH_ADMIN_PASS=changemeThe ANTHROPIC_TOKEN line can be left empty if you use Claude Code, because the daemon auto-detects those credentials. For a first real check, build from source and run with debug output, which the README shows as the manual path.
git clone https://github.com/onllm-dev/onwatch.git && cd onwatch
cp .env.example .env
./app.sh --build && ./onwatch --debugWith one key configured, you should see the poller log a request for that provider and the dashboard report a usage figure for the current cycle. Compare that figure against the provider's own usage page before you configure the remaining providers. If the two disagree, the integration is the problem, not your budget.
Docker deployment, the /data volume and the 64 MB memory limit
The Docker path is documented as a compose file that builds the image, maps host port 9211 to container port 9211 by default through ONWATCH_PORT, and mounts ./onwatch-data at /data. The environment block forces ONWATCH_DB_PATH to /data/onwatch.db so the database lands on the mounted volume rather than inside the container. The compose file sets a memory limit of 64M with a 32M reservation, which is tighter than the README's under-50 MB figure and worth knowing before you add providers.
The image runs as a non-root user with UID 65532, and the Dockerfile creates /data owned by that UID. That detail matters for bind mounts: if you create the directory yourself, the compose comment says to pre-create it with mkdir -p ./onwatch-data && chown -R 65532:65532 ./onwatch-data, otherwise SQLite will fail to write. The default image is distroless at roughly 10 to 12 MB, and an Alpine variant with a shell is published as ghcr.io/onllm-dev/onwatch:alpine for cases where you need docker exec. Logs go to stdout, so docker logs -f onwatch is the way to read them.
Where onWatch is the wrong tool
The daemon assumes it can read credentials from the machine it runs on. If your team's keys live in a shared secret manager and developers work on managed laptops where you cannot install a launchd agent or systemd unit, onWatch has no path to those keys and no hosted mode to fall back on. The README describes a local agent, a local database and a local dashboard, and nothing in the repository layout suggests a server-side collector.
Cookie-based integrations are the second limit. OpenCode Go needs a browser cookie named auth copied into .env, and cookies expire. When that value goes stale, the provider silently stops reporting rather than failing loudly, and the README does not document a re-authentication flow for it. The same caution applies to the GitHub Copilot integration, which the README labels Beta. If you depend on either provider for budget decisions, verify the freshness of the data rather than assuming the dashboard is current.
Finally, the README states the project is in active development and that features and APIs may change. The go.mod module path carries a v2 major version, which is the conventional Go signal that the import path is not stable across major releases. Plan for upgrade work.
How onWatch differs from a general metrics stack
The obvious alternative is a general observability setup: a Prometheus server scraping an exporter, with Grafana in front. That approach gives you alerting rules, long retention and a query language, and it fits teams that already run that infrastructure. The difference is what has to be built. A general stack gives you the pipeline and the storage; you still have to write and maintain the per-provider collectors that know how Synthetic, Z.ai, Anthropic and Codex expose usage, and that knowledge is exactly what onWatch ships. The trade is the reverse of the usual one: you give up query flexibility and retention control in exchange for not owning ten provider integrations.
If your requirement is alerting on quota thresholds across a fleet of machines, the general stack is the better fit and onWatch is not. If your requirement is a developer seeing, on their own laptop, that the Codex key is at 80 percent of its cycle while the Anthropic key is untouched, onWatch does that without any infrastructure.
Licence, maintenance and what an upgrade costs
onWatch is GPL-3.0. For individual developers running the binary locally, the practical effect is that you receive the source and can modify it. The obligation that matters is distribution: if you ship a modified onWatch to others, or embed it in a product you distribute, the GPL requires you to make the corresponding source available under the same licence. Running it internally to watch your own keys does not trigger that. This is a description of the licence text, not legal advice, and anyone embedding the daemon in a commercial product should read the LICENSE file and, if the answer is not obvious, ask a lawyer.
The maintenance picture is favourable on the evidence available. The repository is not archived, the last push was on 2026-09-07, and the most recent release is v2.14.1 from the same day, with v2.14.0-beta.4 a few hours earlier. That cadence means upgrades arrive often, and the README's own note that APIs may change is the cost of that cadence. The upgrade path is cheap for the binary installs, since the installer and Homebrew both replace a single file, and the SQLite database persists across versions. For Docker, the compose file builds from source, so an upgrade means pulling the new source and rebuilding rather than pulling a tag. Back up the onwatch.db file before a major version bump, because a v2 to v3 migration is the point where a schema change would land.
Editorial conclusion
Adopt onWatch if you run several AI coding tools on one machine and keep hitting throttling without knowing which key caused it. Skip it if you need a hosted, multi-tenant view of keys on machines you do not control, because the design assumes a local daemon reading local credential files. Before trusting it, run ./onwatch --debug with a single provider key and confirm the dashboard reports a cycle that matches what the provider's own usage page shows.
Frequently asked questions
What is onWatch?
It is a free, open-source AI API quota monitor written in Go. It runs as a background daemon, polls the providers you configure, stores history in SQLite and serves a local dashboard on port 9211.
How do I install onWatch on macOS or Linux?
The README gives a one-line install that downloads the binary to ~/.onwatch/, creates a .env config and sets up a systemd service or launchd agent. Homebrew is also supported with brew install onllm-dev/tap/onwatch followed by onwatch setup.
Which providers can onWatch track?
The README lists Synthetic, Z.ai, Anthropic (Claude Code), Codex, GitHub Copilot, MiniMax, Gemini CLI (legacy), Cursor, Grok, Kimi Code and Antigravity. At least one provider key is required, and any combination can be polled in parallel.
Does onWatch send my usage data anywhere?
The README states zero telemetry and that all data stays on your machine, with history written to a local SQLite database. The dashboard is served locally, and the Docker compose file mounts the database at /data on the host.
How much memory does onWatch use?
The README describes it as a lightweight background agent using under 50 MB of RAM with all providers polling in parallel. The bundled docker-compose.yml sets a container memory limit of 64M with a 32M reservation.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/onllm-dev-onwatch)