Model or dataset
jmuncor/tokentap avatar
jmuncor/tokentap

tokentap: a local HTTP proxy that counts LLM tokens in your terminal

Intercept LLM API traffic and visualize token usage in a real-time terminal dashboard. Track costs, debug prompts, and monitor context window usage across your AI development sessions.

814 stars38 forksPythonMIT

At a glance

What is it?
tokentap sits between your LLM CLI tool and the provider API, forwarding requests while a Rich dashboard shows cumulative token use against a configurable limit. It works today for Claude Code, Codex and MiniMax; Gemini CLI support is blocked by an upstream bug.
Who is it for?
Adopt tokentap if you run Claude Code or OpenAI Codex from a terminal and want per-request token counts and a saved prompt archive without configuring TLS certificates. Skip it if your workflow depends on Gemini CLI, since the README states that Gemini CLI ignores custom base URLs under OAuth, or if you need a long-lived accounting system rather than a per-session view.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 100 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

The problem tokentap targets: invisible token spend inside CLI coding agents

Claude Code, Codex and similar CLI agents send requests you never see. The provider bills per token, the context window fills silently, and by the time a session ends you have no record of what was sent. tokentap exists for that gap. It is a local HTTP proxy plus a terminal dashboard, aimed at developers who run LLM CLI tools interactively and want a running count during the session rather than a bill at the end of the month. The README frames the value in four bullets: track token usage per request, monitor context windows with a fuel gauge, save every prompt as markdown and JSON, and require no certificate setup. That last point matters. Tools in this space often ask you to install a CA certificate so the proxy can read TLS traffic. tokentap avoids that by having the CLI tool talk plain HTTP to localhost, then forwarding to the provider over HTTPS itself.

How the proxy and dashboard fit together

The architecture is two terminals. In the first, tokentap start launches an HTTP proxy on localhost:8080 alongside the dashboard and the prompt archive. In the second, tokentap claude sets ANTHROPIC_BASE_URL to http://localhost:8080 and then runs the claude binary. Requests flow from the CLI to the local proxy over plain HTTP, and the proxy forwards them to api.anthropic.com over HTTPS. Because the CLI is pointed at a base URL rather than a socket, no certificate is involved on the client side. For OpenAI-compatible providers the routing changes shape. The README documents path-prefix routing for MiniMax: tokentap run --provider minimax sets OPENAI_BASE_URL to http://localhost:8080/minimax/v1, a request arrives at /minimax/v1/chat/completions, the proxy strips the prefix and forwards to https://api.minimax.io/v1/chat/completions. That prefix is what lets one proxy process serve more than one upstream. Token counting itself comes from tiktoken, listed in pyproject.toml alongside aiohttp, rich and click. The dashboard renders through Rich; the proxy is async via aiohttp. The gauge is colour-coded: green below 50 percent of the limit, yellow from 50 to 80, red above 80. The limit defaults to 200000 tokens and is set with -l.

Installing tokentap and capturing your first Claude Code session

The README gives a single pip command for installation, with a source install as the alternative. Python 3.10 or newer is required, and pyproject.toml confirms requires-python >=3.10 with classifiers through 3.13.

bash
pip install tokentap

After installation the tokentap console script is available, mapped in pyproject.toml to tokentap.cli:main. Start the proxy and dashboard in one terminal. The README states you will be prompted to choose where captured prompts are saved, and that -p sets the proxy port (default 8080) while -l sets the token limit used by the fuel gauge (default 200000).

bash
tokentap start -l 200000

In a second terminal, launch Claude Code through the wrapper. The README says this sets ANTHROPIC_BASE_URL to http://localhost:8080 and then runs claude.

bash
tokentap claude

As you work, the dashboard adds a row per request with a timestamp, provider, model name and token count, and the gauge tracks cumulative usage against the limit. On exit the README shows a session summary line in the form "Session complete. Total: 84,231 tokens across 12 requests." Every intercepted request is also written to the directory you chose, once as human-readable markdown with metadata and once as the raw JSON API request body. For an OpenAI-compatible tool, the same pattern uses tokentap run with a provider name; the README lists anthropic, openai, gemini and minimax as accepted values for --provider.

The Gemini CLI gap and other limits worth knowing before you commit

The README is unusually direct about a broken path. Gemini CLI is listed as "Blocked by upstream issue" in the supported providers table, and the known issues section explains why: Gemini CLI ignores custom base URLs when using OAuth authentication, with a link to google-gemini/gemini-cli issue 15430. The README says tokentap's Gemini support will work automatically once that issue is fixed. Until then, tokentap gemini is not a path you can rely on, and the project cannot fix it from its own side. Two other constraints follow from the design. First, the proxy is plain HTTP on localhost, so anything on the machine that can reach port 8080 can send requests through it; the README does not document authentication on the proxy. Second, the prompt archive writes full request bodies to disk, so the directory you pick inherits whatever sensitivity your prompts have. The README does not document redaction, retention limits or an option to skip archiving. Token counting is also an estimate in the sense that it is computed locally with tiktoken rather than read back from the provider's usage field, and the README does not describe reconciliation against provider-reported counts.

How tokentap differs from provider dashboards and from mitmproxy setups

The obvious alternative is the usage dashboard each provider already offers. Those report after the fact, on a billing cycle, aggregated across every key on the account, and they cannot show you which prompt in a live session pushed you over a threshold. tokentap's difference is position: it sits in the request path, so it can attribute tokens to a single request and display them while the session is still running. The second alternative is a general interception proxy such as mitmproxy, which the repository topics list alongside tokentap's own tags. mitmproxy is a full traffic inspection tool that decrypts TLS through a CA certificate you install and trust, and it gives you scripting hooks for arbitrary protocols. tokentap trades that generality for a narrower job: it does not decrypt anything, it relies on the client being configurable via a base URL, and in exchange the setup is a pip install and two commands. If you need to inspect headers, rewrite bodies or capture non-LLM traffic, mitmproxy is the better instrument. If you want a token counter and a prompt archive for a CLI coding agent, the certificate step is overhead you do not need.

Maintenance, release cadence and what the MIT licence covers

The repository is not archived. Its last push was on 2026-06-21, which is roughly three months before the date of this article. Two releases are listed: v0.1.0 on 2026-02-01 and v0.1.1 on 2026-04-03. pyproject.toml carries the classifier "Development Status :: 4 - Beta" and pins version 0.1.1, so the published package and the repository agree. The dependency set is short and mainstream: aiohttp, rich, tiktoken and click, with pytest, pytest-asyncio, build and twine under the dev extra. Upgrade cost is therefore close to the cost of those four libraries, and the CLI surface is small enough that a breaking change would be visible in the commands table. On licensing, tokentap is MIT, which permits commercial and private use and modification provided the copyright notice and permission notice are retained. That is a statement about the licence text, not legal advice; if you redistribute tokentap inside a product, read the LICENSE file in the repository yourself. One thing the README does not document is a rollback path if a new version changes proxy behaviour, so pinning the version in your own environment is the only lever the material describes.

Editorial conclusion

Adopt tokentap if you run Claude Code or OpenAI Codex from a terminal and want per-request token counts and a saved prompt archive without configuring TLS certificates. Skip it if your workflow depends on Gemini CLI, since the README states that Gemini CLI ignores custom base URLs under OAuth, or if you need a long-lived accounting system rather than a per-session view. Before trusting the numbers, run tokentap start -l with your model's real context limit and confirm the dashboard's totals match the provider's own usage report for one request.

Frequently asked questions

What is tokentap used for?

tokentap intercepts LLM API traffic from CLI tools and shows token usage in a live terminal dashboard. It also saves every intercepted prompt as markdown and JSON so you can review what was sent.

How do I install tokentap?

The README gives pip install tokentap, or a source install via git clone followed by pip install -e . Python 3.10 or newer is required.

Which LLM CLI tools does tokentap support?

The supported providers table lists Anthropic (Claude Code), OpenAI (Codex) and MiniMax as supported, and Google (Gemini CLI) as blocked by an upstream issue. Gemini CLI ignores custom base URLs when using OAuth authentication.

Does tokentap need a certificate to read HTTPS traffic?

No. The README states there are no certificates and no setup. The CLI tool is pointed at a plain HTTP proxy on localhost, and the proxy forwards to the provider over HTTPS.

What is the default proxy port and token limit?

The README documents -p, --port with a default of 8080 and -l, --limit with a default of 200000 for the fuel gauge. Both are options on tokentap start.

Official sources

  1. jmuncor/tokentap on GitHub
  2. License: MIT
  3. Project website
  4. README
  5. Releases
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/jmuncor-tokentap.svg)](https://hysenlabs.com/projects/jmuncor-tokentap)
Community notes

Community notes