ingest turns a repository into one markdown file, and its token count is a scaled guess
Parse files (e.g. code repos) and websites to clipboard or a file for ingestions by AI / LLMs
At a glance
- What is it?
- A Go tool that traverses directories, optionally strips function bodies with Tree-sitter, and hands the result to the clipboard or an OpenAI compatible API. Because no Claude tokeniser exists, the offline count is an estimate with a correction factor that differs by model generation.
- Who is it for?
- This fits someone preparing a repository for a model in one command, since the tree view, the diff and the glob filter together are the part you would otherwise assemble by hand. It does not fit anyone who needs an exact context size, because the default count is an approximation with a documented error band, and the exact path needs an Anthropic key.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 51 days ago.
- What is it written in?
- Mainly Go, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
One command, and the count lands in the same terminal
The interface is a single binary with flags and a list of paths.
Run it with no arguments and it takes the current working directory, which is the fastest way to see what it produces. The output is a spinner while the directory is traversed and a tree built, then an approximate token figure, then confirmation that the text reached the clipboard.
ingest [flags] <paths>and installing it is one command if you have the Go toolchain:
go install github.com/sammcj/ingest@HEADAround that core there are selectors for what to include. A glob include filter narrows a run to one language, a diff flag adds a repository diff to the prompt, an output flag writes to a named file, and a save flag writes into a subdirectory named after the project under the home directory. Several paths can be given at once, and individual files mix with directories.
There is also a JSON output mode, customisable templates, shell completions for three shells, and exports to the console for piping.
The design assumption is that you want a prompt, not a report. Everything else in the feature list, web crawling and PDF conversion, feeds the same output.
The token count is offline, and the documentation disagrees about downloads
Token counting is the feature with the most nuance attached, because the honest answer is that it cannot be exact offline.
One section states that the two vocabularies it uses are compiled into the binary, so the tool works offline and never downloads or caches anything. A later section states that the first time it runs, it downloads a small tokeniser file for offline use.
Those cannot both be current, and which one is true depends on the build rather than on the flags. That is worth knowing before relying on the offline claim, and it is cheap to check: run it once with the network unavailable and see whether it still produces a count.
The counting itself is straightforward once a vocabulary is chosen. Counting for a non-Claude model uses that model's own vocabulary with no adjustment. Counting for a Claude model is different, because the vendor has never published a Claude tokeniser, so the count comes from the newer vocabulary and is then scaled.
Two correction factors, and they are not close together
The scaling is documented as two factors, and the split is by model generation rather than by vendor.
For earlier Claude models, the newer vocabulary produces roughly 18 percent fewer tokens for the same text, and the estimate is scaled by 1.18.
For Claude Opus 4.7 and later, the change goes the other way: the tokenizer used by those models produces roughly 30 percent more tokens for the same text than the earlier ones, so those are scaled by 1.53.
The documentation states the effect of the earlier factor as a measured improvement, reducing average estimation error from around 17 percent to around 2 percent.
Both are explicitly described as approximations that vary with content, which is the honest framing: a 2 percent average error still means a count that is wrong by hundreds of tokens on a large prompt, and that matters when you are deciding whether a repository fits in a context window.
So the count is a planning aid. Two flags exist for people who need more: one disables the correction entirely and reports the raw vocabulary count, the other switches to the vendor endpoint.
The exact path calls the vendor endpoint, and it is free
The alternative counting path sends the text to the vendor's counting endpoint with the same model flag used offline.
It needs a key, and three environment variables are accepted in a defined order, with the first non-empty one winning. The documentation notes the endpoint is free, which removes the usual objection to measuring text you were going to send anyway.
Two behaviours make this usable rather than annoying. Requests for multiple files are sent in parallel batches of four, which matters because a report that lists the largest files needs a count per file. And if the call fails for any reason, the tool falls back to the offline estimate rather than erroring out.
That fallback is worth noting when reading results. A line in the output can mean an exact count or a scaled estimate, and the tool prints which path it used, so the distinction is visible rather than silent.
The other end of the pipeline has the same convenience. A flag passes the generated prompt straight to any OpenAI compatible API for processing, using a suffix stored in your configuration file, or one supplied inline on the command line.
Tree-sitter compression keeps signatures and drops bodies
The compression feature is labelled experimental, and the reduction is aggressive.
What survives is the skeleton: package or module declarations, import statements, function and method signatures with their bodies removed, class definitions with method bodies removed, type definitions, and comments.
What goes is the implementation. For a JavaScript file that means the import line and the class outline remain while each method body becomes an elision marker.
Seven languages are listed as supported: Go, Python, JavaScript including arrow functions and modern module syntax, Bash, C and CSS.
The trade is clear and worth stating plainly. A prompt built this way tells a model what the interfaces are and what the file claims to do, but it cannot tell the model what the code does. That is the right shape for a question about architecture or for finding the right file, and the wrong shape for anything requiring the actual logic.
The repository root carries a development plan document for this feature, which is consistent with it being the newest and least settled part of the tool.
The recommended install tracks HEAD, so the curl path lags
Three installation routes are offered, and the recommended one installs from the module at the current head rather than a tagged release.
That matters for the release list. The most recent tags sit in a narrow range and two of them are weeks apart, so a binary installed through the module installer and one installed from a release archive are not necessarily the same code.
The second route pipes a shell script into an interpreter from a raw file URL. The documentation recommends against it, with the stated reason that it is harder to update, which is the honest description of what a curl pipe costs you.
The third is manual: take the archive from the releases page and move the binary into a directory on your path.
The two installation notes and the usage examples together make the flags predictable. Include a glob, add a diff, choose a model for the count, save into a per-project file under the home directory, or send the whole thing to a model with a prompt suffix.
The cross-compile target covers 64-bit Intel only
The build file is small and one line in it is worth knowing about.
The multi-platform target compiles for three combinations, all of them 64-bit Intel: Linux, macOS and Windows, each with the amd64 architecture suffix and a Windows executable extension.
There is no arm64 entry. So a machine on Apple silicon or on arm64 Linux is not covered by that target, and the practical route there is to install with the Go toolchain, which builds for the host.
The rest of the file is conventional. Version information comes from the git describe output and the build timestamp from the system clock, and both are stamped into the binary at link time with the symbol table and debug information stripped. There are targets to build, clean, test, fetch dependencies, run, install and uninstall into the Go workspace bin directory, and one that prints the version.
The lint target is the unusual one. It formats, runs a linter, and then runs a separate modernisation pass from the language tooling with the fix flag applied to source and tests, so a lint run here can rewrite files.
Editorial conclusion
This fits someone preparing a repository for a model in one command, since the tree view, the diff and the glob filter together are the part you would otherwise assemble by hand. It does not fit anyone who needs an exact context size, because the default count is an approximation with a documented error band, and the exact path needs an Anthropic key. Two things to check. The offline tokeniser section contradicts itself about whether anything is downloaded on first run, so test it rather than trusting the claim. And the cross-compile target covers 64-bit Intel machines only, so on an arm64 machine install from source with the Go toolchain instead.
Frequently asked questions
What does the ingest tool actually produce?
A single markdown file built from a directory of plain text files, including a tree view of the structure, with glob include and exclude filters. Output can go to the clipboard, to a named file, to a per-project file under the home directory, or to the console, with an optional JSON form.
How does ingest count tokens without an API key?
It counts offline with a compiled vocabulary. Because no Claude tokeniser has been published, Claude counts come from the newer vocabulary and are scaled: by 1.18 for earlier Claude models and by 1.53 for Opus 4.7 and later, which tokenises more densely.
Can ingest give an exact token count?
With the anthropic flag it calls the vendor's counting endpoint using the same model flag. That needs a key from one of three accepted environment variables, costs nothing, and falls back to the offline estimate if the call fails.
What does ingest keep when it compresses code?
Package and module declarations, imports, function and method signatures without bodies, class definitions without method bodies, type definitions and comments. The feature is marked experimental and covers Go, Python, JavaScript, Bash, C and CSS.
Can ingest send the result directly to a model?
Yes. The llm flag passes the generated prompt to any OpenAI compatible API, using a suffix from your configuration file, or one supplied inline with the prompt flag.
How is ingest installed?
Installing the Go module at head is the recommended route. A curl command piping a shell script is offered but described as harder to update, and the manual route is to move the binary from a release archive into a directory on your path.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/sammcj-ingest)