# Distill: A Token-Efficient Coding Agent Harness with a Three-Tier Model System

> Distill is a Rust-built coding agent harness and terminal UI that reduces token costs by routing bounded tasks like extraction and compression to a cheaper utility model, leaving the main model for reasoning. It works with Grok, Codex, and any OpenAI-compatible provider.

**samuelfaj/distill** — Distill large CLI outputs into small answers for LLMs and save tokens!

- Repository: https://github.com/samuelfaj/distill
- Stars: 691 · Forks: 44
- Language: TypeScript
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/samuelfaj-distill

## The Token Problem Distill Addresses

Coding agents consume tokens quickly. Every tool result, every file read, every shell command output adds to the context. The accumulated context drives up costs and eventually hits model context limits. Distill's description names this directly: it distills large CLI outputs into small answers for LLMs and saves tokens.

The approach is architectural. Rather than sending all output verbatim to the main model, Distill routes extraction, summarization, and compression tasks to a dedicated utility model. That model is designed for bounded tasks where the output is a processed version of the input, not a reasoning step. The main model receives the compressed result instead of the raw output, keeping its context smaller and the cost lower. For long-running sessions that accumulate many tool results, this difference compounds over the course of a session.

## Three-Tier Model Architecture

Distill defines three model tiers with distinct roles. The main model runs every session and every step and is the only required tier. The reasoning model is optional and handles planning and review for steps the main model cannot do alone. The utility model handles bounded tasks: extraction, summaries, and compression.

Each tier is configured separately. The main model picker also accepts an effort setting, with `auto` as the default. The README shows example commands for setting each tier:

```text
/model gpt-6-luna auto
/reasoning-model gpt-6-sol
/utility-model openrouter-qwen37 auto
```

All three tiers accept any model that speaks the OpenAI-compatible protocol. Tier configuration is accessible via `/tiers` on the home screen, or by using `/tiers main`, `/tiers reasoning`, or `/tiers utility` directly. The design separates model selection from task routing: you choose which models fill each role, and Distill decides which tier handles each operation.

## Installing and Updating Distill

On Mac and Linux, the installation uses a shell script:

```sh
curl -fsSL https://raw.githubusercontent.com/samuelfaj/distill/main/install.sh | sh
export PATH="$HOME/.local/share/distill/bin:$PATH"
distill --version
```

On Windows, the equivalent is a PowerShell command:

```sh
irm https://raw.githubusercontent.com/samuelfaj/distill/main/install.ps1 | iex
$bin = "$env:LOCALAPPDATA\distill\bin"; [Environment]::SetEnvironmentVariable('Path', "$bin;" + [Environment]::GetEnvironmentVariable('Path', 'User'), 'User')
distill --version
```

Updating an installation made with the release installer uses a single command:

```sh
distill update
```

First-time setup requires two login steps after installation: login with OpenRouter and login with a subscription (Grok or Codex, or an OpenAI-compatible provider). The README describes these as the second and third steps before getting to token savings.

## Provider Compatibility and Subscription Model

Distill works with Grok and Codex subscriptions by name, and accepts any model that speaks the OpenAI-compatible protocol through OpenRouter or a direct endpoint. This means models hosted locally or on alternative providers can fill any of the three tiers, as long as they accept the same API shape.

The token savings mechanism depends on the utility model being cheaper per token than the main model. If all three tiers use the same model, the three-tier architecture adds orchestration overhead without cost benefit. The README's step list says "Save money" as the last step after logging in, suggesting the assumed setup has a cheaper utility model doing the compression work.

Documentation in the repository's `docs/` directory covers additional detail: `docs/local-models.md` for local model configuration, `docs/jev-routing.md` for how Jev routes work, and `docs/token-saver.md` for the mechanics of where the token savings come from.

## Implementation: Rust Under the Hood

The repository's primary language is listed as TypeScript, but the Cargo.toml at the root reveals that Distill is a Rust workspace. The workspace contains dozens of crates covering the agent lifecycle, chat state, codebase graph, context compaction, MCP integration, hooks, authentication, and the TUI itself. The Cargo.toml patches `async-openai` with a fork at a specific revision, which suggests Distill depends on unreleased behavior in the OpenAI client library.

The `rust-toolchain.toml` pins the Rust toolchain version. Builds use `build.sh`. The workspace structure is auto-generated according to the comment in Cargo.toml; individual crates have their own Cargo.toml files.

The presence of `distill-compaction-transcript`, `distill-codebase-graph`, and `distill-fast-worktree` crates aligns with the token-saving architecture: the compaction crate likely handles the compression pass, the codebase graph crate manages the code understanding layer, and the fast worktree crate supports git-based operations.

## Comparison with Aider

Aider is an open-source coding agent that also runs in the terminal and works with various LLM providers through an OpenAI-compatible interface. Both tools handle code-aware agent sessions in the CLI. Aider does not have a dedicated utility model tier for context compression. It sends full context to a single configured model and relies on that model's context window. Distill's utility model tier is the architectural difference: it offloads compression and extraction to a cheaper model before those results reach the main model's context. Teams already using Aider and satisfied with its context management would not find Distill's primary argument compelling unless token costs are a specific concern.

## License and Maintenance Status

Distill is Apache-2.0 licensed. A NOTICE and THIRD-PARTY-NOTICES file are present in the repository, which is expected given the forked async-openai dependency. A SECURITY.md documents how to report vulnerabilities. A CONTRIBUTING.md is also present.

The project is actively maintained. The last push was on 2026-09-22, and three releases were published in quick succession in the days just before: v2.0.15 on 2026-09-25, v2.0.16 on 2026-09-26, and v2.0.17 on 2026-09-27. That release cadence suggests active ongoing development. The version scheme (v2.0.x) with patch-level bumps indicates stability work rather than new features, though the README does not include release notes for individual versions. The `SOURCE_REV` file at the repository root tracks the source revision separately from the Cargo.lock and may relate to reproducible build tracking.

## Conclusion

Distill is worth evaluating for developers who use Grok or Codex subscriptions and want to reduce per-session token costs through context compression. The three-tier model architecture is the mechanism to verify first: does the utility model compress accurately enough that the main model gets sufficient context, or does the compression lose details that matter for the agent's next step? The docs at `docs/token-saver.md` and `docs/jev-routing.md` in the repository explain both mechanisms in detail. The install scripts are straightforward, and `distill --version` confirms a working binary before committing to a configuration. Distill is actively maintained, with v2.0.17 released on 2026-09-27, two days before this review.

## FAQ

### What does Distill do differently from a standard coding agent?

Distill routes extraction, summarization, and compression tasks to a separate utility model tier instead of sending all CLI output directly to the main model. This keeps the main model's context smaller and reduces per-session token costs.

### Which models does Distill support?

Distill works with Grok and Codex subscriptions and with any model that speaks the OpenAI-compatible protocol. Each of the three tiers (main, reasoning, utility) can be set independently via /model, /reasoning-model, and /utility-model commands.

### How does Distill save tokens?

The utility model tier handles bounded tasks like extraction and compression, producing a smaller result that the main model receives instead of the raw output. The repository's docs/token-saver.md documents the specific mechanics.

## Sources

- [Issues](https://github.com/samuelfaj/distill/issues)
- [README](https://github.com/samuelfaj/distill/blob/main/README.md)
- [Releases](https://github.com/samuelfaj/distill/releases)
- [samuelfaj/distill on GitHub](https://github.com/samuelfaj/distill)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/samuelfaj-distill
