# Ante: A Self-Contained Coding Agent Harness Built in Rust

> Ante is a single compiled Rust binary that runs any coding agent task in a terminal, connecting to any API-hosted model or running fully offline through an embedded llama.cpp engine, with no runtime dependencies.

**AntigmaLabs/ante** — Ghost in your shell. Ante is a self-contained agent harness with a highly optimized core. It works like Claude Code or Codex, with none of their dependencies or model constraints. 

- Repository: https://github.com/AntigmaLabs/ante
- Website: https://antigma.ai
- Stars: 1,992 · Forks: 69
- Language: Rust
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/antigmalabs-ante

## What Ante Solves and Who It Is For

The gap Ante fills is this: Claude Code and Codex are model-coupled tools. You get the tool's chosen model, its API keys, its billing, and its runtime. Ante inverts that. It is a harness without a model subscription, a single compiled Rust binary that runs in a terminal and connects to whichever model its operator points it at. The README states the design goal plainly: an agent you can verify, afford, and run anywhere.

The practical targets are engineers who need one of three things: a coding agent that works in an air-gapped environment with no internet at all, a CI pipeline where a 15MB binary install is acceptable but a full Node.js runtime is not, or a local-first workflow where inference cost and data privacy make a commercial API unsuitable. The README also identifies a second audience: teams building a custom agentic harness. The examples/mini-tui directory in the repository shows a minimal TUI shell built on top of the same core, demonstrating how Ante can serve as the engine for a higher-level tool.

Ante does not treat the system prompt as proprietary. A settings profile can define the full agent behaviour, including a replacement system prompt, so the operator controls what the agent knows and how it operates. This positions it for organisations that need an agentic coding tool but cannot share their system prompts with a third-party vendor.

## Architecture: A Single Process with Embedded Tooling

Ante is hand-written Rust. Tools that most coding agents shell out to, such as grep and git queries, are compiled directly into the binary and run inside a single process. Local inference uses a pinned, managed build of llama.cpp that Ante controls internally; the user does not install llama.cpp separately or manage its version. The binary is approximately 15MB compressed and has no runtime dependencies beyond what the operating system provides.

The README documents a resource comparison across 20 parallel tasks run in Docker. Ante used approximately 7 times less peak memory, about 9 times less average CPU, and roughly 5 times less disk I/O than Claude Code in the same scenario. The raw comparison table and benchmark methodology are available at docs.antigma.ai.

Ante is also evaluated on Terminal-Bench 2.1, which covers 89 tasks with 5 trials each. The README reports a score of 83.9% with the open-weight DeepSeek V4.1 Flash model (370 out of 445 trials, Ante 0.preview.98, approximately $18 of inference). Each published result pins the exact Ante build and links the raw run for independent audit. This evaluation approach, running the harness against multiple model families rather than coupling it to one, reflects the stated design philosophy: the harness and the model should evolve together but not be bound together.

## Installing Ante and Running First Tasks

Installation is a single curl command that downloads the correct binary for the host platform:

```bash
curl -fsSL https://ante.run/install.sh | bash
```

The script places the binary on PATH. Two variants handle specific deployment needs. Passing `nightly` selects the nightly release channel, and setting `ANTE_INSTALL_DIR` places the binary in a directory already on PATH:

```bash
# Install a specific release channel
curl -fsSL https://ante.run/install.sh | bash -s -- nightly

# Install into a directory already on PATH
curl -fsSL https://ante.run/install.sh | ANTE_INSTALL_DIR=/usr/local/bin bash
```

macOS and Linux are supported. The README recommends WSL on Windows, which adds its own layer of operational setup.

Running `ante` with no arguments opens the interactive TUI. For CI and one-off tasks, the headless mode accepts a prompt as a command-line flag:

```bash
git diff | ante -p "review this for security issues"
```

The server mode (`ante serve`) exposes the agent over stdio, a Unix socket, or WebSocket so that editor plugins can connect to it. A gateway mode (`ante gateway`) runs Ante as a Slack or Discord bot when combined with the separate ante-gateway component. These four modes, interactive TUI, headless, server, and gateway, cover the common deployment shapes without requiring a different binary for each.

## Offline Operation with Embedded llama.cpp

Pointing Ante at a GGUF model file switches it to fully offline mode. The embedded llama.cpp handles inference inside the same process, with no network call required:

```bash
ante --offline-model ~/.ante/models/Qwen3.5-9B-Q4_K_M.gguf \
  -p "add error handling to src/main.rs"
```

The llama.cpp build is pinned and managed by Ante; there is no separate installation step. The README does not document which GGUF quantizations or context lengths the embedded engine supports in detail. That information is at docs.antigma.ai/local/offline rather than in the repository itself, which means the repository alone is not a complete reference for offline configuration.

A companion project called nanochat-rs, described in the README as a small GPT inference core written in pure Rust on candle, is published separately. The README is explicit that nanochat-rs is a study project and is not part of the main binary; it is published because in-process inference is where local models become interesting for agents. Engineers interested in how in-process inference works in Rust have that codebase as a reference, but it is not a supported component of Ante itself.

The offline capability is a meaningful differentiator from tools that require a live API connection for every call. Air-gapped machines, secure development environments, and workloads where inference cost needs to be predictable are all cases where the offline path is useful.

## Multi-Agent Coordination with /term

The /term command opens a named tmux session and launches a specified agent CLI inside it. Typing any of the following in the main Ante session creates a new named terminal:

```
/term ante
/term claude
/term codex
```

The main Ante instance gains a shared view of that terminal. It can delegate a task to the session, read the session's progress, and send follow-up prompts. Both the user and the main Ante instance can type in the session simultaneously. Detaching the viewer leaves the agent running; reconnecting Ante after a restart picks up the same live session.

The README calls these persistent interactive subagents. Each named terminal is an independent agent running its own context. Forking a conversation, to branch from a saved state, uses the target agent's native fork command rather than a feature in Ante itself. Ante's role in that operation is coordination and handoff, not the branch mechanism.

This feature requires tmux and whichever agent CLIs the user wants to run in those sessions. tmux is not a standard installation on most cloud instances or minimal container images, so this feature is effectively unavailable in those environments without additional setup.

## Trade-offs, Limitations, and Beta Constraints

The README marks the current release as a beta preview and states explicitly that breaking changes and incomplete functionality should be expected. That warning appears at the top of the README, not in a footnote. Windows is not supported at the binary level; the WSL path adds its own operational overhead.

The offline mode's detailed documentation lives at docs.antigma.ai rather than in the repository. The README does not specify what model sizes fit within a given hardware budget, or what the embedded llama.cpp version's context length limits are. Operators who want to run large models locally will need to consult external documentation to determine whether their hardware and model combination is workable.

The /term orchestration feature depends on tmux. A team deploying Ante in headless CI will not be able to use /term without adding tmux to the environment.

An operator who needs a first-party supported coding agent with a known, tested model should use Claude Code or Codex. Ante's value is resource efficiency, model freedom, and offline capability. It is not a supported product with a stable API between releases.

## Maintenance, Licensing, and the Upgrade Path

The last push to the repository was on 2026-09-17. Three releases shipped in the week ending 2026-09-28: v0.2.4 on 2026-09-23, v0.2.5 on 2026-09-25, and v0.2.6 on 2026-09-28. That pace indicates active development, though the beta warning means no release guarantees API stability.

The source code is Apache-2.0 licensed. The binary distribution is subject to an additional BINARY-TERMS.md file in the repository. The Apache-2.0 source licence does not automatically govern the distributed binary. Operators who plan to redistribute Ante or include it in a product should read BINARY-TERMS.md before doing so. The repository carries both files at the top level.

The upgrade path for the binary is a re-run of the install script. There is no version manager or lock file for the binary itself, so CI pipelines that need a reproducible build should pin the download to a specific release tag rather than pulling the latest each time.

## Conclusion

Ante suits engineers who need a coding agent without a model dependency: CI pipelines where install size matters, air-gapped machines, or cost-sensitive workloads where local inference is preferable to API calls. It is not a good fit for teams that need a stable, production-supported tool: the beta flag is genuine, and breaking changes between releases are expected. Before shipping Ante in a pipeline, verify the BINARY-TERMS.md obligations beyond the Apache-2.0 source licence and confirm that the embedded llama.cpp version supports your target model's quantization.

## FAQ

### Does Ante work on Windows?

The README states that macOS and Linux are the supported platforms at the binary level. On Windows, the README recommends running Ante through WSL, the Windows Subsystem for Linux.

### How do I run Ante without an internet connection?

Use the --offline-model flag with a path to a local GGUF model file. Ante's embedded llama.cpp engine handles inference entirely on the local machine with no API key or network connection required.

### How does Ante differ from Claude Code in model support?

Ante is model-agnostic: it works with any API provider or any local GGUF model. Claude Code is built around Anthropic's models. Ante also ships as a single compiled Rust binary with no runtime dependencies, whereas Claude Code requires a Node.js runtime.

## Sources

- [AntigmaLabs/ante on GitHub](https://github.com/AntigmaLabs/ante)
- [License: Apache-2.0](https://github.com/AntigmaLabs/ante/blob/main/LICENSE)
- [Project website](https://antigma.ai)
- [README](https://github.com/AntigmaLabs/ante/blob/main/README.md)
- [Releases](https://github.com/AntigmaLabs/ante/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/antigmalabs-ante
