rivet-dev/sandbox-agent: run coding agents in a sandbox and drive them over HTTP
Run Coding Agents in Sandboxes. Control Them Over HTTP. Supports Claude Code, Codex, OpenCode, and Amp.
At a glance
- What is it?
- sandbox-agent puts a Rust server inside the sandbox and exposes one HTTP plus SSE API over Claude Code, Codex, OpenCode, Cursor, Amp and Pi. Good for platform teams that want to swap agents without rewriting their integration; weaker where you need a stable surface, since the newest published line is a release candidate.
- Who is it for?
- Adopt sandbox-agent if you are building a platform that runs coding agents on behalf of other people and you want one HTTP surface plus a normalized event stream instead of one integration per agent. Do not adopt it if you need a pinned, long-supported API today, or if your agents already run comfortably on the developer's own machine.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 103 days ago.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem sandbox-agent solves, and who ends up using it
Three assumptions break when coding agents move off a laptop. The first is that the agent process and the code it edits share a machine. The README states that coding agents need isolated environments, and that existing SDKs assume local execution. The second is that one agent is enough. Claude Code, Codex, OpenCode, Cursor, Amp and Pi each carry their own API, event format and behaviour, so swapping one for another means rewriting the integration layer. The third is that a session is durable. Transcripts live inside the sandbox, and when the process ends they are gone.
sandbox-agent is aimed at the people who feel all three at once: teams building a hosted coding product, an internal developer platform, or a CI job that delegates edits to an agent. If you run one agent on your own machine and read its output in a terminal, this project adds a network hop and a schema you did not ask for. The audience is narrower than the feature list suggests.
One adapter per agent, behind a single HTTP and SSE surface
The architecture is an adapter pattern with a process boundary in the middle. The README describes the server as a universal adapter between your client application and the coding agents; each agent gets its own adapter that translates between the universal API and the agent-specific interface. The server itself is a Rust daemon, and the CLI is the same binary plus an npm wrapper, so the HTTP endpoints and the command line mirror each other rather than diverging.
Two modes exist. Embedded mode runs agents locally as subprocesses, which is what the TypeScript SDK uses when you call SandboxAgent.start(). Server mode runs the daemon as an HTTP server from any sandbox provider, and the client connects to a baseUrl with a token. The Rust workspace in Cargo.toml shows the split in code: separate crates for agent management, agent credentials, an opencode adapter, an opencode server manager, and an acp-http-adapter. The acp-http-adapter crate is the interesting one, because it suggests the universal surface is built on the Agent Client Protocol rather than a bespoke format. That is a design choice worth checking against the API reference if you care about interop beyond this project's own SDK.
The universal session schema normalizes every agent's event format for storage and replay. The README names Postgres, ClickHouse and Rivet as targets, and the examples directory contains persist-postgres and persist-sqlite, so persistence is demonstrated rather than only described.
Installing the server and creating a first session
The README gives three installation paths. The server path is a curl install script, run inside the sandbox. The README shows the 0.4.x line:
curl -fsSL https://releases.rivet.dev/sandbox-agent/0.4.x/install.sh | shAfter that, start the daemon with a token, a host and a port. The README uses port 2468 throughout, and the token is read from an environment variable:
sandbox-agent server --token "$SANDBOX_TOKEN" --host 127.0.0.1 --port 2468The README notes that agent binaries are installed lazily on first use, and that you can preinstall them instead so the first request is not delayed:
sandbox-agent install-agent --allFor local work there is an explicit escape hatch. The README documents disabling auth with --no-token, which is fine on a loopback interface and a bad idea anywhere else.
The SDK path is npm. The README pins the 0.4.x range, and Bun users get an extra step because the package ships native binaries:
npm install [email protected]Once installed, the SDK example connects to a running server and opens a session. Note the agent name and the agentMode field:
import { SandboxAgent } from "sandbox-agent";
const client = await SandboxAgent.connect({
baseUrl: "http://127.0.0.1:2468",
token: process.env.SANDBOX_TOKEN,
});
await client.createSession("demo", { agent: "codex", agentMode: "default" });
await client.postMessage("demo", { message: "Hello from the SDK." });
for await (const event of client.streamEvents("demo", { offset: 0 })) {
console.log(event.type, event.data);
}The offset: 0 argument is the part to notice. streamEvents takes a starting offset, which is what makes replay after a reconnect possible rather than a fresh subscription. If you are wiring this into a queue or a webhook consumer, that offset is the piece you persist alongside the session id.
The CLI wrapper offers the same operations from a shell, which is useful for smoke-testing without writing TypeScript:
npm install -g @sandbox-agent/[email protected]
sandbox-agent api sessions create my-session --agent codex --endpoint http://127.0.0.1:2468 --token "$SANDBOX_TOKEN"
sandbox-agent api sessions send-message-stream my-session --message "Hello" --endpoint http://127.0.0.1:2468 --token "$SANDBOX_TOKEN"Finally, the built-in Inspector serves a UI on the same port. The README gives http://localhost:2468/ui/ as the example, which is the fastest way to confirm that events are arriving before you write any client code. The justfile also exposes SANDBOX_AGENT_SKIP_INSPECTOR=1 for builds that should not ship the UI.
Where the model strains: agent coverage, versioning and the auth boundary
The agent list is the first thing to check against reality. The repository description names Claude Code, Codex, OpenCode and Amp. The README's opening paragraph adds Cursor and Pi, and the Cargo.toml workspace description names only Claude Code, Codex, OpenCode and Amp. Those three statements do not agree. Coverage is likely per-adapter and uneven, and the README does not publish a capability matrix showing which features work for which agent. If your workflow depends on a specific agent's permission model or tool-calling behaviour, verify it against the API reference before you build on it.
The second strain is versioning. The most recent releases listed are v0.5.0-rc.3 from 2026-03-30, v0.4.2 from 2026-03-26 and v0.5.0-rc.2 from 2026-03-25. The README, the SDK install commands and the curl install URL all point at 0.4.x, while the newest published tag is a release candidate. So the documented, stable-looking path and the newest code are not the same thing. The README does not document a rollback procedure, and it does not state a support window for the 0.4.x line. A platform team pinning this dependency has to decide that for itself.
The third strain is the security boundary. The token is a shared secret passed on the command line, and the daemon has to reach the agent binaries it manages. Running with --no-token removes the only documented authentication control, so the host and port flags become the entire boundary. The README does not describe per-session authorization or scoping, which means a token that can create a session in one place can likely create one anywhere the server listens. That is a normal shape for a single-tenant sandbox, and a problem for a multi-tenant one. The README does not address multi-tenancy directly, so treat that as unverified.
How it compares with a plain Docker setup or the sandbox vendors' own SDKs
The obvious alternative is doing it yourself: bake the agent into a container image, run it with docker run, and talk to it over SSH or a mounted socket. That works, and it is what most teams start with. The README argues the failure mode directly, saying SSH breaks TTY handling and streaming. Whether that bites you depends on your client. A batch job that runs an agent once and reads a log file will not care. An interactive product that streams partial output to a browser will care a great deal, and that is where a structured SSE stream with an offset beats a pseudo-terminal.
The second alternative is the sandbox provider's own SDK. The examples directory includes directories for e2b, daytona, modal, cloudflare, vercel, sprites, boxlite and agentcomputer, which tells you the intended relationship: sandbox-agent runs inside those environments rather than replacing them. The difference in approach is layer, not function. A provider SDK gives you the machine and its lifecycle; sandbox-agent gives you the agent protocol on top of it. Choosing between them is a category error, but choosing whether to add the second layer is not. If you only ever run one agent and you are happy parsing its native output, the provider SDK plus a shell command is less machinery.
The third alternative is embedding the agent's own SDK in your application process. That is the fastest route to a prototype and the worst route to isolation, because the agent then executes inside your application's environment. The README's first stated problem is exactly this: you cannot let AI execute arbitrary code on your production servers. If your application already runs in a throwaway container, the argument weakens; if it runs next to your database, it does not.
Licence, maintenance and the cost of keeping up
The licence is Apache-2.0, declared in both the repository metadata and the Cargo.toml workspace package section. That is a permissive licence with an explicit patent grant, and it imposes no copyleft obligation on the code you write around the client. The repository does not ship a separate licence file for the agent binaries the server installs, and the README does not discuss how those agents are licensed. If you redistribute a sandbox image with agents preinstalled via sandbox-agent install-agent --all, that is the question to take to your own counsel, because this project's licence does not answer it.
Maintenance signals are mixed. The repository is not archived, and the last push was on 2026-06-19. The release history shows a 0.4.2 stable tag and two 0.5.0 release candidates within five days of each other in late March 2026. The workspace is a pnpm and Turbo monorepo with a Rust workspace, Biome, Vitest, lefthook and a justfile of release targets, so the build surface is broad. Upgrading means moving the Rust daemon, the TypeScript SDK, the CLI wrapper and the agent adapters together, and the npm packages ship platform-specific native binaries through optional dependencies. The README shows that Bun users must run bun pm trust for five separate @sandbox-agent/cli-* packages, which is a small but real friction point on every fresh install. Budget for the upgrade as a coordinated change, not a version bump.
Editorial conclusion
Adopt sandbox-agent if you are building a platform that runs coding agents on behalf of other people and you want one HTTP surface plus a normalized event stream instead of one integration per agent. Do not adopt it if you need a pinned, long-supported API today, or if your agents already run comfortably on the developer's own machine. Before committing, verify two things yourself: which agent adapters your workflow actually needs, and whether the universal session schema carries the fields you plan to persist. The repository's own README still labels Gigacode and OpenCode SDK and UI support as experimental, so treat those as the least settled parts.
Frequently asked questions
What does sandbox-agent do?
It is a server that runs inside a sandbox and exposes an HTTP plus SSE API for controlling coding agents such as Claude Code, Codex, OpenCode, Cursor, Amp and Pi. A client application connects remotely to stream events, handle permissions and manage sessions through one interface instead of one integration per agent.
What is sandbox-agent?
The README describes it as a server that runs inside your sandbox, with each coding agent getting its own adapter that translates between a universal API and the agent-specific interface. It ships as a Rust daemon, a TypeScript SDK, a CLI and a built-in Inspector UI.
How is sandbox-agent different from just running Docker?
Docker gives you the isolated machine; sandbox-agent runs inside it and gives you a structured agent protocol on top. The README argues that SSH breaks TTY handling and streaming, which is the gap the HTTP and SSE surface is meant to close.
Which version of sandbox-agent should I install?
The README's install commands for the SDK, the CLI and the curl script all point at the 0.4.x line, while the newest published tag is v0.5.0-rc.3. The README does not state a support window for 0.4.x or a rollback procedure, so that choice is yours to make.
Does sandbox-agent support MCP?
The examples directory contains mcp and mcp-custom-tool directories, and the repository carries an .mcp.json file at the top level. The README itself does not document an MCP feature, so treat the examples as the reference point.
Can I run sandbox-agent without authentication?
Yes. The README documents sandbox-agent server --no-token --host 127.0.0.1 --port 2468 for local use. The README does not describe any other authorization control beyond the token, so disabling it leaves host and port as the only boundary.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/rivet-dev-sandbox-agent)