Model or dataset
mattpocock/sandcastle avatar
mattpocock/sandcastle

Sandcastle: orchestrating coding agents in sandboxes with sandcastle.run()

Orchestrate sandboxed coding agents in TypeScript with sandcastle.run()

8,191 stars880 forksTypeScriptMIT

At a glance

What is it?
Sandcastle is a TypeScript library that runs AI coding agents inside Docker, Podman or Vercel sandboxes and merges their branch commits back. It suits teams already scripting agents in TypeScript, and it assumes you have a container runtime and a Claude token.
Who is it for?
Adopt Sandcastle if you already drive coding agents from TypeScript scripts or CI and want each run isolated in a container with its commits landing on a branch you can review. Skip it if you have no container runtime available, if you are not using a Claude-family agent, or if you expect the library to plan and review the agent's work for you: it starts the agent and hands back the commits.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 92 days ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

The problem Sandcastle solves: agents that write to your working tree

Running a coding agent directly in your checkout is awkward. The agent edits files, creates commits and sometimes leaves the tree in a state you have to untangle by hand. If you want two agents working at once, they collide on the same files. Sandcastle's answer is to give each agent its own sandbox and its own branch, then merge the resulting commits back. The README frames this in three steps: you invoke agents with a single sandcastle.run(), Sandcastle handles sandboxing with a configurable branch strategy, and the commits made on the branches get merged back. The audience is narrow and clear. This is a library for people who write TypeScript and want to fan out agent work programmatically, whether that is several AFK agents at once, a review pipeline, or their own agent orchestration. It is not a chat interface and it is not a hosted service. The repository ships a CLI entry point named sandcastle plus a programmatic run() export, so the same package covers the scaffold-and-run path and the embed-it-in-your-own-tooling path.

How sandcastle.run() splits an agent across sandboxes and branches

The mechanism is a provider abstraction. run() takes two required options: an agent and a sandbox. The agent comes from claudeCode(), which accepts a model string and an optional second argument for provider-specific settings such as effort level. The sandbox comes from a provider factory: docker(), podman(), vercel(), noSandbox(), or one you write yourself. Docker and Podman are bind-mount providers; Vercel is an isolated provider backed by Firecracker microVMs through @vercel/sandbox; noSandbox runs the agent directly on the host and skips container isolation entirely. All four are accepted by run(), createSandbox() and interactive(), and the worktree methods wt.run(), wt.interactive() and wt.createSandbox() accept the same set. One asymmetry is worth noting: wt.interactive() defaults to noSandbox() when no sandbox is specified, so the interactive worktree path is unisolated unless you say otherwise. The return value is a result object, not a stream of text. It carries iterations (with an optional sessionId per iteration), commits as an array of { sha } entries, and the target branch name. That shape tells you what the library considers its output: how many agent iterations ran, which commits were produced, and where they landed. Reviewing the diff is your job, not the library's.

Installing Sandcastle and running a first agent

Sandcastle needs Git and a sandbox provider. Docker Desktop is described as the most common local choice; Podman is the rootless alternative; Vercel covers the cloud case. Install the package as a dev dependency:

bash
npm install --save-dev @ai-hero/sandcastle

Then scaffold the project directory. The init command creates a .sandcastle directory with the files the library expects:

bash
npx @ai-hero/sandcastle init

Before the first run you need credentials. The README says to fill in CLAUDE_CODE_OAUTH_TOKEN in .sandcastle/.env, which you obtain by running claude setup-token on the host. If you would rather use an Anthropic API key, uncomment and fill in ANTHROPIC_API_KEY instead. The example file is copied into place first:

bash
cp .sandcastle/.env.example .sandcastle/.env

With that done, execute the scaffolded entry file. The README shows main.ts, and notes that main.mts also works:

bash
npx tsx .sandcastle/main.ts

The generated script is the real starting point. The README's example of the same call in library form looks like this:

typescript
import { run, claudeCode } from "@ai-hero/sandcastle";
import { docker } from "@ai-hero/sandcastle/sandboxes/docker";

await run({
  agent: claudeCode("claude-opus-4-8"),
  sandbox: docker(),
  promptFile: ".sandcastle/prompt.md",
});

What you should see is a run that returns a result object: the number of iterations, the commits created, and the branch they were made on. The prompt itself lives in .sandcastle/prompt.md, which is why the call points at a file rather than an inline string. Inline prompts are also accepted, as the provider examples in the README show with prompt: "...".

Mounts, UID mismatches and the limits of the sandbox model

The docker() factory takes configuration that matters more than it first appears. imageName lets you point at your own image. containerUid and containerGid override the UID and GID passed to the container's --user flag, which default to the host values; the README states these must match the UID baked into the image, and that a pre-flight check catches mismatches. That check is a convenience, not a fix: if your image was built with a different user, you will be adjusting one side or the other before anything runs. Mounts let you bring host directories into the sandbox, which is the documented way to reuse package manager caches such as ~/.npm. hostPath accepts absolute, tilde-expanded and relative paths resolved from the current working directory; sandboxPath accepts absolute and relative paths resolved from the sandbox repo directory. There is also a SELinux volume label option taking "z" (the default, shared), "Z" (private) or false, which the README says is a no-op on non-SELinux systems. The larger limitation is categorical. Sandcastle isolates where the agent runs, not what it decides to do. A bind-mount provider still shares directories with the host, so a mount that includes something sensitive is a mount that includes something sensitive. And the no-sandbox provider exists precisely because container isolation is sometimes unwanted, which means the library will happily run an agent with no isolation at all if you ask it to. Nothing in the result object tells you whether the commits are good.

Sandcastle compared with running agents from a CI job

The obvious alternative is a plain CI pipeline that checks out the repository, installs the agent CLI, and runs it against the working tree. That approach gives you logs, artifacts and scheduling for free, and it needs no new abstraction. The difference is what happens to the agent's output. In a CI job the agent edits the checkout in place; you get a diff and whatever the job decides to do with it. Sandcastle introduces a branch strategy and a sandbox provider as first-class concepts, then reports commits and the target branch in its return value, so the merge-back step is part of the model rather than something you assemble from git commands in a shell script. The trade is that you now depend on a container runtime being present wherever the script runs, and on the provider abstraction being one the library ships or one you write. The README documents two factory functions for custom providers, createBindMountSandboxProvider and createIsolatedSandboxProvider, which is the escape hatch if your environment is neither Docker, Podman nor Vercel. A second alternative is simply running the agent CLI by hand in a scratch clone. That is cheaper for one-off work and stops being practical the moment you want several agents running at the same time.

Maintenance, licence and what upgrading costs

The repository is not archived. Its last push was on 2026-06-29, the same day v0.12.0 was released, which is the most recent of the three releases listed here (v0.10.0 on 2026-06-18 and v0.9.0 on 2026-06-16). The version numbers move in the 0.x range, and the presence of a .changeset directory in the repository root indicates changesets is used to manage releases; the package.json release script runs changeset publish. In practice that means minor versions can carry breaking changes, and the changelog in the repository is where you check before bumping. The package is published as @ai-hero/sandcastle, with a bin entry named sandcastle, and the exports map currently exposes the root plus ./sandboxes/docker, ./sandboxes/vercel, ./sandboxes/podman, ./sandboxes/daytona and ./sandboxes/no-sandbox. Note that daytona appears in the exports map but not in the README's provider table, so treat it as present in the build and undocumented in the prose. The licence is MIT, which permits commercial use and modification provided the copyright notice and permission notice are retained; that is a summary of the licence terms, not legal advice, and you should read the LICENSE file if the distinction matters to your organisation. Upgrade cost is mostly the provider factories: docker() and friends take their configuration inline, so a change to accepted options shows up at the call site rather than in a config file.

Editorial conclusion

Adopt Sandcastle if you already drive coding agents from TypeScript scripts or CI and want each run isolated in a container with its commits landing on a branch you can review. Skip it if you have no container runtime available, if you are not using a Claude-family agent, or if you expect the library to plan and review the agent's work for you: it starts the agent and hands back the commits. Before committing to it, run npx @ai-hero/sandcastle init, confirm the generated .sandcastle/main.ts executes under npx tsx, and check that the UID and GID baked into your sandbox image match the host values the pre-flight check compares against.

Frequently asked questions

How do I use Sandcastle to run a coding agent?

Install @ai-hero/sandcastle as a dev dependency, run npx @ai-hero/sandcastle init to scaffold a .sandcastle directory, fill in CLAUDE_CODE_OAUTH_TOKEN or ANTHROPIC_API_KEY in .sandcastle/.env, then execute the generated entry file with npx tsx .sandcastle/main.ts. The script calls run() with an agent such as claudeCode("claude-opus-4-8") and a sandbox provider such as docker().

What is Sandcastle by mattpocock?

It is a TypeScript library, published as @ai-hero/sandcastle, for orchestrating AI coding agents in isolated sandboxes. It ships built-in providers for Docker, Podman and Vercel, and the README describes it as provider-agnostic with support for custom providers.

Which sandbox providers does Sandcastle support?

Docker, Podman, Vercel and a no-sandbox option are built in, imported from paths such as @ai-hero/sandcastle/sandboxes/docker. Docker and Podman are bind-mount providers, Vercel is an isolated provider using Firecracker microVMs via @vercel/sandbox, and noSandbox runs the agent directly on the host.

What does sandcastle.run() return?

It returns a result object containing iterations, an array of per-iteration results with an optional sessionId, commits as an array of { sha } entries, and the target branch name. The number of commits and the branch are what you use to review the agent's work afterwards.

Does Sandcastle require Docker?

No. Docker Desktop is described as the most common option for local development, but Podman and Vercel are also built in, and the README documents createBindMountSandboxProvider and createIsolatedSandboxProvider for writing your own provider. Git is listed as a prerequisite alongside a sandbox provider.

Official sources

  1. Issues
  2. License: MIT
  3. mattpocock/sandcastle on GitHub
  4. README
  5. Releases
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/mattpocock-sandcastle.svg)](https://hysenlabs.com/projects/mattpocock-sandcastle)
Community notes

Community notes