Model or dataset
Th0rgal/sandboxed.sh avatar
Th0rgal/sandboxed.sh

sandboxed.sh: a self-hosted mission backend for autonomous coding agents

Safe runtime for autonomous on-chain AI agents: isolated sandboxes, Library skills, encrypted secrets.

513 stars53 forksRustLicense varies

At a glance

What is it?
sandboxed.sh runs Claude Code, OpenCode, Codex, Gemini and Grok agents inside isolated Linux workspaces, and exposes them over MCP to a coordinator that decides what to build. The split is deliberate, and it is also the main thing to weigh before adopting it.
Who is it for?
Adopt sandboxed.sh if you already have an MCP-capable coordinator and you want agent work confined to throwaway Linux workspaces with a versioned Library and encrypted secrets. Do not adopt it if you want a single binary that both decides and executes, or if your host cannot give the container privileged mode and host cgroups for systemd-nspawn.
Can I use it commercially?
Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Rust, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem: agents that edit your real machine

Most coding-agent setups give the model a shell on the machine you care about. The agent reads your SSH keys, writes into your working tree, and installs whatever it decides it needs. sandboxed.sh takes the opposite position: the agent never runs untrusted code itself. Each unit of work is handed to a mission runner, which executes it in a throwaway workspace and streams structured results back.

The README frames this as a two-part system. sandboxed.sh is the mission-execution backend, the half that actually builds things in isolation. The other half is a coordinator that decides what to do and when. The project's own coordinator is a fork of Hermes, plus a bundled Electron desktop app, but the README states that any MCP-capable assistant works. That is the intended audience: people who already have something that can call MCP tools and want the execution side to be contained.

The vision section lists three scenarios: handing off a GitHub issue end to end, running multi-day operations unattended, and keeping sensitive data local. The second and third are the ones that justify the container boundary. If you only ever run one agent against one repository on a laptop, the isolation buys you less than the operational overhead costs.

Missions, projects and the control conversation

The architecture separates durable state from execution. A project is the durable unit of work, stored in `projects.db` on the sandboxed.sh host and served at `/api/projects/*`. It carries a mode (`active`, `blocked`, `paused`), an autonomy grant covering merge authority, budget and parallelism, plus tracks and open decisions.

A mission is one unit of autonomous execution: an agent in an isolated workspace or container running a harness such as Claude Code or Codex, tagged with a project and track. Controllers are coordinator-side cron jobs that wake on a schedule, read their control conversation, GitHub and project state, then dispatch missions. The coordinator talks to sandboxed.sh over MCP and receives results over SSE or webhooks.

The detail worth noticing is that controllers are supposed to write structured project state through MCP tools such as `list_projects`, `update_project_status`, `set_project_grant` and `link_mission_to_project`, rather than prose. The README also describes a state ingestor that folds controller status trailers from deliveries into the project record, so text-only updates still keep the roster current. That fallback exists because structured-only reporting is hard to enforce in practice, and the project acknowledges it by building the looser path anyway.

The target model is documented as portfolio, project, track, attempt, action, receipt, evidence in `docs/AGENT_CONTROL_PLANE.md`. The README's rule of thumb is blunt: in-conversation subagents are for quick reasoning and decomposition; anything needing a real filesystem, git, builds or a PR gets dispatched as a mission.

Installing sandboxed.sh with Docker Compose

The repository ships a `docker-compose.yml` that builds the image locally and maps container port 80 to host port 3000. Two named volumes persist state: `sandboxed-data` at `/root/.sandboxed-sh` and `claude-auth` at `/root/.claude`. An SSH key mount for private git repositories is present but commented out.

bash
docker compose up

After the build, the dashboard is reachable on port 3000. The compose file reads `.env` if present, and marks it not required, so the service starts with defaults. Copy `.env.example` to `.env` before you rely on anything configurable.

bash
cp .env.example .env

The defaults that matter most are `PORT=3000`, `MAX_ITERATIONS=50`, `STALE_MISSION_HOURS=24` and `MAX_PARALLEL_MISSIONS=1`. That last one is the throttle: out of the box, one mission runs at a time. If you expect parallel work, that is the value to change, and it is also where resource contention will show up first.

Container workspace support through systemd-nspawn is off by default. The compose file has the two lines that enable it, `privileged: true` and `cgroup: host`, commented out with a note. The native install path is documented separately in `docs/install-native.md`, and the README points to `docs/install-docker.md` for the container route. The `LIBRARY_PATH` default is `/root/.sandboxed-sh/library`, and `LIBRARY_REMOTE` falls back to an official template repository when no settings file exists.

The Library and encrypted secrets

Skills, tools, rules, agents and MCPs live in a single git-backed repository called the Library. That choice has a real consequence: the configuration your agents run with is versioned, diffable and reviewable the same way code is. You can point `LIBRARY_REMOTE` at your own repository, and the README notes the remote is preferably configured through the dashboard Settings page, with the environment variable acting as the initial default when no settings file exists.

Secrets are handled in the Rust backend with `aes-gcm`, `pbkdf2`, `sha2` and `hmac`, and the project description calls them encrypted secrets. Authentication uses `jsonwebtoken`. What the README does not document is a key-rotation procedure or a recovery path if the encryption key material is lost. Before you put production credentials into the secrets store, that gap is worth resolving against the source rather than assuming.

The Library also interacts with strong skill isolation. `.env.example` mentions `OPENCODE_CONFIG_DIR` as optional and points to `INSTALL.md` for the strong-isolation mode. If you run agents with different trust levels, read that section rather than leaving the default.

Where sandboxed.sh is the wrong tool

The coordinator split is the biggest constraint. sandboxed.sh does not decide what to do. If you want one process that reads an issue, plans, edits files and opens a PR, you are running half a system and you will need to supply the other half. The README is explicit that the agent never runs untrusted code itself and that a separate coordinator decides what and when, so this is a design boundary, not a missing feature.

The isolation story also depends on the host. systemd-nspawn workspace support requires `privileged: true` and `cgroup: host` in the compose file, which the file itself leaves commented out. On a shared host, or any environment where you cannot grant those, you are limited to the weaker default. That is a genuine deployment constraint, not a tuning knob.

Resource limits are conservative by default. `MAX_PARALLEL_MISSIONS=1` and `MAX_ITERATIONS=50` mean a single mission runs at a time and stops after fifty iterations. Multi-day unattended operations, which the vision section describes, will not look like that until you raise both. And because the whole thing is self-hosted, the operational surface is yours: a Rust service, a Next.js dashboard, a SQLite database in `projects.db`, and whichever AI CLIs the image installs.

Finally, the licence status is not stated in the repository metadata available. No licence identifier is listed, so anyone evaluating it for commercial use should check the repository directly before depending on it.

How it differs from running an agent inside a container yourself

The obvious alternative is a plain container: build an image with the agent CLI, mount a working directory, run it. That gets you filesystem isolation and nothing else. You still own the scheduling, the result capture, the state record of what ran and what it produced, and the mapping from a unit of work back to a project.

sandboxed.sh supplies those layers. Missions are first-class objects tagged with project and track. Results stream back over SSE or webhooks. Project state lives in a queryable database rather than in your head or in a chat log. The Library makes agent configuration a git repository. The MCP registry is optional and only needed when a mission requires extra tool servers such as desktop or playwright.

The trade-off is coupling. A hand-rolled container plus a shell script has no schema, no mode field, no autonomy grant, and no controller protocol to learn. sandboxed.sh has all four, and the README's own comparison table shows how much of the system's behaviour is defined by the coordinator side. If your workflow is one agent, one repository, one command, the container is less machinery for the same isolation. If your workflow is many missions across several projects with a need to know what ran, sandboxed.sh is solving a problem the container does not address.

Maintenance, upgrades and what to check first

The last push to the default branch was on 2026-09-10, and the most recent release listed is v0.12.0 from 2026-05-16, with v0.11.5 and v0.11.1 before it in May. The gap between the latest tagged release and the last commit is roughly four months, which suggests development continues on `master` between tags. The repository is not archived.

Upgrade cost is shaped by the build. The Dockerfile is a multi-stage build: a Rust stage pinned to `rust:1.91-bookworm`, a dashboard stage on `oven/bun:1`, and a runtime stage carrying the AI CLIs. The Rust stage runs `cargo generate-lockfile 2>/dev/null || true` and a stub-source build with errors suppressed before copying real sources. Those suppressions mean a broken dependency resolution can pass silently at that step and surface later. If you build from source, watch that stage rather than trusting the exit code.

The crate version in `Cargo.toml` is 1.3.0 while the release tags are at v0.12.0, so version numbers in the repository do not track the release tags. Pin to a tag or a commit rather than to a version string when you deploy.

On licensing, no licence is stated for this repository. That is not a legal opinion, just a fact you should resolve before shipping it inside a company. Check the repository for a licence file and read it, because the answer determines whether the encrypted-secrets and Library features can be used in a commercial deployment at all.

Editorial conclusion

Adopt sandboxed.sh if you already have an MCP-capable coordinator and you want agent work confined to throwaway Linux workspaces with a versioned Library and encrypted secrets. Do not adopt it if you want a single binary that both decides and executes, or if your host cannot give the container privileged mode and host cgroups for systemd-nspawn. Before committing, verify the licence terms in the repository, confirm whether the build is reproducible given the `2>/dev/null || true` steps in the Dockerfile, and check whether the coordinator half it expects is something you are willing to run.

Frequently asked questions

What is sandboxed.sh used for?

It is a self-hosted mission-execution backend for autonomous AI agents. A coordinator drives it over MCP, and it runs each unit of work as a mission inside an isolated Linux workspace using a harness such as Claude Code, OpenCode, Codex, Gemini or Grok.

How do I install sandboxed.sh?

The repository ships a docker-compose.yml that builds the image and maps container port 80 to host port 3000, with the command `docker compose up`. The README also links a Docker guide at docs/install-docker.md and a native guide at docs/install-native.md.

Does sandboxed.sh decide what the agent should work on?

No. The README describes it as the mission-execution backend of a two-part system, where a separate coordinator decides what to do and when. The project's own coordinator is a Hermes fork, but the README states that any MCP-capable assistant works.

What is the Library in sandboxed.sh?

It is a git-backed repository holding skills, tools, rules, agents and MCPs. It is configured through the dashboard Settings page, with the LIBRARY_REMOTE environment variable acting as the initial default when no settings file exists.

Does sandboxed.sh need privileged mode to run?

Not for the default setup, but container workspace support via systemd-nspawn requires the `privileged: true` and `cgroup: host` lines in docker-compose.yml, which ship commented out. Without them you are limited to the weaker default isolation.

Official sources

  1. Issues
  2. Project website
  3. README
  4. Releases
  5. Th0rgal/sandboxed.sh on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/th0rgal-sandboxed-sh.svg)](https://hysenlabs.com/projects/th0rgal-sandboxed-sh)