agentsh: a policy-enforced shell that can redirect an agent, not just deny it
Execution-Layer Security (ELS) for AI agents — policy-enforced shell with audit.
At a glance
- What is it?
- Execution-layer security for coding agents, built in Go, that intercepts file, network, process and signal activity including whole subprocess trees and decides per operation. Its distinguishing feature is a redirect decision that swaps a command or a write target and returns success, which changes how an agent behaves rather than merely blocking it, and its documentation publishes a security score for every platform including the ones that are not ready.
- Who is it for?
- agentsh is worth evaluating if you let an agent run shell commands on a machine you care about, because it closes the gap that command approval cannot: everything the command spawns. Two things to weigh before deploying it.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 17 days ago.
- What is it written in?
- Mainly Go, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The platform scores are published, including the bad news
The first thing this project does, before describing its features, is publish a security score for each platform and tell you not to use two of them in production. Linux provides full enforcement and is given a 100 percent score. Windows under WSL2 also gets 100 percent, described as full Linux-equivalent enforcement, which follows from the fact that it is Linux. Native macOS enforcement through the Endpoint Security Framework and a network extension scores 90 percent and is labelled Alpha, with an explicit warning that it works end to end, that file, process and network events do flow through the system extension to the Go policy engine, but that you should expect rough edges and breaking changes between releases. Native Windows through a minifilter driver with app container scores 85 percent and is pending driver signing, so the documented position is that only WSL2 is fully supported for production there. The recommendation is Linux for production today. Publishing a score for the platforms that do not work yet is a better signal than omitting them.
Redirect is the decision that changes agent behaviour
Every comparable tool can deny an action, and the documentation is clear that denial is the weak version. A denied command produces an error, the agent tries a variation, gets denied again, and eventually either gives up or finds a path the policy did not think of. Redirect takes the same policy slot and swaps the operation instead. The documented example routes downloads: a rule matching curl and wget with a redirect decision sends them to a wrapper command with an audit flag, so the agent gets the file it asked for through a path you control. A second example catches writes under the home and temporary directories and redirects them to a scratch directory inside the workspace. The stated effect is that the agent sees a successful operation rather than an error. That is the power and the cost in one sentence: you get steering instead of a retry loop, and you also get an agent whose account of what it just did does not match what happened, which matters if anything downstream reads its narration as a record. The rule is short enough to see the shape of it:
command_rules:
- name: redirect-curl
commands: [curl, wget]
decision: redirect
message: "Downloads routed through audited fetch"
redirect_to:
command: agentsh-fetch
args: ["--audit"]Runtime enforcement, because the tool boundary is where controls stop
The argument for this whole category of tool is in the section explaining why it exists, and it is a short one. Agent workflows end up running arbitrary code: a package install, a test suite, a script somebody wrote. The usual safety control is to ask a human for approval before a command runs, but that control stops at the tool boundary, which means it cannot see anything that happens inside the command it just approved. An approved test run that writes to your home directory, or an install that runs a post-install script, looks identical to an approved test run that does neither. Enforcing at runtime means the subprocess tree is still governed, still logged, and still routed to a human when the policy requires it. The interception surface listed is correspondingly wide: file open, read, write and delete; network connections and DNS; process start and exit; terminal activity; model API requests with data loss prevention and usage tracking; database traffic through declared services; and signal delivery, which is enforced on Linux and merely audited on the other two platforms.
Five decisions, and only one of them is deny
The policy engine makes five decisions per operation, and the set is more interesting than the count. Allow and deny are the obvious pair. Approve is a human confirmation step, described as an explicit human OK rather than a timeout or a default-allow. Soft delete is the one that is unusual: it permits the operation in a way that does not commit the change, which is a materially different thing from denying and one that is hard to retrofit. Redirect is the fifth, and the one the project leads with. What is not in the list is as informative as what is. There is no per-operation rate limit, no quota, no time-bound approval and no notion of a capability that persists beyond the operation, so every decision is scoped to a single intercepted event. Whether that is the right granularity depends entirely on whether the policy you can express in those terms is the policy you actually want.
A real PostgreSQL parser sits in the dependency list
The dependency list explains the feature claims better than the prose does, and one entry in particular deserves attention. A genuine PostgreSQL query parser is a direct dependency, which is what makes per-statement classification of database traffic possible rather than per-connection. That matters because the interesting attack on a database from an agent is a single statement smuggled inside an otherwise approved query, and a connection-level check cannot see it. The rest of the list maps onto the same kind of claim: an eBPF library for kernel-level observation, a FUSE library for the filesystem interception layer, a pty library for terminal capture, a DNS library for name resolution visibility, and a set of secret backends spanning a vault client, Google Cloud key management and secret manager, one password manager, Azure key vault, and the AWS key management and secret manager services. Also present are OIDC and WebAuthn libraries and a one-time-password library, which together suggest the approval UI is authenticated rather than local-only.
The container asks for SYS_ADMIN and an unconfined AppArmor profile
The compose file is the most honest artifact in the repository, and it is worth reading as a security document rather than a deployment convenience. To function, the container adds the SYS_ADMIN capability, annotated as required for the FUSE layer and namespaces, and NET_ADMIN for network namespace and iptables work. It maps the FUSE device in from the host. And it sets an unconfined AppArmor profile, annotated as required for namespace operations, with a note that SELinux systems need the label disabled as well. Every one of those is a broad grant, and the person deploying this is being asked to hand an enforcement tool more privilege than most workloads get. That is a defensible trade for a tool whose purpose is kernel-adjacent interception, but it is a trade, and it should be made deliberately. The upside of running it in a container is that the damage from a policy mistake is bounded to the container rather than the host.
detect reports what will enforce, not what the kernel supports
There is a diagnostic command that exists because the difference between those two things is where this kind of tool usually lies to you. The detect command probes the host and reports which enforcement primitives are actually available, naming seccomp, Landlock, FUSE, eBPF, ptrace and cgroups, then groups them into per-domain protection scores alongside the selected security mode. The specific case the documentation calls out is restricted hosts, including sandboxed development environments and microvm-class platforms, where the seccomp user-notify listener cannot be installed. On such a host the command reports the mode that will actually enforce rather than what the kernel merely claims to support, which is the difference between a protection score you can rely on and one that is aspirational. There is also a form that emits a configuration tuned for the host it just probed, so the policy file matches the enforcement the machine can really deliver.
A shell shim starts the daemon, and macOS needs four build steps
Two operational details are documented more clearly than most projects manage. The first is autostart: you are told explicitly that you do not need to start the server yourself, because the first command through the tool, or any shell that has been shimmed, launches a local server using the shipped configuration file or whatever an environment variable points at, and that server keeps the FUSE layer and the policy engine alive for the session lifetime. A single environment variable turns the behaviour off if you would rather manage the lifecycle yourself. The second is the macOS build, which is not a cross-compile. The make targets split into a Go build, a Swift build, assembling an application bundle, and signing the bundle, plus a separate wrapper binary that needs cgo and a Darwin target. The repository also carries a large continuous integration surface, with a separate container file per distribution being tested and a committed audit findings document at the root.
Editorial conclusion
agentsh is worth evaluating if you let an agent run shell commands on a machine you care about, because it closes the gap that command approval cannot: everything the command spawns. Two things to weigh before deploying it. The redirect decision is a behavioural choice as much as a security one, since the agent is handed a success response while the operation is quietly rerouted, which keeps an agent on a paved path but also means its account of what it did is not literal. And the container deployment asks for the SYS_ADMIN capability plus an unconfined AppArmor profile, so the enforcement tool itself holds broad privilege on the host. The documentation's own advice is to use Linux for production; the last push was on 2026-09-15 and the newest tag is v0.20.5 from 2026-06-23.
Frequently asked questions
What is agentsh?
It is an Apache-2.0 licensed Go tool for execution-layer security on AI agents. It acts as a drop-in shell and exec endpoint that intercepts file, network, process and signal activity including whole subprocess trees, enforces a policy you define per operation, and emits structured audit events in either human-readable or compact JSON form.
Which platforms does agentsh support for production use?
The project's own recommendation is Linux, which is given a 100 percent security score for full enforcement. Windows under WSL2 also scores 100 percent as Linux-equivalent. Native macOS through the Endpoint Security Framework scores 90 percent and is in Alpha, with breaking changes expected between releases, and native Windows through a minifilter driver scores 85 percent and is pending driver signing, so only WSL2 is fully supported there for now.
What decisions can an agentsh policy make?
Five: allow, deny, approve for an explicit human confirmation, soft delete, and redirect. Redirect is the distinctive one, swapping the command or the write destination and returning guidance so the agent stays on the intended path, though the agent is given a successful result rather than an error.
What does agentsh detect do?
It probes the host for the enforcement primitives that are actually available, naming seccomp, Landlock, FUSE, eBPF, ptrace and cgroups, and reports per-domain protection scores plus the selected mode. On restricted hosts where the seccomp user-notify listener cannot install, it reports the mode that will really enforce rather than what the kernel merely supports, and it can emit a configuration tuned for that host.
How do you run agentsh in Docker?
The shipped compose file adds the SYS_ADMIN capability for the FUSE layer and namespaces plus NET_ADMIN for network work, maps the FUSE device from the host, and sets an unconfined AppArmor profile, with a label override noted for SELinux systems. It exposes port 18080 for HTTP and 9090 for gRPC, and mounts a workspaces directory plus a read-only config directory.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/canyonroad-agentsh)