AgentSafe: an authority boundary that binds to the exact action, with two of its three systems hosted
Reference architecture for AI agents that propose actions but cannot authorize them — immutable intent capture, an independent Decionis policy verdict (ALLOW/ESCALATE/BLOCK), verified human approval, and a SafeExecutor that consumes a single-use intent-bound grant.
At a glance
- What is it?
- AgentSafe is an Apache 2.0 reference architecture that intercepts consequential agent actions and asks a separate authority for one of three verdicts before forwarding anything. The thesis is that authority must bind to the action rather than the caller, and the implementation ships a three-way boundary test that runs on every packaged binary. Two of the three systems it depends on, the decision control plane and the human verification layer, are hosted services rather than code in the repository.
- Who is it for?
- AgentSafe is the clearest statement I have read of a specific and under-addressed failure mode: an agent whose identity is real, whose credential is current, and whose request nobody authorised. Most authorization work binds a decision to who is asking, and that binding is exactly what an autonomous agent defeats.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Authority binds to the action, not to the caller
The whole design reduces to one distinction, stated in the readme as a single sentence: decision intelligence determines what an AI wants to do, and execution authority determines whether it is permitted to happen. The interesting half is who the second is bound to. This boundary binds the decision to the exact action and never to the identity that proposed it. The stated reason is that the realistic adversary is not a forged credential but a valid one, an agent whose identity is real, whose credential is current and whose request is not what anyone authorised. As agents get faster and more autonomous, identity becomes a weaker proxy for authority, which is the argument for re-binding. That failure mode has a name here and a test attached to it: the Compromised Principal Test, described in the threat model as a threat and in the documentation as a test the boundary passes on every pull request, with three ways to run it in under a minute and a document saying what each one proves. A threat model that names its own central assumption and then writes a test for it is rarer than it should be.
Two of the three systems are hosted, and the substitutes admit what they are
A first-time reader meets three names and only one of them is in this repository. Decionis is the control plane, called the Independent Execution Authority, and it evaluates the organisation's policy for one captured intent and answers one of three things: allow, escalate or block. An allow comes with the single-use execution grant the request runs on, an escalate with the human ceremony it needs, and every decision with a signed dossier recording what was proposed, what was decided and why. It runs at a hosted address and is not in the tree. Presence is the adaptive human verification layer: when the verdict is escalate, a verified person on their own device approves that exact action, and the signed record is evidence the authority re-checks before issuing a grant, never authority on its own. Also hosted, also not in the tree. What is in the tree is AgentSafe, the boundary itself, which is Apache-2.0 and explicitly decides nothing. The honesty of the substitutes is the notable part: the local demo authority stands in for the hosted one with a synthetic policy and says so on every line it prints, and the loopback double for the human ceremony simulates the flow and is described as proving nothing about a real one. A separate document states the seam between the open repository and what the company operates.
Shadow first, and the local fixture can only tighten a hosted decision
The configuration is the clearest statement of the philosophy. There are two modes. In shadow, which is the default, the local fixture's verdict governs execution while the hosted verdict and its dossier are recorded beside it, so you get real decisions logged without real authority exercised. In enforcement, the hosted decision governs and its grant is claimed with the authority, and the fixture's role inverts: it can only tighten the hosted decision, never loosen it. That asymmetry is the design in one clause, because it means the local component is a second opinion that can veto and cannot authorise, which is the safe direction for a component that ships in a repository someone might clone. Two more defaults reinforce it. A provisional key, minted automatically for anyone who runs the quickstart, evaluates in shadow only, and a hosted call that exceeds its timeout budget is recorded as fail-closed rather than treated as absent. So the failure paths converge on the safe answer: no authority reachable means refuse, and a weak credential means observe.
The boundary test sends the same request three ways, and the refusal row is cut off
There is one command that tells you what the gateway does before you put it in front of anything:
agentsafe testIt sends the same consequential requests three ways at a synthetic target that records what reached it: directly, as an agent with nothing in the way; through the gateway in shadow; and through the gateway in enforcement. Nothing real is called and none of your data is read. Exit code zero means the boundary held and one means there is exposure, and the release smoke test runs this on every packaged binary. The printed table is where the design becomes concrete. A read reaches 200 in all three columns but is annotated as not consequential in enforcement, which is how the boundary declines to get in the way of traffic. A payment within policy reaches 201 directly, is marked as one that would be allowed in shadow, and in enforcement is forwarded once, returns 201 and produces a dossier. A payment above the human ceiling reaches 201 directly, and then the shadow and enforcement columns are empty in the version of the text collected here, which means the one row that demonstrates a refusal is the row you cannot see. The readme notes that the caller line is the point, since nothing was refused for who asked, only for what was asked.
The test also dials a real internal host to check it cannot go around the gateway
One flag turns the boundary test from a demonstration into an audit. Given an address, the test connects to a real system of record from where you are standing and reports whether it answers without the gateway, which is exactly what an agent could reach by going around the boundary entirely. That is a much more uncomfortable check than the synthetic one, because the answer is usually yes: agents with network access rarely have their egress restricted, and a boundary that only governs requests routed through it is a convention rather than a control. The readme is upfront that this is the last step of an activation-milestone adoption path, and it frames the run as a report you can consume as one line of JSON for automation. The hosted variant does the same exercise with the real authority deciding in shadow at the same synthetic target, producing one signed dossier per consequential request and fetching the first record back with the run's own key. A workspace can be created in one command with no account, and the hosted test is said to take about as long as the local one.
Homebrew builds start at v0.2.0 and the formula lands by its own pull request
The install story is unusually well specified, and one detail will save you an afternoon if you are debugging a version mismatch. The published install forms, including the Homebrew tap and formula, are produced by the release workflow from version 0.2.0 onward; before that, the tap pattern does not apply. After each release the formula reaches the master branch by its own separate pull request, so the tap and the repository move on different clocks and there is a documented window where the tag exists and the formula has not landed. The readme says plainly that the same commands also run from a clone, which is what the quickstart demonstrates, and that the quickstart needs no account: without a key the gateway starts with the local demo authority. Five further install paths are documented alongside the quickstart, covering macOS through Homebrew, Linux, Docker, Kubernetes and a hosted option. The adoption sequence it recommends is longer than the install instructions, and that is the right emphasis: discover, install, test the boundary, see what is exposed, run in shadow, enforce, deploy.
Twenty-two governance documents and scripts that check their own evidence
The root of this repository is mostly prose about process, and some of it is executable. Alongside the architecture, threat model, roadmap, maintenance, onboarding, governance, security, licensing and trademark documents there are files covering dependency policy and dependency licences, evidence, the evaluation path, fixture provenance and publication signoffs, plus two machine-readable rule files for discovery and coding, a licence policy file and two large plain-text files intended for language models. The scripts are where the interesting part is. One generates a corpus of decision dossiers and can check it against what is committed, so the evidence the project publishes is reproducible rather than hand-written. Another verifies an individual dossier. Another checks that every fixture has recorded provenance, which is how you tell a real captured interaction from a synthesised one. There are licence inventory and licence check scripts, a release metadata check, a dependency coverage check, a discovery check, and a fuzz script scoped to one package. A repository that ships a generator for its own audit trail and a check that the generator agrees with the committed output is doing something most projects skip.
The manifest has a literal newline inside a dependency version override
The workspace manifest is private, versioned in step with the release tag, and pinned to a package manager version and a Node floor, which is normal. One override is not. Seven dependencies are forced to specific versions, and one of those version strings contains a literal newline between two copies of the same number, so the value is a version followed by a line break followed by the version again. Whatever the resolver does with that string, it is not what the author intended, and it is the kind of typo that survives because the build still succeeds. The override list otherwise looks like a careful response to advisories rather than a reflex: several of the entries are transitive packages pinned at patched versions, one is scoped to a specific parent version, and one is a build tool's own dependency. Given that the same repository ships a dependency policy document, a licence inventory generator and a licence check script, a malformed override in the root manifest is the one thing in the packaging story that has clearly not been through the same process.
Editorial conclusion
AgentSafe is the clearest statement I have read of a specific and under-addressed failure mode: an agent whose identity is real, whose credential is current, and whose request nobody authorised. Most authorization work binds a decision to who is asking, and that binding is exactly what an autonomous agent defeats. Binding it to the exact action, with a single-use grant, a signed dossier and an optional human in the middle, is a better fit for the problem, and the project is unusually honest about where it stops. Two of its three systems are not here: the decision authority and the human verification layer are hosted services, and the loopback substitutes say on their face that they are substitutes. Before you adopt it, read the open-core document, because that seam decides whether you can run this in an air-gapped environment at all. The engineering is careful in ways that suggest real deployment: shadow before enforcement, a fixture that can only tighten a hosted decision, a timeout that records fail-closed, and provenance checks on the fixtures themselves. Confirm three things before you start: whether your environment can reach the hosted authority, whether a provisional key limited to shadow mode is enough for how you intend to stage this, and whether the direct-connection check the test performs is going to alarm your security team when it dials a real internal host to see whether it answers without the gateway.
Frequently asked questions
What does AgentSafe actually intercept?
Consequential actions. An agent, application or tool sends its HTTP request to AgentSafe instead of the target, AgentSafe captures the action as an intent, asks the control plane for a verdict of allow, escalate or block, and then forwards exactly the authorized request once on a claimed single-use grant, holds it for a person, or refuses it, leaving a chained record. AgentSafe decides nothing itself.
Are all three systems in the AgentSafe repository?
No. AgentSafe is the execution boundary and it is Apache-2.0. The decision control plane and the human verification layer are hosted services outside the repository, and the local substitutes say on their output that they are substitutes: the demo authority uses a synthetic policy, and the loopback double for the human ceremony proves nothing about a real one.
What is the difference between shadow and enforcement mode in AgentSafe?
In shadow, which is the default, the local fixture's verdict governs while the hosted verdict and its dossier are recorded alongside. In enforcement the hosted decision governs and its grant is claimed with the authority, and the fixture can only tighten that decision, never loosen it. A provisional key evaluates in shadow only, and a hosted call past its timeout budget is recorded as fail-closed.
How do I test an AgentSafe boundary before deploying it?
Run agentsafe test. It sends the same consequential requests three ways at a synthetic target, directly, through the gateway in shadow and through the gateway in enforcement, and exits zero if the boundary held or one if there is exposure. Given a host it also checks whether that host answers without the gateway, which is what an agent could reach by going around the boundary.
Why does AgentSafe bind authority to the action instead of the identity?
Because the stated realistic adversary is not a forged credential but a valid one: an agent whose identity is real, whose credential is current and whose request is not what anyone authorised. As agents get faster and more autonomous, identity becomes a weaker proxy for authority, so the grant is bound to the exact action instead.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/decionis-agent-safe-pipeline)
Community notes
Thanks for a careful read. Two updates since 10 Sept: the current release is v0.1.4 (2026-09-10), and the repository now carries a banking execution-authority profile (profiles/beap/v0.1, its site at banking.decionis.com), an onboarding journey for four workflow families (ONBOARDING.md), and a canonical-source statement — copies at other hosts, including reverse proxies of github.com, are not maintained by Decionis.