Sponsio checks a tool call in 0.005 ms and calls no model
Deterministic safety solutions for probabilistic AI agents
At a glance
- What is it?
- SponsioLabs/Sponsio is an agent contract layer that intercepts tool calls before they run and enforces a runtime rule against what already happened. A contract check is deterministic, measured in microseconds, and slow enough only if you compare it to an LLM judge.
- Who is it for?
- Fit for a team running an agent against other people's data, where a prompt telling the model not to do something is not a control, and where you need the check to be fast enough to sit on every single call. A poor fit if you want the rule drafted automatically and enforced verbatim, because the documentation is explicit that determinism lives in enforcement and not in drafting, and the generated contract is a draft to review.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 5 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 8, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The check happens before the tool call runs
The core mechanism is interception rather than instruction. Sponsio checks an agent's tool calls before they run.
What makes that useful is that a rule can look at what already happened. The example given is telling: checking the policy before issuing a refund is one rule instead of a paragraph of prompt. That reframes the problem from persuasion to evaluation, and it is why the project can claim to call no model at all.
An agent contract is defined as a runtime rule checked at every agent action, and the project says it is backed by formal methods, with a dedicated document on that. The keyword list in the package file gives the family a name: contracts, LTL, runtime verification.
So the enforcement surface is deliberately narrow. It is not a wrapper around a model call, it is a gate in front of one specific action type. Anything the agent does that is not a tool call is outside the layer, and the integrations list makes that boundary visible by naming frameworks rather than agent behaviours.
Microseconds against a judge that takes hundreds
The performance comparison is the argument, and it is given with the numbers attached.
One contract takes p50 of 0.0052 milliseconds to check. The heaviest workload in the benchmark, 19 contracts on every call, takes 0.139 milliseconds. p99 stays near 1 millisecond on every workload measured.
Set against an LLM-as-judge guardrail taking 50 to 800 milliseconds, the stated speedup is between 5,000 and 60,000 times, and the important qualifier is attached to it: it calls no model.
That distinction is the whole design. A judge needs a network round trip and cannot give the same answer twice, so it is either slow or inconsistent. A deterministic contract is fast and reproducible, which means it can sit on every single agent action rather than on a sampled or periodic check.
The cost of the same choice is coverage. Whatever cannot be expressed as a rule over observable state is what this layer cannot check, and that is the boundary the hosted version with an LLM judge exists to cover.
Three benchmark suites, reported separately
The evaluation claims come from three external suites, and it is worth keeping them apart rather than averaging them.
On ODCV-Bench, twelve frontier models across eighty trajectories, unguarded models cheat in 11.5 to 66.7 percent of runs. With Sponsio, 95.6 percent of misalignment is avoided on average, and 24 of 36 high-risk scenarios are at 100 percent. On the financial audit fraud finding scenario specifically, frontier models commit fraud in 16 of 24 trials and Sponsio blocks 18 of 19.
On RedCode-Exec, 1,410 cases, the combined figure is 98.9 percent: bash at 98.3 and python at 99.4, with the python number lifted from 92.4 by a four-iteration self-improvement loop. The false positive count is zero on a 60-file clean-code audit, which is the other half of the claim and the one that decides whether you can leave it on.
Then the open-core caveat. These are the open-core numbers. The hosted Cloud version adds an LLM judge layer, which takes ODCV-Bench to about 99 percent and RedCode-Exec to 99.4 percent. Two tiers, two sets of numbers, and the project is explicit about which is which.
Twenty-two bundles compose 48 patterns
Authoring a contract per rule would not scale, so the shipped surface is bundled.
Twenty-two contract bundles ship out of the box, organised by tier: always-on, per-tool, or per-incident. Each bundle is a YAML pack composed from Sponsio's deterministic patterns, and dropping one into your config guards the agent against a known failure class in one line with no per-contract authoring.
The configuration shape is a bundle inclusion list under the agent:
agents:
my_agent:
workspace: "/srv/my-bot"
include:
- sponsio:capability/destructive # gate irreversible actions
- sponsio:capability/shell # if your agent runs commands
- sponsio:capability/filesystem # if your agent touches filesUnderneath the twenty-two bundles are 48 underlying patterns, which are the primitives the bundles compose. The three shown are the capability tier, named after what they gate rather than after an attack.
The project names bundle authoring as the most useful contribution it could receive, and asks for an issue with an incident, a CVE or a pattern. That is a useful signal about where the project sees its gaps: the pattern library, not the enforcement core.
Enforcement is local, the console is where a person publishes
There are two halves and they are deliberately separated.
Enforcement is local and needs no account. The code path is three lines: build a guard from a config file, an agent identifier and a mode, then attach it.
import sponsio
import sponsio.bridge
guard = sponsio.Sponsio(config="sponsio.yaml", agent_id="mailer", mode="enforce")
run = sponsio.bridge.attach(guard)The hosted side at app.sponsio.dev is for watching runs and keeping a rulebook somewhere a person reviews it before it arms. The publishing flow is the part with a governance property: pushing a rulebook uploads it as a draft and does not arm it. A person publishes it in the console, and pulling brings the reviewed version back. Sending is best effort, so a console that is down never blocks the agent.
That last property matters more than it sounds. A hosted control plane that can stop your enforcement is a dependency in the hot path, and this one is designed so it cannot be.
Drafting is a convenience, not the guarantee
The natural question about a deterministic rule engine is whether you have to learn a formal language to use it, and the answer is partly yes.
There is a command that takes a plain-English rule and produces a contract you can read back. The documentation immediately qualifies it: treat the output as a starting draft to review and adjust before you enforce. And then the sentence that matters, that the determinism is in how contracts are enforced at runtime, not in how they are drafted.
So the drafting path saves you syntax, not judgement. It is a first draft of a rule you still have to agree with, which is a reasonable place for a model to help and exactly the place where you should not let one make the decision.
Setup follows the same pattern. Two ways in: paste a prompt into your coding agent and let it walk the onboarding flow, or run the CLI yourself. The CLI route installs the package, runs an init that asks what you use, and writes the configuration file. The wizard auto-detects your framework and prints the wrap snippet for it, and the integration list covers LangChain, Claude Agent, the OpenAI Agents SDK, Google ADK, CrewAI, Vercel AI, MCP and custom tool-calling loops, in Python or TypeScript.
A per-customer key is the fix in the current alpha
The release notes tell you what version to install and what was broken.
The current tag is v0.2.0a16, an alpha that has to be installed with a pre-release flag. The changelog entry for it is a bug fix that explains a whole class of deployment failure.
`SPONSIO_PROJECT` names the customer a run belongs to, so a per-customer key needs no code change and cannot get the customer wrong. The failure it fixes is that before this release, the attach call always claimed the default project, and a key scoped to one customer was refused. The stated consequence was that runs from a correctly wired deployment never arrived at all, which is the worst kind of bug for a telemetry path: everything looked configured and nothing was recorded.
The same release extends privacy control to the exporter. The OTLP exporter now obeys the privacy setting too, so the level you set is the level that leaves the machine on either path. Before that, the setting applied to one path and not the other, which is the sort of inconsistency a privacy setting must not have.
Four CLI dependencies, one of them load-bearing
The dependency list is short, and the comment attached to it is the interesting part.
There are four runtime dependencies: a command line library, a terminal formatting library, a prompt library, and PyYAML. The last one is called out as a hard dependency rather than an optional one, with the reason spelled out: the config loader, the CLI host and plugin install path, the rc file handling, and the plugin scan and append all import it on the core code path.
The failure mode that motivated the pin is documented too. A base install, whether with pip or with pipx, must ship it, or onboarding crashes with a module-not-found error. That is an issue number in the tree, which is a reasonable way to record why a dependency stopped being optional.
The rest of the package metadata is conventional and slightly conservative. Python 3.10 or newer, a development status of beta, Apache 2.0, and three optional dependency groups: one for LangGraph, one for the model providers used by the scan command, and more beyond what the visible file shows. The keywords include LTL, which is the clearest single word about what the contracts are.
Editorial conclusion
Fit for a team running an agent against other people's data, where a prompt telling the model not to do something is not a control, and where you need the check to be fast enough to sit on every single call. A poor fit if you want the rule drafted automatically and enforced verbatim, because the documentation is explicit that determinism lives in enforcement and not in drafting, and the generated contract is a draft to review. Before you rely on the headline numbers, read which benchmark each one came from and whether the layer you want is the open-core one or the hosted LLM judge, since the two are reported separately and the hosted one is better.
Frequently asked questions
What is Sponsio used for?
It intercepts an agent's tool calls before they run and enforces a runtime contract against them. A rule can look at what already happened, so checking a policy before issuing a refund is one rule rather than a paragraph of prompt. Each check takes microseconds and calls no model, and it works with LangChain, Claude Agent, OpenAI Agents, Google ADK, CrewAI, Vercel AI, MCP or a custom loop.
How do I install and start Sponsio?
Install the alpha with pip install --pre sponsio, or the TypeScript SDK with npm install -D @sponsio/sdk@alpha, then run sponsio init in your project, which asks what you use and writes sponsio.yaml. Enforcement is local and needs no account: build a guard from that config with an agent id and a mode, then attach it with sponsio.bridge.attach.
What benchmarks does Sponsio report, and how good is it?
Three suites, reported apart. On ODCV-Bench, twelve frontier models across eighty trajectories, 95.6 percent of misalignment is avoided on average with 24 of 36 high-risk scenarios at 100 percent. On RedCode-Exec, 1,410 cases, the combined figure is 98.9 percent with zero false positives on a 60-file clean-code audit. Those are the open-core numbers; the hosted Cloud version with an LLM judge reaches about 99 percent and 99.4 percent.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/sponsiolabs-sponsio)