# Adrian monitors your agent with a local model, and the cross-step memory is in RAM

> A runtime security layer that reads both what an agent does and why it decided to, then optionally blocks the action. The reasoning analysis is done by a model, in this case a local Gemma served by a bundled llama.cpp container with an 8K context. The design decisions are unusually well documented, down to the fact that the sliding window holding cross-step context is in memory and does not survive a restart.

**secureagentics/Adrian** — Open-source runtime AI agent security tool - monitors and controls AI agents, catching malicious tool use, prompt injection, and policy drift in real time, before the agent acts.

- Repository: https://github.com/secureagentics/Adrian
- Website: https://secureagentics.ai
- Stars: 578 · Forks: 93
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/secureagentics-adrian

## The monitor is a model too, with a bounded context

This is the architectural fact that matters and it is easy to miss. Adrian's differentiator is reading the agent's reasoning, not just its tool calls, and reasoning is read by another model. In the self-hosted configuration that model is a local Gemma, served by a llama.cpp container inside the same compose stack, and the engine posts to its chat completions endpoint verbatim without appending a path. The configuration states the context window explicitly, defaulting to 8192, with a comment that a larger value buys more history at the cost of more video memory. So the monitor has its own comprehension limit, its own latency, and its own failure modes. An action justified by reasoning that exceeds the window will be classified with only the tail of that reasoning in view, which is a different question from whether the action was correct.

## Cross-step detection lives in a ring buffer that a restart erases

Single-turn classification is easy; the hard case is an attack split across steps. The configuration describes the mechanism in unusual detail. Each tuple of session, invocation, and agent identifier keeps a ring buffer of the last sixteen classified turns, and the classifier prepends them as alternating user and assistant pairs so that detections needing cross-step context can fire. Three caveats are written into the same file rather than a footnote. The buffer is in memory, so restarting the process loses the window and the multi-step picture with it. The lifetime is one day, which bounds how long a slow, patient sequence can stay invisible. And the lock that serialises read, classify, publish, and push is process-local, which means two replicas of the backend keep separate locks and separate windows with nothing reconciling them. None of that is a bug report; it is a map of what the design cannot see.

## The recommended install is a markdown file handed to your agent

The quickstart opens by suggesting you feed a file to a coding agent and let it install the product for you, describing it as a hands-off sixty-second install, with a link to the guide and a video walkthrough. The instruction that follows is worth taking seriously in a security product: always review the instructions manually. That is an honest admission about how the product is expected to be adopted, and it also means the first thing an untrusted agent reads on your behalf is a document from a repository you have not audited. The alternative paths are ordinary. You install the SDK and sign up for the hosted dashboard to generate an API key:

```sh
pip install adrian-sdk
```

Then you configure the agent's remit, choose audit or block mode, set alerting channels, list behaviours you accept as known risks, and let events appear in the dashboard within seconds, classified by severity. Or self-host. Three routes, and only the last one keeps the data on your machine.

## The quickstart wraps a single model call, not an agent

The claim is that two lines bracket your code. The example shows an initialise call with an API key, a model invocation, and a shutdown call, and the invocation is a single call to a chat model with a prompt about finding underpriced recent IPOs. That is the wiring, not the product. There is no tool call, no loop, no multi-step trajectory, which means the example never exercises the one capability the README leads with, since single-call reasoning needs no cross-step window and no action blocking. The prose does say the SDK auto-instruments LangChain and LangGraph and that installing LangChain pulls in the graph library, so both agent constructors are covered. A fuller runnable version with environment checks is referenced in the examples directory. Worth reading before you conclude the two lines are all you need to review.

## The dependency pin carries a verification date, and it is stale

Beneath the install command there is a footnote doing something most projects skip: it records the exact versions the integration was last checked against, and the date it was checked. Four packages are named with full version numbers, plus the note that installing LangChain brings in the graph library, so the single-agent and the agentic constructors are both covered by that one install. Supported ranges are then given as a bounded window on those packages. The stamp says the last verification was on 2026-06-24, which is a little over three months before this. For a tool whose job is to sit inside another library's call path, that date is the most useful line in the section, because a framework minor release can change what an instrumented callback sees without changing anything in this repository.

## The image is build-only, on purpose, because names get squatted

One line in the compose file carries a comment longer than the setting it explains. The backend service sets its pull policy to build rather than pull, and the stated reason is to stop someone registering that image name on a public registry from having their image silently pulled and run in your stack. The image name is the fully local-looking name the project would naturally reach for, which is exactly the name someone else can take. Two other details in the same service are worth noticing. The environment file is declared optional, specifically so the bootstrap command can run on a fresh clone before the file exists and write it. And the compose file is organised into profiles, with a setup profile whose subcommands include bootstrapping, resetting a password, and pointing at a different model file, so those one-shot operations do not require the long-running stack to be up.

## The container port is pinned so the published mapping has something to reach

There is a long comment explaining why the port inside the container is hardcoded while the host-side port stays configurable. The reasoning is a good lesson in container networking: if the process inside the container reads the host-side port value, it tries to bind that port in its own network namespace, and the published mapping then forwards to a port where nothing is listening. So the internal value is pinned, the host value only controls the mapping, and the two are deliberately decoupled. The volumes follow the same pattern of thought. Data is mounted read-write because it holds the SQLite database, and the models directory is mounted read-only, which is the right default for a multi-gigabyte model file the service consumes but never modifies.

## Two distributions share a name, and one file is called CLA.md

The developer Makefile documents a local convention that explains a class of otherwise baffling bug reports: every Python entry point runs through a project-local virtual environment, specifically so the bundled SDK install cannot collide with a system-wide wheel of the same name already on the package index. Both exist. The SDK is published, and it is also in this tree, so without the convention a developer who has the published wheel installed globally gets the wrong code with no obvious symptom. The file layout shows the same attention. There is a licence header file alongside an editor config and a pre-commit configuration, which together imply headers are enforced rather than remembered. And there is a file named with a three-letter abbreviation where you would expect a longer name, sitting next to the contributing guide and the security policy.

## Conclusion

Use it if your agent can take actions you would have to undo, because remit-based judgement catches things a prompt-injection classifier never saw, and the e-commerce password-reset example is the case that sells it. Two limits to plan around. The classifier runs on a model with a bounded context, so reasoning longer than that window is judged without its own history. And the cross-step sliding window lives in memory, so an attack engineered to straddle a restart is outside what it can see, which argues for keeping sessions short enough to audit rather than assuming continuity.

## FAQ

### What is Adrian?

An open-source, Apache-2.0 runtime security monitoring and control engine for AI agents, aligned with the AARM framework. It analyses both activity logs covering tool calls, actions, and outputs, and the agent's reasoning traces, in order to detect malicious, misaligned, or out-of-remit behaviour, and it can optionally intervene while an action is in flight.

### What kinds of agent attacks does Adrian catch?

Prompt injection and jailbreaks, both direct and indirect; tool poisoning and unsafe or off-policy tool calls; data exfiltration and secret or credential leakage; and privilege escalation and out-of-remit actions. The claimed advantage is judging each action against a working understanding of what your specific agent is meant to do, correlated across the session, rather than only against categories seen in training data.

### How do I install Adrian?

Install the SDK with pip install adrian-sdk, then install LangChain and the provider matching your model, after which the SDK auto-instruments LangChain and LangGraph. Alternatively add the Claude Code plugin with two slash commands, which classifies every tool call in the terminal without code changes, or self-host the whole stack with docker compose.

### Can I self-host Adrian?

Yes, with offline, data-sovereign deployment. The repository ships the Go backend with its WebSocket, dashboard API and inference engine, the dashboard front end, the Python SDK, and a llama.cpp container serving a local Gemma mode. Bringing the stack up uses a compose profile that requires the NVIDIA Container Toolkit, and the model file has to be present on the host before the container can serve it.

### How much context does Adrian's own classifier see?

The bundled local model's context window defaults to 8192, with the configuration noting that a larger value buys more history at the cost of more video memory. Cross-step detection is supported by a sliding window of sixteen classified turns per session, prepended as alternating user and assistant pairs, kept in memory only, so a restart loses it.

## Sources

- [Issues](https://github.com/secureagentics/Adrian/issues)
- [License: Apache-2.0](https://github.com/secureagentics/Adrian/blob/main/LICENSE)
- [Project website](https://secureagentics.ai)
- [README](https://github.com/secureagentics/Adrian/blob/main/README.md)
- [secureagentics/Adrian on GitHub](https://github.com/secureagentics/Adrian)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/secureagentics-adrian
