# AgentHound: an offensive security framework for MCP, A2A and the agentic stack

> AgentHound collects agent credentials and AI service inventory from a single foothold, then turns the result into a Neo4j attack graph. Here is what the collector does, how to install it, and where the design trades safety for evidence.

**adithyan-ak/AgentHound** — Offensive security framework for AI agent infrastructure - recon, credential looting, model exfiltration, poisoning, and attack-path analysis across MCP, A2A, gateways, and AI services. BloodHound for the agentic stack.

- Repository: https://github.com/adithyan-ak/AgentHound
- Website: https://agenthound.io
- Stars: 442 · Forks: 83
- Language: Go
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/adithyan-ak-agenthound

## What AgentHound is for, and who it is written for

AgentHound is a command-line offensive security framework aimed at the infrastructure that surrounds AI agents rather than the models themselves. The README frames it as "BloodHound for the agentic stack", and the comparison is structural: BloodHound took Active Directory and turned it into a queryable graph of who can reach what. AgentHound takes MCP servers, A2A agent cards, agent client configuration files, model gateways, inference servers, vector stores, MLOps services and notebooks, and does the same thing.

The intended user is a red-team operator or a penetration tester who already has a foothold on a host. The README states the workflow plainly: "Drop one static collector onto a compromised host and run one scan." The collector needs no database or server connection on the compromised machine, which matters because the analysis server usually lives somewhere the operator controls, not somewhere the target can see.

The secondary user is the defensive engineer who wants to know what an attacker would find. The same artifact that drives attack-path queries also records exposed secrets, suspicious instructions and poisoned context in agent configuration files, so a blue team can read the findings without running the graph server at all.

It is not a scanner you point at a network from the outside. Discovery is foothold-first: local agent configs, instruction files, environment-backed secrets, loopback services, active interfaces and configured endpoints all feed one scan.

## One collector, one artifact, one graph: how the pipeline actually fits together

The architecture splits into two binaries and one intermediate file. The collector runs on the target. The server runs wherever you want the graph. The artifact is the contract between them, and the README describes it as a "continuously checkpointed JSON artifact" that is ingest-valid before collection even begins.

The collector's data flow is unusual in how much it reuses its own output. Concrete secrets are saved as usable material, deduplicated by value hash, and associated with every observed source. A bearer token found in an agent config is immediately available as a candidate for validating a discovered MCP resource in the same scan. That is why the default scan is active: the README argues that "foothold access may disappear at any moment", so waiting for a second pass is a luxury the operator may not have.

Validation is differential rather than a reachability ping. For an eligible MCP resource the collector first performs an anonymous control read. If that exact read succeeds, it records public access without presenting a credential. If it fails, the collector follows with an authenticated read, and a denied control plus an allowed credentialed read becomes "Verified During Scan" evidence tied to that credential and that resource. The distinction matters in a report: reachable is an inference, a matched pair of reads is a demonstration.

Against eligible ContextForge-managed tools the collector goes further and performs a reversible mutation. It writes a scan-specific description marker, observes it through MCP, restores the original immediately, and independently confirms restoration before any other work continues. That is a real write to a target system, and the README flags it as such in the authorization notice at the top.

The server side is a Go binary that talks to Neo4j through the official driver. It joins trust relationships, authentication, credential reuse, tool capabilities, sensitive resources, protocol boundaries and observed proof into paths an operator can query: reachability, execution, exfiltration, impersonation, shadowing, poisoning and tainted data flow. PostgreSQL is also in the dependency list via pgx, so the server stack is not Neo4j alone.

## Installing AgentHound and running a first scan

The repository ships an install.sh at the top level, and the Makefile exposes build targets for each component. The README's quickstart section is the intended entry point; what follows is what the repository files show about the build and run surface.

Building from source requires Go, and go.mod pins a specific toolchain:

```
go 1.25.13
```

If you prefer to build rather than run the installer, the Makefile separates the two binaries. These targets also run a preflight check that verifies required tools are present at the expected major versions before the build starts:

```bash
make build-collector
make build-server
```

The preflight script is worth knowing about because it is the first thing that can fail. It can be bypassed when you know your environment is fine:

```bash
AGENTHOUND_SKIP_PREFLIGHT=1 make build-collector
```

There is also a Docker path. The Makefile defines `docker`, `docker-collector` and `docker-server` targets, plus `up` and `down` for the compose stack, and a `preflight-docker-compose` gate that checks the compose prerequisites first.

```bash
make docker
make up
```

For the analysis server you need a graph database. The Go module list includes `github.com/neo4j/neo4j-go-driver/v5`, so Neo4j is the graph backend the server is built against. The Makefile has a `seed` target, which is what you would run to populate an empty instance with the schema the server expects. The README does not document the connection string format or the default port in the excerpt available here, so check the docs site before you wire it up.

For the collector, the README's own framing is the operating instruction: one static binary, dropped on the host, one scan. The `--stealth` flag switches the same workflow to read-only collection. Run that first on any host where a mutation would be noticed, and compare the artifact against a default run before you decide which mode your engagement needs.

## The active default is the trade-off you have to accept

AgentHound's most consequential design decision is that the default scan writes. The README is explicit that it "performs active credential validation, model invocation, and reversible mutation when their prerequisites are present", and the authorization banner at the top of the page says to run it only against systems you own or are authorized to assess.

This is not a documentation disclaimer bolted on for legal cover. It reflects the mechanism. Credential validation means presenting a captured token to a live service. Model invocation means sending a request to an inference server. The ContextForge round trip means writing a description marker into a live tool and restoring it. Each of those can appear in application logs, in model provider billing, in audit trails, or in an alerting pipeline that a defender is watching.

The `--stealth` flag is the mitigation, and it is a genuine one: the README says it "switches the same workflow to read-only collection when OPSEC requires it". But read-only collection loses the differential proof that makes the tool's findings strong. You get inventory and captured secrets; you do not get the matched anonymous-denied, credentialed-allowed pair that turns a guess into evidence. That is the trade, and it is the right one to make deliberately rather than by default.

A second limitation is scope. The discovery surface is enumerated and specific: twelve agent client configuration formats, MCP, A2A, LiteLLM, Ollama, vLLM, MLflow, Jupyter, Qdrant and web interfaces. An agent framework outside that list is not covered. If your stack runs a bespoke orchestration layer with its own credential store, AgentHound will see the MCP servers it talks to but not the layer itself.

Recovery is handled, but it is handled by the collector, not by an external supervisor. Recovery state is persisted before mutation, cleanup runs even after cancellation, and unresolved restoration can be retried from the same artifact. That means the artifact is load-bearing: lose it and you lose the ability to complete a restoration you started.

## Where AgentHound stops and a graph platform like BloodHound begins

The obvious comparison is BloodHound, and the README invites it. The difference is in what gets collected and how.

BloodHound's collectors are domain-joined and directory-aware. They run LDAP queries and remote procedure calls against Active Directory, and the graph is built from directory objects and their ACL relationships. The collection model assumes a Windows domain with a consistent schema to enumerate.

AgentHound has no such schema to lean on. MCP servers are configured in JSON files in a dozen different client formats. A2A agents publish cards at whatever URL the operator chose. LiteLLM gateways hold master keys and virtual keys with no directory behind them. There is no central authority that knows the full set of agents, so discovery has to be local-first and opportunistic: read the configs on the host, pull secrets out of the environment, probe loopback ports, then use what you found to probe further.

That also changes what a path means. In BloodHound a path is a chain of directory permissions. In AgentHound a path can cross protocol boundaries. The README lists cross-protocol pivots as a first-class output, and the A2A section mentions "cross-protocol pathing" alongside impersonation and confused-deputy analysis. A path might start at a credential in a Cursor config, unlock an MCP server, and end at an Ollama instance whose models are reachable from the same network position.

If your target is a conventional Active Directory environment, BloodHound is the right tool and AgentHound will find almost nothing. If your target is a fleet of agent clients talking to MCP servers through a LiteLLM gateway, the reverse is true.

## Maintenance, versioning and licence

The repository is not archived, and the last push was on 2026-09-10. Recent releases are 1.0.0 on 2026-07-28, 1.1.0 on 2026-08-20 and 1.1.1 on 2026-08-23, so the release cadence over that window is roughly monthly with a patch shortly after each minor.

One versioning detail deserves attention before you pin a dependency. The go.mod file carries an explicit retraction:

```
retract [v0.5.0, v1.0.0]
```

The accompanying comment states that v0.5.0 through v1.0.0 are "superseded unsupported publications", that the historical v1.0.1 is also unsupported, and that a numeric Git tag cannot technically retract a canonical module version, so the instruction is to use the explicit revision @1.0.0. If you import the Go module rather than running the released binaries, read that comment before choosing a version. The retraction range and the recommended revision are not the same thing, and the comment is where the distinction is explained.

The licence is Apache-2.0, which permits commercial and closed-source use with the usual notice and patent-grant terms. The Makefile shows the project takes dependency licensing seriously: `GO_LICENSES` is pinned, and `ALLOWED_LICENSES` is set to a specific allowlist of Apache-2.0, MIT, BSD-2-Clause, BSD-3-Clause, ISC, MPL-2.0, Unlicense and Zlib. That is an allowlist for the project's own dependency check, not a statement about your obligations. If you redistribute AgentHound or a derivative, the Apache-2.0 terms apply to you, and the copyleft-adjacent licences on that list have their own conditions. This is not legal advice; read LICENSE and the dependency licences yourself.

Upgrade cost is modest for the binaries. The collector is static and self-contained, and the artifact format is the compatibility surface between collector and server, so the pairing that matters is collector version to server version. The preflight script is the guardrail here: `make preflight-build` and the per-target equivalents fail early with a readable message rather than a missing-command error deep in the build chain. Run the preflight before you upgrade a lab instance, not after.

## Conclusion

Adopt AgentHound if you are running authorized red-team work against MCP servers, agent clients or model gateways and you need evidence rather than reachability guesses. Do not adopt it for continuous production monitoring or as a read-only posture tool; the default scan is active and can invoke models and mutate tool descriptions. Before your first engagement, verify three things on a lab host: that install.sh resolves the release you expect, that --stealth produces the artifact your workflow requires, and that the analysis server can reach your Neo4j instance on 7687.

## FAQ

### What does AgentHound actually do on a target host?

It runs a foothold-first scan: it reads local agent configuration files, environment-backed secrets, loopback services and configured endpoints, then uses the credentials it finds to validate access to reachable AI services. Everything is written into one JSON artifact. The default scan is active, so it can validate credentials, invoke models and perform a reversible mutation on eligible ContextForge-managed tools.

### Is AgentHound safe to run against production systems?

The README says to run it only against systems you own or are authorized to assess, because it performs active credential validation, model invocation and reversible mutation. The `--stealth` flag switches the same workflow to read-only collection when OPSEC requires it, at the cost of the differential proof that authenticated reads produce.

### What database does the AgentHound analysis server use?

The Go module list includes the official Neo4j Go driver, so the attack graph is stored in Neo4j. The Makefile also has a `seed` target for loading the expected schema into an empty instance. The excerpt of the README available here does not document the connection string format or the default port.

### Which AI services and agent clients can AgentHound discover?

The README lists MCP servers, A2A agent cards, twelve agent client configuration formats including Claude Desktop, Claude Code, Cursor, Windsurf, VS Code, Cline, Continue, Zed, JetBrains, Kiro, Amazon Q and Augment, plus LiteLLM, Ollama, vLLM, MLflow, Jupyter, Qdrant and web interfaces. Anything outside that enumerated set is not covered.

## Sources

- [adithyan-ak/AgentHound on GitHub](https://github.com/adithyan-ak/AgentHound)
- [License: Apache-2.0](https://github.com/adithyan-ak/AgentHound/blob/main/LICENSE)
- [Project website](https://agenthound.io)
- [README](https://github.com/adithyan-ak/AgentHound/blob/main/README.md)
- [Releases](https://github.com/adithyan-ak/AgentHound/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/adithyan-ak-agenthound
