# Beelzebub: an LLM-driven deception runtime for SSH, HTTP, TCP, TELNET and MCP

> Beelzebub is a Go honeypot framework that answers attackers with LLM-generated text instead of canned banners. It is low-code, plugin-extensible, and licensed GPL-3.0, which decides a lot about where you can put it.

**beelzebub-labs/beelzebub** — A secure low code deception runtime framework, leveraging AI for System Virtualization.

- Repository: https://github.com/beelzebub-labs/beelzebub
- Website: https://docs.beelzebub.ai
- Stars: 2,188 · Forks: 214
- Language: Go
- License: GPL-3.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/beelzebub-labs-beelzebub

## What Beelzebub is for, and who ends up running it

A classic honeypot answers with a fixed script. The attacker types a command, gets the same string every time, and within a few probes works out that the service is fake. Beelzebub takes the opposite position: the response is produced at request time by a language model, so the decoy can carry on a conversation that stays internally consistent for longer. The README describes it as a "deception runtime" that deploys adaptive decoys across SSH, HTTP, TCP, TELNET and MCP, and says the goal is to keep an attacker engaged long enough to collect tactics, techniques and procedures.

The audience is narrow and technical. This is not an endpoint agent you install on a laptop. It is infrastructure you stand up on a host or a Kubernetes cluster, expose on real ports, and watch. The topics attached to the repository (cloudsecurity, cybersecurity, honeypot, llm-security, mcp-honeypot, research-project) point at two overlapping groups: blue teams who want early warning and attacker telemetry, and researchers studying how AI agents behave when they meet a hostile service. The MCP service is the clearest signal of the second group. MCP is the protocol AI agents use to call tools, so a decoy MCP endpoint exists to catch prompt injection attempts aimed at an agent rather than at a human operator.

## How the runtime turns YAML into a live decoy

The architecture is a single Go binary that reads two things at startup: a core configuration file and a directory of service definitions. The CLI defaults are ./configurations/beelzebub.yaml and ./configurations/services/. Each service file describes a protocol, a port, and a set of matching rules. The README calls the matching model low-code: regex command matching, with no custom Go required to add a new decoy. When an incoming request matches a rule, the runtime either returns the configured response or hands the request to the LLM integration, which currently lists OpenAI and Ollama as backends.

The plugin layer sits underneath that. The public SDK at pkg/plugin defines three interfaces. CommandPlugin generates text for SSH, TCP, TELNET and HTTP services. HTTPPlugin returns a full response with status code, headers and body. WirePlugin observes and can rewrite matched binary TCP exchanges, and an optional WireSessionCloser releases per-connection state. Wire plugins are not implicit: a TCP service must name them under a wirePlugins key in order, and the README gives a single-entry example. Plugins register themselves through init(), so adding one means compiling it into the binary rather than loading it at runtime. That is a deliberate trade-off: it keeps the process a single static binary, and it means every plugin change is a rebuild.

Observability is built in rather than bolted on. Prometheus metrics are exposed, and events can be streamed to RabbitMQ. The docker-compose.yml publishes port 2112 and labels it Prometheus Open Metrics, which is where a scraper would point.

## Installing Beelzebub and running a first decoy

There are four documented paths: an installer script, a local Go build, Docker, and a Helm chart. The installer is the shortest route and asks whether you want the local or Docker runtime.

```bash
./install.sh     # asks local or Docker, checks prerequisites, and starts it
```

For scripted setups the README gives non-interactive flags. Note the caveat it states: on non-root hosts, local installation does not auto-start when the default configuration includes privileged ports, because binding 22 or 23 as an unprivileged user fails.

```bash
./install.sh --local
./install.sh --docker
./install.sh --local --no-run
```

The local path goes through the Makefile, which checks for Go and git, installs any plugins declared in the configuration, builds the binary, and runs it.

```bash
make start
```

Docker builds an image with the declared plugins baked in, then brings the stack up. The Dockerfile installs plugins with go run . plugin install --no-build before the final build, and passes BEELZEBUB_GITHUB_TOKEN as a build argument for private plugin repositories.

```bash
make docker
```

Before starting anything in a new environment, validate the configuration. This parses the files and reports problems without opening a single port, which is what makes it usable in CI.

```bash
beelzebub validate --conf-core ./configurations/beelzebub.yaml --conf-services ./configurations/services/
```

The runtime itself takes a core config path, a services directory, and a memory ceiling in MiB, with -1 disabling the limit.

```bash
beelzebub run -c ./configurations/beelzebub.yaml -s ./configurations/services/ -m 100
```

In Kubernetes the chart is installed from the repository directory, and the same command with upgrade applies changes.

```bash
helm install beelzebub ./beelzebub-chart
helm upgrade beelzebub ./beelzebub-chart
```

What you should see after a successful start is the process holding the configured ports. The compose file maps 22, 23, 2222, 8080, 8081, 80, 3306 and 2112, so a containerised run will collide with anything already listening on those ports on the host.

## The LLM decoy is also the weakest link

The feature that makes Beelzebub interesting is the one that introduces the most uncertainty. A model generating responses to attacker input is, by construction, being fed hostile text. The README frames prompt injection as something Beelzebub detects against AI agents, but the same class of input arrives at Beelzebub's own LLM-backed services. The documentation does not describe a hardening layer between the incoming command and the model prompt, and it does not document what happens when the model returns something malformed or refuses. Treat the LLM path as an untrusted boundary.

There is a second, quieter failure mode. Because responses are generated, the same command can produce different text on two connections. An attentive attacker who probes the same path twice may notice the inconsistency and disengage. A static honeypot is predictable but consistent; Beelzebub trades consistency for depth. Which one you want depends on whether you are optimising for time-on-target or for clean, comparable telemetry.

Operationally, the memory limit is a real constraint rather than a formality. The default is 100 MiB per the CLI reference, and the README lists per-service memory limits as a production feature. Running several protocols plus an LLM backend inside that ceiling requires measurement on your own hardware. The documentation does not publish a figure for how many concurrent sessions a given limit supports, and this review will not invent one.

Finally, the project is a research project by its own topic list. The last push was on 2026-09-09 and the most recent release, v3.9.1, was tagged on 2026-08-31, so the codebase is moving, but the interfaces and configuration schema should be expected to shift between minor versions. Pin a version in production.

## Beelzebub compared with a classic honeypot such as Cowrie

Cowrie is the obvious reference point for anyone evaluating an SSH and TELNET honeypot, and the difference is architectural rather than cosmetic. Cowrie emulates a filesystem and a shell in Python, with a large body of hand-written behaviour covering what an attacker sees after login. Its responses are deterministic because they are code. Beelzebub keeps a much thinner model of the target system and delegates the conversational part to a language model, which is why its service definitions are YAML with regex rules instead of an emulated filesystem.

The practical consequences run in both directions. Beelzebub covers protocols Cowrie does not, notably HTTP, raw TCP and MCP, and its plugin interfaces let you attach binary protocol handling through WirePlugin, which is how the README's wirePlugins example works. It also streams events to RabbitMQ and exposes Prometheus metrics as documented features rather than community add-ons. Against that, Cowrie's determinism is an asset for research: every session is reproducible, and a surprising response is a bug rather than a model sample. Beelzebub's LLM path cannot offer that guarantee. If your goal is a stable, well-understood SSH decoy with a long track record, Cowrie is the safer default. If your goal is to cover MCP or HTTP surfaces and to study how attackers react to a service that talks back, Beelzebub is aimed at exactly that gap.

## Licence and the cost of staying current

Beelzebub is GPL-3.0. That matters more here than for a library, because the plugin system compiles plugins into the same binary. The README states that plugins register via init() and are installed into the build, and the Dockerfile installs them before the final go build. A plugin linked this way is part of the distributed program. If you write a proprietary plugin and ship the resulting binary to customers, the licence terms are worth reading with your own counsel, because the usual library-friendly reading does not obviously apply. Internal deployment inside one organisation is a different question from redistribution, and this article is not legal advice.

Upgrade cost is low in mechanism and moderate in attention. The Makefile derives version, commit SHA and build date through linker flags, so beelzebub version reports exactly what you built. Helm users have a one-line upgrade. The work is in the configuration: service definitions are YAML validated against a schema, and beelzebub validate is the tool that tells you whether your files still match after a version bump. Run it against your own configurations directory before every upgrade, not against the shipped examples. The repository has moved through v3.8.0, v3.9.0 and v3.9.1 in roughly three months, which is a normal pace for a project at this stage and a reason to keep the upgrade in a staging environment first.

## Conclusion

Adopt Beelzebub if you run a security research or threat-intelligence function and can give it an isolated network segment plus an LLM endpoint you are willing to point at attacker input. Do not adopt it if you need a supported commercial product, if GPL-3.0 conflicts with how you ship software, or if you cannot accept that a decoy service can be talked into saying something wrong. Verify three things before deployment: that beelzebub validate passes on your own service directory, that the memory limit in --mem-limit-mib suits the number of services you enable, and that the ports listed in docker-compose.yml do not collide with anything already bound on the host.

## FAQ

### How do I install Beelzebub?

The README lists four routes: ./install.sh, which asks whether you want the local or Docker runtime, make start for a local Go build, make docker for a container image, and helm install beelzebub ./beelzebub-chart for Kubernetes. The installer also accepts --local and --docker for non-interactive use, and --local --no-run to build without starting.

### What is Beelzebub?

It is an open-source deception runtime written in Go that deploys decoy services across SSH, HTTP, TCP, TELNET and MCP. The README describes it as going beyond passive honeypots by engaging attackers with LLM-generated responses and detecting prompt injection against AI agents.

### How do I use Beelzebub once it is installed?

Start the runtime with beelzebub run, which takes a core configuration path, a services directory and a memory limit, then let it hold the configured ports. Before starting, beelzebub validate parses the same files and reports configuration errors without opening any service.

### Which protocols can Beelzebub emulate?

The README lists SSH, HTTP, TCP, TELNET and MCP. Each has its own deception service, and TCP services can additionally enable installed wire plugins by name under a wirePlugins key in the service configuration.

### What licence is Beelzebub released under?

GPL-3.0, per the repository licence file. Because plugins are compiled into the same binary through the init() registration described in the README, the licence question is worth raising with counsel before you distribute a build containing your own plugin code.

## Sources

- [beelzebub-labs/beelzebub on GitHub](https://github.com/beelzebub-labs/beelzebub)
- [License: GPL-3.0](https://github.com/beelzebub-labs/beelzebub/blob/main/LICENSE)
- [Project website](https://docs.beelzebub.ai)
- [README](https://github.com/beelzebub-labs/beelzebub/blob/main/README.md)
- [Releases](https://github.com/beelzebub-labs/beelzebub/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/beelzebub-labs-beelzebub
