Phantom by Ghostwright: an AI co-worker that gets its own machine
An AI co-worker with its own computer. Self-evolving, persistent memory, MCP server, secure credential collection, email identity. Built on the Claude Agent SDK.
At a glance
- What is it?
- Phantom is a TypeScript agent built on the Claude Agent SDK that runs on its own VM, keeps persistent memory, and can build tools for itself. Here is how it installs, what it actually requires, and where the design trade-offs sit.
- Who is it for?
- Adopt Phantom if you can give it a dedicated VM and you want an agent that keeps state between sessions and can install its own tooling. Do not adopt it on a personal workstation, and do not adopt it if you need a hosted service with no Docker socket exposure.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 105 days ago.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The disposable-agent problem Phantom is aimed at
Most agent tooling assumes a session. You open a chat, the model answers, the tab closes, and the next conversation starts from zero. Phantom's README frames this directly: agents today are disposable, and every session is day one.
The project's answer is to give the agent a machine of its own. Not a sandbox that resets, but a persistent host where it installs packages, runs databases, and keeps memory across restarts. The README describes the target user implicitly: someone who wants a long-running collaborator reachable over chat rather than a query interface. The stated surface area is Slack, a web chat at /chat, its own email address, and webhooks. It is not positioned as a library you embed in your own application.
That framing matters for evaluation. If your need is a single-turn code assistant, Phantom is heavier than the problem. It carries a container stack, a vector store, an embedding service, and an evolution pipeline. The value only appears when the agent runs for weeks and accumulates context.
Persistent memory, MCP tools and the self-evolution loop
The architecture visible in the repository is a Bun TypeScript process (src/index.ts) that wraps the Claude Agent SDK, with Qdrant for memory and Ollama for embeddings. The docker-compose.yaml overrides QDRANT_URL to http://qdrant:6333 and OLLAMA_URL to http://ollama:11434, which tells you the two supporting services are treated as part of the deployment rather than optional extras.
Memory is the piece that separates this from a stateless wrapper. Qdrant holds vectors, Ollama produces the embeddings, and the agent reads back prior context on later runs. The README's ClickHouse story is the clearest illustration of the intended pattern: the agent installed ClickHouse on its own VM, loaded a dataset, built a dashboard and a REST API, then registered that API as an MCP tool so future sessions and other agents could query it. The MCP server is not decoration. It is how capabilities the agent creates become durable.
The second mechanism is the evolution pipeline. The README says both the main agent and every evolution judge flow through the configured provider, which means the self-improvement loop is not a separate hardcoded path. The phantom-config volume in docker-compose.yaml is mounted at /app/phantom-config and described as persistent state that survives docker compose down and up. That volume is where evolved artifacts would live, separate from the base config volume.
One thing the README does not document is how a bad evolution is rolled back. There is no rollback section in the documentation, so treat the evolved state as something you back up yourself.
Installing Phantom with Docker and reaching first contact
The README calls Docker the recommended path. Three commands fetch the compose file and the environment template, then start the stack. The only required credential in .env.example is ANTHROPIC_API_KEY for the default setup.
curl -fsSL https://raw.githubusercontent.com/ghostwright/phantom/main/docker-compose.user.yaml -o docker-compose.yaml
curl -fsSL https://raw.githubusercontent.com/ghostwright/phantom/main/.env.example -o .env
# Edit .env - add your ANTHROPIC_API_KEY, Slack tokens, and OWNER_SLACK_USER_ID
docker compose up -dAfter the containers come up, Qdrant starts for memory and Ollama pulls the embedding model. The README points at the health endpoint to confirm the agent booted.
curl http://localhost:3100/healthThe port comes from the PORT variable, which defaults to 3100 in docker-compose.yaml. If Slack is configured, the README says the agent DMs you when it is ready. Without Slack, the web chat lives at /chat, and .env.example notes that a magic link email requires OWNER_EMAIL plus RESEND_API_KEY. If Resend is not set, a bootstrap token is printed to the container logs instead, which is the path to use on a first run before you have email wired up.
Switching the model backend is a YAML edit rather than a code change. The README gives this example for Z.AI:
# phantom.yaml
model: claude-opus-4-7
provider:
type: zai
api_key_env: ZAI_API_KEY
model_mappings:
sonnet: glm-5.1Set ZAI_API_KEY in .env and restart. The README states that tools, memory and the evolution pipeline stay the same, and that only the brain changes. The provider list also includes OpenRouter, Ollama, vLLM, LiteLLM and any Anthropic Messages API compatible endpoint.
The Docker socket mount is the real security boundary
Phantom mounts /var/run/docker.sock into its own container so it can spawn sibling containers, for example for sandboxed code execution. The README is unusually direct about what that means: the socket grants root-equivalent access to the Docker daemon, and a compromised Phantom process could create, modify or destroy any container on the host.
That is not a bug to be patched later. It is the mechanism that makes the demo stories possible. An agent that installs ClickHouse and runs an observability pipeline needs to control containers. The README's own mitigations are to run Phantom on a dedicated machine or VM rather than a personal workstation, and not to expose it further. Take that seriously. On a shared host, one bad tool call or one prompt injection reaching the agent has the blast radius of your Docker daemon.
The compose file adds two smaller operational details worth knowing. The container runs non-root because, per the Dockerfile comments, the Claude Code CLI refuses --dangerously-skip-permissions as root, and socket access is granted through group_add with the host's Docker GID. The default is 988, and the README tells you to override DOCKER_GID if your host's docker group differs. It also sets oom_score_adj to -500 so the OOM killer prefers qdrant and ollama over the main process, and cpu_shares to 2048 so the reflection subprocess does not starve. Those are signs of a project that has been run under memory pressure, not just written.
What Phantom is not good at
Phantom is the wrong tool when you cannot give it a dedicated host. The socket mount plus the agent's own ability to install software means co-tenancy with anything you care about is a bad fit. A laptop running other work is exactly the deployment the README warns against.
It is also the wrong tool if you want a managed service. Everything here is self-hosted: Qdrant, Ollama, the agent container, and the volumes. There is no hosted tier described in the documentation, and the homepage link points at a site rather than a control plane. You are responsible for upgrades, backups of the phantom_config and phantom_evolved volumes, and the health of three services instead of one.
The provider story has a caveat too. Even when you switch to a non-Anthropic backend, .env.example says to keep a Claude model ID in the model field. That means the configuration surface is not fully provider-neutral, and the provider block is doing translation work rather than the agent being written against a generic interface. The README states Anthropic stays the default and existing deployments need no configuration changes, which reads as a deliberate compatibility choice rather than an oversight.
Finally, the README does not document rollback of evolved state, does not describe a resource budget for the agent's self-installed software, and gives no upgrade procedure beyond restarting the stack. Those gaps are on you to close.
How this differs from a plain Claude Agent SDK service
The closest comparison is a straightforward service built on the Claude Agent SDK, which is the dependency Phantom itself uses at ^0.2.77. A plain SDK service gives you the agent loop and tool calling. You supply the process, the state store, and the deployment. Nothing persists unless you build persistence.
The difference is that Phantom ships the persistence layer as part of the product: Qdrant for vectors, Ollama for embeddings, a mounted config volume, a separate evolved-config volume, and an MCP server so tools the agent writes stay available. It also ships the communication layer, with @slack/bolt for Slack, telegraf for Telegram, imapflow and nodemailer for email, and a React chat UI built from chat-ui/ into public/chat.
A second comparison is a workflow automation platform with an LLM node. Those give you deterministic graphs and auditability. Phantom gives you the opposite: an agent that decides to install a database because it judged that useful. The README's own stories are the evidence, and they are also the warning. The trade is control for capability, and you should pick based on which one your environment can afford to lose.
Editorial conclusion
Adopt Phantom if you can give it a dedicated VM and you want an agent that keeps state between sessions and can install its own tooling. Do not adopt it on a personal workstation, and do not adopt it if you need a hosted service with no Docker socket exposure. Before you deploy, verify the DOCKER_GID value against stat -c '%g' /var/run/docker.sock on the host, confirm the health endpoint answers at http://localhost:3100/health, and decide whether the Anthropic default or the provider block in phantom.yaml matches your budget.
Frequently asked questions
Does Phantom require an Anthropic API key?
No. Anthropic is the default and .env.example lists ANTHROPIC_API_KEY as the required credential for that setup, but the provider block in phantom.yaml supports Z.AI, OpenRouter, Ollama, vLLM, LiteLLM and custom Anthropic Messages API compatible endpoints. Even with a non-Anthropic provider, the example keeps a Claude model ID in the model field.
What port does the Phantom web chat run on?
The chat interface is served at /chat, and docker-compose.yaml maps the container port 3100 to the host using the PORT variable, which defaults to 3100. The README points at http://localhost:3100/health to check that the agent booted.
Why does Phantom mount the Docker socket?
The mount lets Phantom spawn sibling containers, for example for sandboxed code execution. The README states plainly that this grants root-equivalent access to the Docker daemon, so a compromised Phantom process could create, modify or destroy any container on the host, and recommends running it on a dedicated machine or VM.
How do I log into Phantom if Slack is not configured?
Use the web chat at /chat. With OWNER_EMAIL and RESEND_API_KEY set, Phantom sends a magic link email; without Resend, .env.example says a bootstrap token is printed to the container logs instead.
What does Phantom use for persistent memory?
Qdrant stores the vectors and Ollama produces the embeddings. docker-compose.yaml overrides QDRANT_URL to http://qdrant:6333 and OLLAMA_URL to http://ollama:11434, and the phantom_config and phantom_evolved volumes keep state across docker compose down and up.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/ghostwright-phantom)