Minecraft Agent Swarm: A Local-LLM Team That Earns Its Own Progress
A self-improving swarm of local-LLM agents that mine, smelt, build, farm, and fight their way through Minecraft as a coordinated team. Built on mineflayer + Ollama.
At a glance
- What is it?
- This project coordinates five Minecraft bots through a single Ollama model, with a hybrid skill system that includes runtime skill generation. It is aimed at streamers and AI tinkerers who want to see self-improvement without cloud APIs.
- Who is it for?
- Adopt this project if you are a streamer or an AI hobbyist who wants a visible, entertaining demonstration of multi-agent coordination on a local model, and you are comfortable debugging TypeScript and Python. Do not adopt it if you need a stable, production-grade automation tool for Minecraft or if you lack a GPU that can run a 20B MoE model at usable speed.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 7 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 6, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What the Swarm Actually Solves
The project addresses a narrow but real problem: getting multiple LLM-driven agents to cooperate on a long-horizon task without a central orchestrator or cloud dependency. The five bots, Atlas, Flora, Forge, Mason, and Blade, each have a distinct role and a leash radius that keeps them near base. They share a resource stash and a Team Bulletin, which is a simple text-based coordination channel. The README is explicit that the bots earn everything in-game: no item handouts, no teleport rescues. That constraint forces the system to improve through better prompts, skills, and action logic rather than through scripted shortcuts. This is for someone who wants to watch AI agents struggle and learn, not for someone who needs a reliable Minecraft farm automation tool. The target audience is clearly a streamer or an AI researcher who enjoys debugging emergent behavior.
Event-Driven Brain and the Single MoE Model
Each bot runs an event-driven brain in src/bot/brain.ts. The brain switches between four modes: idle for ten seconds triggers strategic thinking, hostiles or damage trigger reactive mode, an action completion triggers the critic, and chat triggers a reply. The brain feeds a context string built from World Context, Team Bulletin, and TECH TREE into the LLM, which returns an action. The notable design choice is that one Ollama model, gpt-oss:20b, serves all decision types. The README says this was chosen by A/B trial over qwen3.6 and nemotron-3-nano for structured-output quality and speed. Using a single model avoids VRAM eviction thrash between two resident models, which is a practical concern for local setups. The trade-off is that a 20B MoE model, even with active parameters around 3.6B, still requires a 32GB GPU according to the README's speed note. That is a significant hardware requirement, and it means the project is not accessible to users with modest GPUs.
Hybrid Skills: From Hand-Crafted to Runtime Generated
The skill system is the core of the self-improvement claim. It has three layers: hand-crafted TypeScript skills, 57 Voyager-style JavaScript skills, and dynamic skill generation at runtime. The Voyager loader runs JS skills in a vm sandbox, which is a security boundary but not a strong one. The skill generator uses the LLM to create new JS skills, and the reliability tracker in src/skills/reliability.ts records team-wide success rates and can auto-retire broken skills. This is a real feedback loop: if a skill fails often, it gets retired, and the LLM can generate a replacement. The README does not specify how the generator decides when to create a new skill, only that it exists. This is the most interesting part of the project, but also the riskiest. Generated code running in a sandbox can still cause unexpected behavior, and the README does not describe any containment beyond the vm sandbox. For a streamer, that unpredictability is part of the show. For a user who wants stable automation, it is a liability.
Getting It Running: Commands, Config, and the .env File
Setup requires Node.js 20+, Ollama with gpt-oss:20b pulled, a Minecraft Java server version 1.21.4 with five player slots, and Python 3.10+ for the neural combat server. The install is straightforward: git clone, npm install, and pip install -r requirements.txt. Configuration lives in a .env file. You set MC_HOST, MC_PORT, five usernames, MC_VERSION, and MC_AUTH=offline. For the LLM, you set LLM_PROVIDER=ollama, OLLAMA_HOST, OLLAMA_MODEL, and OLLAMA_FAST_MODEL. The README notes that gpt-oss:20b is used for both to avoid VRAM eviction. There is also an option to use any OpenAI-compatible API by setting LLM_PROVIDER=openai, OPENAI_BASE_URL, OPENAI_API_KEY, OPENAI_MODEL, and OPENAI_FAST_MODEL. The README warns that OPENAI_MODEL has no default and startup fails without it. It suggests running curl against the /v1/models endpoint to list valid IDs. The example IDs in the README were current in August 2026, which is a future date relative to this review, so treat them as illustrative. The project also includes a neural combat server in neural_server.py, a Python heuristic or VPT policy server that runs a 50ms tick loop over TCP. That is a separate component from the LLM brain, and it adds a second language to debug.
Streaming Features and the Cost of Complexity
The project is built for live streaming. It includes a Mission Control dashboard on port 3010, per-bot 3D viewers using prismarine-viewer, OBS overlays via WebSocket, TTS for bot thoughts, and Twitch integration. That is a lot of moving parts. The dashboard and viewers are useful for an audience, but they also increase the surface area for bugs. The README lists a safety filter in src/safety/filter.ts that blocks harmful chat and thoughts, which is a sensible addition given that the bots can chat. The trajectory capture in src/bot/trajectory.ts logs prompt, decision, and outcome, which is meant for fine-tuning. This is a thoughtful feature, but the README does not describe how to use those logs for actual fine-tuning. You get the data, but no pipeline to convert it into training examples. That is a gap. The maintenance cost is high: you need to keep Node, Python, Ollama, and a Minecraft server in sync, and the model IDs can go stale. The license is MIT, which is permissive, but the project depends on mineflayer and Ollama, each with their own licenses. The README does not give release notes or a changelog, and there are no recent releases, so you are working from the default branch.
Limitations and Failure Modes
The most obvious limitation is hardware. The README mentions roughly 200 tokens per second on a 32GB GPU for gpt-oss:20b. If you do not have that, the bots will be slow and the experience will suffer. The second limitation is the reliance on a single Minecraft version, 1.21.4. Mineflayer is version-sensitive, so upgrading the server or the game will likely break the project until the dependencies catch up. Third, the self-improvement loop is not guaranteed to converge. The reliability tracker can retire broken skills, but the generator could also produce new broken skills. There is no mention of a rollback mechanism or a cap on generated skills, so the skill library could grow with junk. Fourth, the neural combat server is a separate Python process. If it crashes, the combat bot may act erratically, and the README does not describe fallback behavior. Finally, the project is a proof of concept, not a polished product. The README is honest about that, but a potential adopter should expect rough edges. The wrong tool for this is a user who wants a set-and-forget automation. The bots will make mistakes, and the whole point is to watch and intervene.
Alternatives: Voyager and Mineflayer Alone
The closest alternative is Voyager, the original project that inspired the 57 JavaScript skills. Voyager, from NVIDIA, uses a single agent with an automatic curriculum and a skill library, but it does not have a multi-agent team or a shared stash. The key difference is that Voyager is single-agent and uses GPT-4, typically a cloud model. This project replaces that with a local MoE model and adds five agents. If you want the self-improving skill loop without the coordination complexity, Voyager is a simpler starting point. Another alternative is to use mineflayer directly, without any LLM. You can write scripted bots that mine and build reliably, but you lose the adaptive decision-making. The trade-off is reliability versus flexibility. This project sits in the middle: it uses mineflayer for the low-level actions and an LLM for the high-level decisions. If you only need deterministic behavior, mineflayer scripts are easier to debug and have no GPU requirement. The choice depends on whether you value emergent behavior over predictable output.
Editorial conclusion
Adopt this project if you are a streamer or an AI hobbyist who wants a visible, entertaining demonstration of multi-agent coordination on a local model, and you are comfortable debugging TypeScript and Python. Do not adopt it if you need a stable, production-grade automation tool for Minecraft or if you lack a GPU that can run a 20B MoE model at usable speed. Before running, verify your Ollama installation, pull the exact model ID, and confirm your Minecraft server version matches 1.21.4. Also check the README's note that the OpenAI model IDs may go stale, and test your own API access if you plan to use that path. The project is a proof of concept with a specific design trade-off: it favors self-improvement over reliability, so expect to invest time in monitoring and tuning.
Frequently asked questions
Is swarm AI free?
Running it locally is, since `LLM_PROVIDER` defaults to `ollama` against a model served on localhost, and the project is MIT licensed. The configuration also accepts any OpenAI-compatible HTTP endpoint instead, which is where cost enters, and the notes call the cheap-model split the single biggest cost lever.
Do Minecraft NPCs use AI?
The README does not describe vanilla NPCs. What it describes is five bots built on mineflayer whose decisions come from a local model, with Atlas, Flora, Forge, Mason and Blade each holding a role, a memory file and a leash radius between 100 and 500 blocks.
How many bots does minecraft-agent-swarm run?
Five, one per role: Atlas as Scout and Explorer at 500 blocks, Blade as Combat and Guard at 300, Forge as Miner and Smelter at 250, Mason as Builder at 150, and Flora as Farmer and Crafter at 100. The environment template sets MC_USERNAME through MC_USERNAME_5 to match those five names.
Has minecraft-agent-swarm shown that its agents learn?
No. The research notes state that model learning, cost savings and robotics transfer have not been demonstrated, and describe what has passed, including a scripted replay and a clean-checkout reproduction, as qualifying test infrastructure rather than a model performance result.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/jesserweigel-minecraft-agent-swarm)