Numen: an AI companion that plays Minecraft as a real server-side player
住在 Minecraft 里的 AI 同伴——召唤它、跟它说话,它自己规划并动手:挖矿、建造、种地、战斗、合成。
At a glance
- What is it?
- Numen puts a language-model-driven companion into your Minecraft world and drives it through native player code paths. It is a serious piece of engineering with a thin documentation surface and a per-player API key model.
- Who is it for?
- Numen suits players who already pay for an LLM API key and want a companion that mines, builds and fights under survival rules, and mod authors willing to write skills in Markdown or plugins against NumenGateway. It is the wrong choice if you want a zero-config NPC, if you cannot accept per-step model latency, or if you need a stable API: every release is labelled beta.
- Can I use it commercially?
- Yes, with conditions. LGPL-3.0 is a weak copyleft licence: you can use it inside commercial and closed-source software, but if you distribute changes to its own files, you must publish those changes under the same licence.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly Java, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Numen actually is, and who it is for
Numen places one or more AI companions in a Minecraft world. You tell a companion what to do in natural language, by typing or by holding V to speak, and it decomposes the request into dozens of steps, plans a route, picks tools and adapts as it goes. The README is explicit that it is not a chatting NPC: the companion is a server-side fake player (ServerPlayer), so mining, walking, swinging a sword and opening chests all run through native player code paths and share the same rules as redstone, mob AI and other mods.
That single design decision defines the audience. Because the companion is a real player on the server, it works on multiplayer servers, its actions are validated server-side, and you can only drive your own companion. The README states the server only needs Numen installed while each client fills in its own key. The other half of the audience is mod authors. Numen ships every built-in tool and skill written against public APIs only, and the same capability is exposed to third parties through NumenGateway and bundled skill directories.
What it is not: a hosted service. The agent loop runs on the owner's client using the owner's API key. The README frames this as a cost decision (each player pays for their own usage, server owners do not subsidise everyone) and a privacy one (the key is stored locally and sent directly to the chosen backend).
The four-part architecture: body, eyes, hands, feedback loop
The README describes the system as four parts, and the split is worth understanding because it explains most of the behaviour you will observe.
The body is a ServerPlayer. The eyes are a perception API: self and world state, ranged block and entity scans, recipe queries, single-block inspection, and reading the contents of a machine without opening its GUI (items, fluids, energy). The hands are an action API: movement, mining, placing, combat, driving arbitrary container and machine GUIs, inventory management, structure and biome location. The feedback loop is the interesting part. Every tool return, success or failure, is rendered into a sentence that teaches the model how Minecraft works. The README gives the example that punching iron ore yields a message along the lines of needing at least a stone pickaxe. The model decides its next step from that real environmental feedback rather than from a static prompt.
Pathfinding is the component with the most documented constraints. The README says movement defaults to not modifying the world: walls, floors, other players' houses and terrain stay as they are. When no clean route exists, the companion lists candidate routes with prices attached (what it would mine, what it would place), and the model either picks one with goto route:<id> or changes destination before opening a path. plan_route only computes and does not walk, and every receipt states honestly what was mined and placed along the way. The README credits Baritone's public mechanisms (weighted A*, partial path commitment, execution-time cost rechecking) while stating that no Baritone source was copied, ported or rewritten, and that Baritone drives a client-side local player whereas Numen drives a server-side fake player through server APIs.
Spatial awareness is fed to the model as an egocentric semantic character grid rather than a coordinate list, a format principle the README attributes to Gao et al., arXiv:2410.08500, adapted to a block world in three dimensions.
Installing Numen and running a first task
The README's quick start is four steps. Install the mod (on Fabric you also need Fabric API) and launch once. Press G, go to Settings, then Model, pick a provider and paste your own key. Click the + in the left column of the panel, name the companion and press Enter. Click its avatar to open the chat and describe the task.
The build badges claim Minecraft 1.20.1 through 26.2, loaders Fabric, NeoForge and Forge up to 1.20.4, and Java 17, 21 or 25. Recent releases are tagged per Minecraft version, for example v0.1.3-1.21.8-beta, so pick the artifact matching your game version rather than assuming one jar covers everything.
Ten provider presets are built in: OpenAI, Anthropic, DeepSeek, Kimi, Zhipu GLM, Doubao, Qwen, MiniMax, SiliconFlow and OpenRouter. Any OpenAI-compatible backend also works. The README notes Anthropic uses its native protocol rather than a compatibility shim.
If you are building from source rather than downloading a release, the README gives one command per loader from the repository root:
./gradlew :core:fabric:buildSubstitute :core:neoforge:build for NeoForge. The same section notes that changing engine internals means depending on core rather than only the engine, and points to api/README.md for the dependency coordinates.
Once a companion exists, the panel has three pages: Chat (conversation plus a live plan panel), Items (a read-only character sheet styled like the vanilla inventory) and Settings. Settings itself is split into ten entries: model, voice input, voice output, persona, profile, skin, theme, skill library, external brain and MCP. The README notes you can skip the panel entirely: hold R for a companion wheel, aim your crosshair at the companion, press Y for a minimal text box, or hold V to talk push-to-talk style and release to send the transcription. Keybinds live under Options, Controls, Numen.
Teaching it a mod: skills in Markdown, plugins for capability
The README draws a clean line between what works out of the box and what does not. Because the companion is a real player, it can mine and place modded blocks, open modded containers and read machines that expose a standard capability without any per-mod adaptation. What it cannot read is gameplay logic. The README names concrete examples: AE2 channel arithmetic, Create stress limits that halt a whole line when exceeded, and progression gates where an item only matters at a certain tier.
Two extension paths exist, both writable by the community. A skill is a Markdown file dropped into config/numen/skills/. No code. Five examples ship with the mod: the Nether, blaze rods, ender pearls, strongholds and the dragon fight. A plugin is a mod that wires another mod in. Installing a Create plugin makes the companion use Create; installing an AE2 one makes it understand AE2. A plugin can do two things: register tools through NumenGateway (the README's example is reading what is inside a machine), and ship skills inside its own jar so players get tools and gameplay knowledge together.
The README also covers the reverse direction. Numen can start a local MCP server so external AI clients such as Claude Desktop or Cursor drive the companion in your world. While that is enabled, the built-in brain stops entirely, because one body cannot have two brains. The endpoint and access token are on that settings page, and the token is randomly generated by default.
There is a third option for authoring tools directly. Plugins and compatibility mods may use any licence, including closed source, and the README states that separately distributed works using Numen through the API are not bound by LGPL. The engine lives in the api/ directory and is published as a separate coordinate:
repositories { maven { url = 'https://raw.githubusercontent.com/Dwinovo/numen-maven/main' } }
dependencies { modCompileOnly "com.dwinovo.numen:numen-api-fabric-1.21.1:<version>:api" }The README warns that the syntax differs per loader and that engine changes require depending on core instead.
Where Numen is the wrong tool
Every step goes through an LLM inference. The README's own FAQ answers the latency question by saying the faster the model, the smoother the experience, and that this area is still being optimised. That is a real constraint, not a footnote: a task that a human performs in ten seconds can take noticeably longer because each action depends on a round trip to a model. If you want instant, deterministic automation, a scripted NPC or a command block contraption will beat this every time.
Second, the release history is uniformly beta. All three recent releases carry the -beta suffix in both the tag and the version string. That is not a reason to avoid it, but it is a reason not to build a server around a fixed tool contract yet.
Third, the model is the intelligence. The README is candid that a faster, smarter model produces a better companion, which means behaviour quality varies with whichever backend you paste a key for. A weak or heavily rate-limited endpoint produces a companion that plans badly, and nothing in the mod fixes that.
Fourth, documentation is uneven. The README is long and unusually specific about design intent, but it does not document rollback, and it does not list the complete tool set: the closing note tells readers to look in the source tree starting at core/common/src/main/java/com/dwinovo/numen/ for the full tool list and architecture. If you need a written specification before adopting something, this is not it yet.
Finally, macOS voice input has an awkward dependency. The microphone permission declaration lives in the launcher's .app Info.plist, and the mod runs in a Java subprocess that cannot add it. The README recommends Prism Launcher specifically because it declares microphone permission, and notes the launcher is only the authorisation entry point; recording and recognition still happen in Numen and the service you configure.
How Numen differs from Baritone and from scripted NPC mods
The most instructive comparison is the one the README makes itself. Numen's pathfinding borrows Baritone's public mechanisms: weighted A*, partial path commitment, execution-time cost rechecking. But the two sit on opposite sides of the client-server line. Baritone is a client mod controlling the local player; Numen drives a server-side fake player, and movement, mining and placing all go through server APIs. That difference is why Numen works on a multiplayer server where the server only needs the mod installed, and why its actions are validated the way any player's actions are. It is also why the README can state that no Baritone source was copied, ported or rewritten.
Against scripted NPC mods the difference is the decision layer. A scripted NPC follows authored behaviour; Numen's next action comes from a model reading environmental feedback. That makes it flexible on tasks nobody scripted and unpredictable in exactly the same proportion. It also changes the cost model: a scripted NPC costs nothing per action, while every Numen step is a billed inference against your key. The README's cost guidance is that DeepSeek, Qwen, Kimi and GLM keep a typical task at a few cents, while Claude and GPT are for when you want the most capable planning.
Licence and upgrade cost
The source is LGPL-3.0. The README states the practical consequence plainly: if you distribute a modified version, it must stay open under the same licence. Separately distributed plugins and compatibility mods that use Numen through the API may use any licence, including closed source. Art and assets are all rights reserved under LICENSE-ASSETS, and the names Numen and 言出法随 are also reserved. The project is built on the MultiLoader Template.
Upgrade cost is dominated by Minecraft version churn rather than internal API churn. Releases are tagged per game version, so a server on one version cannot simply take the newest jar. For plugin authors the surface to watch is NumenGateway and the engine coordinate, which is published per loader and per Minecraft version (numen-api-fabric-1.21.1 in the README's example). The repository's top level includes agent/, ai/, api/, core/, plugins/, tools/ and ui/ directories, and the README points builders at core/ for loader-specific builds.
Maintenance is current: the last push was on 2026-09-16, and the repository is not archived. That matters less than the beta labelling, which is the honest signal about API stability.
Editorial conclusion
Numen suits players who already pay for an LLM API key and want a companion that mines, builds and fights under survival rules, and mod authors willing to write skills in Markdown or plugins against NumenGateway. It is the wrong choice if you want a zero-config NPC, if you cannot accept per-step model latency, or if you need a stable API: every release is labelled beta. Before adopting it, verify three things in your own setup: that your launcher declares microphone permission if you are on macOS, that your chosen model returns usable feedback strings, and that the engine coordinate com.dwinovo.numen:numen-api-fabric-1.21.1 actually resolves from the numen-maven repository for your loader.
Frequently asked questions
Does Numen work on a multiplayer server?
Yes. The companion is a real server-side player and its actions are validated one by one by the server, and you can only drive your own companion. The server only needs Numen installed while each client fills in its own API key.
Do I need my own API key to use Numen?
Yes. The agent loop runs on the owner's client and calls the LLM with the owner's API key, so each player pays for their own usage. The key is stored locally and sent directly to the backend you choose, not through any third-party server.
Which model should I pick for Numen?
The README suggests DeepSeek, Qwen, Kimi or GLM to keep a typical task at a few cents, and Claude or GPT when you want the most capable planning. It also states that the faster and smarter the model, the better the companion performs.
How do I teach Numen a mod it does not understand?
Write a Markdown skill and drop it into config/numen/skills/, or install a plugin that registers tools through NumenGateway and ships skills inside its own jar. The mod ships five example skills covering the Nether, blaze rods, ender pearls, strongholds and the dragon fight.
Can Numen break my base or take things that are not mine?
The README states it only does what a real player in survival can do, and that every action is ownership-checked against its owner, so it will not create items from nothing or touch what does not belong to you. Movement also defaults to not modifying the world.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/dwinovo-minecraft-numen)