Model or dataset
undreamai/LLMUnity avatar
undreamai/LLMUnity

LLMUnity: running local LLM characters inside Unity

Create characters in Unity with LLMs!

1,717 stars196 forksC#Apache-2.0

At a glance

What is it?
LLM for Unity packages a llama.cpp-based backend and a RAG layer as a Unity package, so NPC dialogue runs on the player's machine instead of a hosted API. Here is what the package ships, where it constrains you, and what to check before adopting it.
Who is it for?
Adopt LLMUnity if you are shipping a Unity game where NPC dialogue must work offline and you accept bundling model weights with the build. Do not adopt it if your dialogue logic belongs on a server you control, or if you need a model family the bundled LlamaLib backend does not cover.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 154 days ago.
What is it written in?
Mainly C#, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What LLMUnity solves for Unity developers

Unity has no built-in inference path for large language models. The usual workaround is to call a hosted API from C#, which adds a network round trip to every line of dialogue, sends player input to a third party, and stops working the moment the player is offline. LLMUnity takes the other route. It ships an inference backend and a C# API as a Unity package, so a character's response is generated on the device that runs the game.

The audience is narrow and specific: Unity developers building NPCs, chat interfaces or conversational mechanics who want the model local. The README frames the goal as creating "intelligent AI characters that players can interact with", and the package.json samples list confirms the shape of that work (SimpleInteraction, MultipleCharacters, FunctionCalling, ChatBot, KnowledgeBaseGame). This is not a general-purpose LLM serving stack. It is a component you drop into a scene.

The LlamaLib backend and the RAG layer

The architecture has two halves. The inference half is LlamaLib, a standalone C++/C# library that the README says is built on top of llama.cpp and distributed separately from the Unity package. That separation matters: the Unity package is the C# surface, while the native library does the token generation. The README states the backend runs inference on CPU and GPU across Nvidia, AMD and Apple Metal, and that it can also target a remote server instead of the local device.

The second half is retrieval. The package includes a Retrieval-Augmented Generation system for semantic search across your own data, described in the README as using ANN search, with a dedicated RAG sample. In practice this is the layer that lets a character answer from a knowledge base you supply rather than only from its training weights. The package also exposes function calling, listed as a sample with "structured output from the LLM", which is how you route a model response into game logic instead of displaying it as text.

One thing the README does not spell out is memory behaviour under load. Model weights, the KV cache and the ANN index all occupy the same budget on a player machine, and nothing in the documentation states a hard ceiling or a fallback when that budget is exceeded.

Installing LLMUnity and getting a first character to answer

The package is distributed through the Unity Asset Store (the README links the listing) and the repository is a Unity package with the identifier ai.undream.llm. The package.json declares a minimum of Unity 2022.3, specifically 2022.3.16f1, while the README's compatibility line claims testing on 2021 LTS, 2022 LTS, 2023 and Unity 6. Treat the package.json value as the binding constraint when you set up a project.

A typical install from the repository, if you are adding it as a local or git package rather than through the Asset Store, is a Package Manager manifest entry:

json
{
  "dependencies": {
    "ai.undream.llm": "https://github.com/undreamai/LLMUnity.git"
  }
}

After the package resolves, the samples become available through the Package Manager's Samples section. The README and package.json together list these: SimpleInteraction, MultipleCharacters, FunctionCalling, RAG, MobileDemo, ChatBot and KnowledgeBaseGame. Import SimpleInteraction first. It is described as "Simple interaction with an AI character", which is the smallest complete loop: a model, an LLMCharacter component and an input field.

Before the sample will generate anything you need a model file. The README's LLM model management section is where the package documents where models come from, and the MobileDemo sample is described as showing "an initial screen displaying the model download progress", which tells you that downloading weights at first run is a pattern the project supports. The README claims a single line of code is enough to send a prompt, but it does not reproduce that line in the excerpt available here, so read the quick start section on the documentation site before writing your own call.

The first thing to verify after import is not the dialogue quality. It is whether the native LlamaLib binary for your target platform resolved. If the C# components appear in the inspector but nothing generates, the backend library is the first place to look.

Where LLMUnity is the wrong choice

The package's selling point is also its main constraint: the model runs on the player's hardware. A desktop build with a discrete GPU and several gigabytes of spare RAM can host a small quantised model comfortably. A mid-range phone cannot host the same model at the same quality, and the package.json samples acknowledge this by shipping a separate MobileDemo rather than pretending one configuration covers both.

There is a second, sharper limitation. The backend is LlamaLib, built on llama.cpp. That means the set of models you can run is the set llama.cpp supports, and any model format outside that family is out of reach regardless of what the Unity-side API looks like. If your product depends on a specific hosted model's behaviour, LLMUnity will not reproduce it.

Build size is the third. Shipping weights with a game means the download grows by the size of the model, and the README does not document an official compression or streaming story beyond the download-progress screen the MobileDemo sample demonstrates. For a small itch.io release this can be the difference between a quick download and a long one.

Finally, the documentation surface is split. The README points to undream.ai/LLMUnity for docs, and the repository carries Migration.md and CHANGELOG.release.md alongside the standard changelog. That is a sign of active versioning, but it also means you should read the migration notes before moving between major versions rather than assuming the API is stable across them.

How LLMUnity differs from calling a hosted model or running Ollama yourself

The obvious alternative is a hosted API: OpenAI, Anthropic or any HTTP endpoint. The difference is not quality, it is where the data and the latency live. A hosted call gives you a larger model than any consumer device can run, at the cost of a network dependency and per-token billing. LLMUnity gives you offline operation and no per-request cost, at the cost of a smaller model and a heavier build.

A closer alternative is running a local server such as Ollama or LM Studio and having Unity talk to it over HTTP. That keeps inference local, which is the same benefit, but it moves the runtime outside the game. The player has to install and run a separate application. LLMUnity's remote server option covers the same topology when you want it, but the package's default posture is in-process, which is the difference that matters for a shipped game: one executable, no external setup.

A third option is writing your own llama.cpp binding. That is what LLMUnity already is, plus a Unity component layer, samples and a RAG implementation. The trade-off is that you inherit the project's release cadence and its API decisions instead of controlling them.

Maintenance, licensing and upgrade cost

The repository is not archived, and the last push was on 2026-04-29. The most recent release listed is v3.0.3 on 2026-03-08, following v3.0.2 in February 2026 and v3.0.1 in January 2026. That is a steady release rhythm across the 3.0 line, and the presence of Migration.md suggests the project takes breaking changes seriously enough to document them. It is not a guarantee of long-term support, and the README does not state a support window or an LTS policy.

The package is Apache-2.0. That is a permissive licence, and it is the licence on the Unity package itself. It does not automatically cover the model weights you ship, which carry their own licences from whoever published them, nor the native LlamaLib library, which is a separate repository. Check both before shipping a commercial build. Nothing here is legal advice; read LICENSE.md and the Third Party Notices file in the repository, and confirm the licence of your chosen model separately.

The upgrade cost is mostly the migration notes plus the model. Moving between major versions of a package that wraps a native library can mean re-importing samples and re-checking which model files your build expects. Budget for that rather than assuming a drop-in replacement.

Editorial conclusion

Adopt LLMUnity if you are shipping a Unity game where NPC dialogue must work offline and you accept bundling model weights with the build. Do not adopt it if your dialogue logic belongs on a server you control, or if you need a model family the bundled LlamaLib backend does not cover. Before committing, open the RAG sample, confirm which model file you intend to ship and how large it is, and read Migration.md against the 3.0.3 changelog so an upgrade from an earlier major version does not surprise you mid-project.

Frequently asked questions

Does LLMUnity cost anything to use?

The package is released under Apache-2.0, and the README describes it as free to use for both personal and commercial purposes. That covers the Unity package; the model weights you ship carry their own licences and must be checked separately.

What Unity version does LLMUnity require?

The package.json declares a minimum of Unity 2022.3, specifically 2022.3.16f1. The README states the project has been tested on 2021 LTS, 2022 LTS, 2023 and Unity 6, so the package.json value is the stricter of the two.

Does LLMUnity need an internet connection to run?

The README states the package runs locally without internet access and that no data leaves your game. It also documents a remote server setup for cases where you want inference to happen elsewhere.

Which models can LLMUnity run?

The backend is LlamaLib, which the README says is built on top of llama.cpp, so the supported model set follows that library. The README's LLM model management section is where the package documents how models are obtained and swapped.

Official sources

  1. License: Apache-2.0
  2. Project website
  3. README
  4. Releases
  5. undreamai/LLMUnity on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/undreamai-llmunity.svg)](https://hysenlabs.com/projects/undreamai-llmunity)