Model or dataset
nobodywho-ooo/nobodywho avatar
nobodywho-ooo/nobodywho

NobodyWho: An On-Device LLM Inference Engine with Bindings for Godot, Flutter, Swift, Kotlin, Python and React Native

NobodyWho is an inference engine that lets you run LLMs locally and efficiently on any device.

1,512 stars91 forksRustEUPL-1.2

At a glance

What is it?
NobodyWho wraps llama.cpp in a Rust core and ships platform bindings so you can run GGUF chat models offline. It is aimed at app and game developers who need local inference, and its main cost is that every binding is a separate release with its own API.
Who is it for?
Adopt NobodyWho if you are shipping a Godot, Flutter, React Native, Kotlin, Swift or Python application and want a GGUF chat model running on the user's device without an API key. Do not adopt it if you need server-side batch inference, a stable cross-language API, or a Python-only stack where llama-cpp-python already covers you.
Can I use it commercially?
Yes, with conditions. EUPL-1.2 is a weak copyleft licence: you can use it inside commercial and closed-source software, but if you distribute changes to its own files, you must publish those changes under the same licence.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Rust, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem NobodyWho targets: local chat inference inside a shipped app

The project describes itself as an inference engine that lets you run LLMs locally and efficiently. That sentence is doing more work than it looks like. Most LLM tooling assumes a server: you send a prompt over HTTP and get tokens back. NobodyWho assumes the opposite. The model file lives on the device, the inference runs on the device, and the application ships the weights or downloads them at runtime.

The README lists the audience implicitly through its platform table: Kotlin, Swift, React Native and Expo, Flutter, Python, and Godot. Those are client-side runtimes. A Godot game that wants an NPC to answer free-form player text, a Flutter app that needs to summarize a note without uploading it, a Swift app that must work offline. The feature list confirms the intent: run locally and offline, no API keys, no hidden fees.

The topics list goes further than the README's headline claim, adding stt and tts alongside llm and slm. The README backs this up with speech-to-text through Whisper and text-to-speech through Kokoro, Pocket TTS and Supertonic backends. So the scope is not just chat completion. It is a local multimodal stack: text, image and audio input, plus synthesized WAV output.

That breadth is the interesting part and also the risk. A project that ships seven language bindings, three TTS backends, a Whisper integration and GPU acceleration through both Vulkan and Metal is spreading its maintenance budget across a lot of surface area. The release history suggests the team is aware: the Swift, React Native and Python bindings each carry their own version tag and their own release date rather than moving in lockstep with the core.

How the engine works: a Rust core, llama.cpp underneath, bindings on top

The README states plainly that NobodyWho is powered by llama.cpp. That is the load-bearing architectural fact. NobodyWho is not a from-scratch inference implementation; it is an orchestration and packaging layer over an existing GGUF runtime, written in Rust.

The repository layout supports this reading. There is a top-level nobodywho/ directory that holds the Rust workspace, and the justfile shows the core living at nobodywho/core, with cargo clippy run from that path. The Python binding sits at nobodywho/python and is built with maturin, which is the standard build tool for Rust-backed Python extensions. The Flutter binding sits at nobodywho/flutter/nobodywho. A Godot crate, nobodywho-godot, is built with cargo build -p nobodywho-godot. So the structure is one Rust core with several thin per-platform crates around it, rather than several independent implementations.

The feature that most distinguishes the engine from a plain llama.cpp wrapper is described as conversation-aware preemptive context shifting, which the README says retains full conversation memory without any message length limits. The mechanism, as named, is context shifting: when a conversation approaches the model's context window, earlier tokens are evicted or moved rather than causing the request to fail. The README does not document the eviction policy, so how much history survives, and in what order, is not something the documentation answers.

Tool calling is the other mechanism worth noting. The README says structured grammars are generated automatically from your function signatures, so you do not write a JSON schema by hand. In practice that means the type signature in your host language is the schema, and the binding generates a grammar that constrains the model's output to match. This is a real design decision: it trades flexibility for reliability, because a grammar-constrained decoder cannot emit a token that would break the structure.

Model loading is uniform across bindings through a path scheme. Kotlin and Swift use hf:// prefixes, Flutter and Python use huggingface:. The README also says models can be loaded from any URL. There is no mention of a local model registry or a caching layer in the README, so where downloaded weights are stored on disk is not documented there.

Installing NobodyWho for Python and running a first chat

Python is the shortest path to a working setup because the README gives a complete install command and a complete example. Install the package from PyPI:

bash
pip install nobodywho

Then construct a chat object with a model path and ask a question. The README's example uses the huggingface: prefix and a Qwen3 0.6B GGUF quantization hosted under the NobodyWho organization:

python
from nobodywho import Chat

chat = Chat("huggingface:NobodyWho/Qwen_Qwen3-0.6B-GGUF/Qwen_Qwen3-0.6B-Q4_K_M.gguf")

response = chat.ask("What is the capital of Denmark?").completed()
print(response) # The capital of Denmark is Copenhagen.

The ask method returns an object, and completed() blocks until the full response is available. That two-step shape implies a streaming path exists on the same object, though the README's example only shows the blocking call. What you should see on a successful run is the printed answer. On the first run, expect a download: the model is fetched from Hugging Face because the path is remote, and the README does not state the download size or the cache location.

The Python binding is versioned separately from the core. The most recent release listed is nobodywho-python-v2.0.0, dated 2026-08-23. If you pin dependencies, pin against that tag rather than assuming the Python package version tracks the repository as a whole.

For the other platforms the README gives install lines rather than full tutorials. Kotlin uses Maven Central with ai.nobodywho:nobodywho-android:2.0.0 for Android or ai.nobodywho:nobodywho:2.0.0 for desktop JVM. Swift uses Swift Package Manager against the nobodywho-swift repository. React Native uses npm install react-native-nobodywho, Expo uses npx expo install react-native-nobodywho, and Flutter uses flutter pub add nobodywho. Godot installs from AssetLib inside the editor, searching for NobodyWho in Godot 4.5 or later, or by importing a release zip through the AssetLib tab. The Godot instructions add a specific warning: make sure the ignore asset root option is set in the import dialogue.

Where NobodyWho is the wrong tool

The clearest limitation is the model format. The README says the engine is compatible with any LLM in the GGUF format. That is a hard boundary, not a preference. If your model of choice is published only as safetensors, or only through a hosted API, NobodyWho cannot load it. Quantization is effectively mandatory, and the README's own examples all use Q4_K_M files.

Context handling has a subtler failure mode. Preemptive context shifting is presented as a way to avoid message length limits, but shifting means information leaves the window. For a short chat that is invisible. For a long document-analysis task where the model must reason over material introduced early in the conversation, eviction can silently degrade the answer. The README does not describe how to observe or control this, so there is no documented way to tell whether a wrong answer came from the model or from history that was shifted out.

The per-platform release cadence is a practical constraint. Swift is at nobodywho-swift-v3.0.0, React Native at nobodywho-react-native-v3.0.0, and Python at nobodywho-python-v2.0.0. Three bindings, three version numbers, three release dates within a few days of each other. That is manageable if you use one binding. It is a coordination problem if you maintain, say, a Flutter app and a Python tool that must behave identically.

The justfile also reveals a maintenance tax that falls on contributors. The check target runs cargo fmt and then fails the build if formatting changed any .rs file, regenerates Python stubs and fails if nobodywho/python/nobodywho.pyi differs, regenerates Flutter doctests and fails if the generated test file differs, and runs ruff with the same diff check. Generated artifacts are committed and must be regenerated before pushing. Anyone sending a patch that touches a public signature has to run the full regeneration chain.

Finally, this is not a server. There is nothing in the README about batching requests across users, request queues, or multi-tenant serving. If you need to serve many concurrent clients from one GPU, a llama.cpp server or vLLM deployment is the right shape and NobodyWho is not.

How NobodyWho differs from llama-cpp-python and Ollama

The obvious alternative for a Python user is llama-cpp-python, which also wraps llama.cpp and also loads GGUF files. The difference is scope. llama-cpp-python exposes the llama.cpp API surface to Python and largely stops there. NobodyWho builds a higher-level Chat abstraction with tool calling, adds multimodal input, and then reimplements that same abstraction across six other runtimes. If you only ever write Python, llama-cpp-python gives you more direct control over sampling parameters and context configuration, and it does not sit behind a second layer of bindings. If you need the same behavior in a Godot game and a Flutter app, NobodyWho is the one that has already done that work.

Ollama is the other comparison worth drawing, and the difference is architectural rather than a matter of features. Ollama runs as a separate local server process and your application talks to it over HTTP on a local port. That model is convenient for development and for desktop tools, and it means one model cache serves every application on the machine. NobodyWho embeds the runtime in the application process. There is no port to configure and no daemon to keep alive, which matters on mobile and inside a game engine where spawning a background server is not viable. The trade-off is that each application carries its own copy of the runtime and its own model files.

For Godot specifically, the alternative is usually writing your own GDExtension against a C or C++ inference library. The justfile shows how much machinery that involves: a separate crate, a build step, and an integration-test project that the comment describes as needing a workaround because the editor's first import segfaults during editor-docs generation at cleanup, after having done its work. The justfile uses a fallback and an explicit check to work around it. NobodyWho has already absorbed that cost, which is the strongest argument for using it in a Godot project.

Licence, maintenance and upgrade cost

NobodyWho is licensed under EUPL-1.2, the European Union Public Licence version 1.2. It is a copyleft licence, and it is not the licence most Rust or Python developers encounter by default. The practical consequence is that EUPL-1.2 carries reciprocity obligations that differ from permissive licences like MIT or Apache-2.0, and it has a compatibility list governing which other licences can be combined with it. Whether those obligations reach your application depends on how you link the library and how you distribute it, which is a question for a lawyer rather than for this article. What is verifiable here is only the identifier: EUPL-1.2, stated in the repository.

Note also that llama.cpp, which the README names as the underlying engine, carries its own licence, and the README does not discuss how the two interact. If licence compatibility matters to your organization, that is a specific thing to check before adoption.

On maintenance, the repository is not archived and the last push was on 2026-09-09. Three binding releases landed in the weeks before that: Swift and React Native at v3.0.0 on 2026-08-22 and 2026-08-24, and Python at v2.0.0 on 2026-08-23. The core repository also carries a justfile with a check target that runs formatting, linting, stub regeneration and doctest regeneration, plus a Godot integration-test recipe. That is a project with a real CI discipline, not a weekend script.

Upgrade cost is dominated by the binding split. Moving from Python v2.0.0 to a future v3.0.0 is a Python-only change, but it may put you out of step with the Swift or React Native binding if you use more than one. The generated Python stubs and Flutter doctests are committed to the repository, which means a breaking signature change produces a visible diff rather than a silent runtime failure. That is a point in the project's favour for anyone tracking upgrades.

Editorial conclusion

Adopt NobodyWho if you are shipping a Godot, Flutter, React Native, Kotlin, Swift or Python application and want a GGUF chat model running on the user's device without an API key. Do not adopt it if you need server-side batch inference, a stable cross-language API, or a Python-only stack where llama-cpp-python already covers you. Before committing, verify two things: which release tag matches the binding you plan to use (the Swift, React Native and Python bindings are versioned independently, at v3.0.0, v3.0.0 and v2.0.0 respectively), and whether the model you intend to ship is available in GGUF on Hugging Face, because the engine loads nothing else.

Frequently asked questions

What is NobodyWho and who is it for?

NobodyWho is an inference engine that runs LLMs locally on a device, built in Rust on top of llama.cpp. It targets application and game developers, with bindings for Kotlin, Swift, Python, Flutter, React Native, Expo and Godot.

How do I install NobodyWho for Python?

Install it from PyPI with pip install nobodywho, then construct a Chat object with a model path such as huggingface:NobodyWho/Qwen_Qwen3-0.6B-GGUF/Qwen_Qwen3-0.6B-Q4_K_M.gguf and call ask followed by completed. The Python binding's most recent release is nobodywho-python-v2.0.0.

Which model formats does NobodyWho support?

The README states the engine is compatible with any LLM in the GGUF format, and all of its examples use GGUF quantizations such as Q4_K_M. Models distributed only as safetensors or behind a hosted API are not covered by that statement.

How do I install NobodyWho in Godot?

In Godot 4.5 or later, open the AssetLib tab and search for NobodyWho, or import a specific version's zip from the GitHub releases page through the same AssetLib tab. The README notes that the ignore asset root option must be set in the import dialogue.

What licence is NobodyWho released under?

The repository states EUPL-1.2, the European Union Public Licence version 1.2. It is a copyleft licence, and the README does not discuss how its terms interact with the licence of the underlying llama.cpp engine.

Official sources

  1. License: EUPL-1.2
  2. nobodywho-ooo/nobodywho on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/nobodywho-ooo-nobodywho.svg)](https://hysenlabs.com/projects/nobodywho-ooo-nobodywho)