# llmman stores models as OCI images and points agents at them

> A Rust command line tool whose central claim is that a model is an OCI artifact rather than a file in somebody else's blob layout: pull from Docker Hub or Hugging Face, transfer straight from one registry to another without touching a laptop, sign with cosign, and serve an Ollama, OpenAI or Anthropic compatible endpoint. Agents are launched against it in one command.

**llmmanorg/llmman** — Run any agent on any model, models stored as OCI images

- Repository: https://github.com/llmmanorg/llmman
- Website: https://llmmanorg.github.io/
- Stars: 532 · Forks: 69
- Language: Rust
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/llmmanorg-llmman

## One command starts a server, fetches a build, loads the model and execs the agent

The headline invocation is short enough to memorise:

```
llmman launch claude --model qwen3.8
```

What it does underneath is four things in sequence, and the README describes each one. It starts a local inference server. It fetches the tested llama.cpp release for your GPU, as a container on Linux when Docker or Podman is available and otherwise as a prebuilt binary. It loads the model. And then it execs an agent against that endpoint, which in the example is Claude Code. So the agent is not wrapped or emulated; it is launched against a server that behaves the way it expects. Run launch with no arguments and it lists the supported agents and whether each one is installed, which is the cheapest way to find out whether the thing you want to use is present. The same shape works for Codex, OpenCode and others, and every command takes a provider flag, so pointing the same agent at a hosted model instead of a local one is a flag rather than a different setup path. The provider names given as examples are Baseten and Groq, and the commands that take the flag are the same three the quick start introduces: launch, run and list. There is also a hybrid mode described for people who want both a local and a hosted model available and picked per request, which is the configuration that makes the agent-facing surface stable while the model behind it changes.

## Models are OCI artifacts, and transfer never lands a copy locally

This is the decision the whole project is built on, and it is worth stating plainly: models are packaged as standard OCI artifacts and stored in any compatible registry, whether that is Docker Hub, GHCR, quay, Harbor or something self-hosted. The consequence the README draws is that the registry, mirroring, access control, retention and signing infrastructure you already operate for container images works for models too, with no new platform. The transfer command is the piece that makes this more than a packaging story:

```
llmman transfer hf.co/unsloth/Qwen3.5-0.8B-GGUF docker.io/owner/model:latest
```

That copies from Hugging Face into your own registry directly, and the emphasised detail is that no copy lands in the local store on the way. For an air-gapped or compliance-bound environment that is the whole feature: the model never touches a laptop. Any source that pull understands, including an OCI registry, an hf scheme and an ms scheme, can be paired with any OCI registry destination. Running a model needs nothing special either, since pull works against Docker Hub directly and a Hugging Face reference works directly too. The registry-agnostic claim is stated in terms of what does not exist: no curated library, no account with llmman, and no gatekeeper. The consequence for an operator is that access control is whatever your registry already enforces, and the consequence for a user is that the choice of model is a URL rather than a name in somebody's index.

## Vanilla everything means no fork, no import step, no private blob format

The word used for this principle is vanilla, and the list of what is not done is more informative than the list of what is. Upstream llama.cpp releases are used as they are, or the llama-server already on your path is used instead, and vllm, sglang and mlx-lm are supported as-is. Unmodified GGUF and safetensors files are served. There is no fork to wait on, no import step, and no private blob format, because the store is a standard OCI Image Layout that other tools can read. That last clause is the one to hold on to: a model this tool has stored remains readable by anything else that speaks OCI, so you are not creating a new silo in exchange for the transfer convenience. The comparison section draws the contrast with Ollama directly, noting that Ollama keeps models in its own blob layout behind its own registry, which is the trade this project is asking you to make in the other direction. Both designs work; they differ in who owns the storage format. The dependency list backs the vanilla claim in a smaller way. The regular expression crate is there for image path detection, and the manifest says in a comment that the logic is a verbatim port of Ollama's own file name extraction regex. The web interface's terminal is served over the web server, with a shell path under the serve command, and the manifest notes that the runtime flag of serve doubles as an environment variable. Digests and signature material are handled with hashing, hex and base64 helpers, and the tar and directory-walking crates are what make unpacking an OCI layout straightforward. None of that is exotic; it is the cost of reading a container format without a container runtime.

## Signing is off by default because there is nothing to trust yet

The verification story follows directly from having no gatekeeper. Because nothing vouches for a model implicitly, the tool verifies the artifact rather than trusting the hub, and the signature format is cosign's, so cosign verify reads what llmman writes and the same check works on Docker Hub, GHCR, quay or an air-gapped mirror. Two commands cover the lifecycle:

```
llmman push docker.io/myorg/mymodel:v1 --sign-key signing.key
llmman verify docker.io/myorg/mymodel:v1 --key signing.pub
```

A trust policy written as a verify directive turns that into an automatic check on every pull, and the README says it can warn or refuse outright depending on the repository. The policy is off by default, and the stated reason is a good one: there is nothing to check against until you have said whom you trust. So the honest default is permissive, which means the safety property only exists if you configure it. That is a deliberate choice consistent with the no-gatekeeper position, and it is the thing to verify in your own environment rather than assume. The signing key is passed as a file path rather than discovered from an agent, which keeps the tool out of the business of holding credentials, and the verify command takes the public half. The dedicated documentation page for verification is linked from the README, so the policy syntax is documented separately from the commands.

## Diffusion repositories run through the same command, including video and audio

The claim about media is that a diffusion model repository works like any other one: run pulls the transformer, its VAEs, the text projection and the text encoder it was trained with. The three examples differ only in flags:

```sh
llmman run unsloth/LTX-2.3-GGUF "Draw a cat"                              # Image saved to: draw-a-cat-<timestamp>.png
llmman run unsloth/LTX-2.3-GGUF --video --seconds 2 "waves on a beach"   # an mp4 with an audio track (needs ffmpeg)
llmman run unsloth/LTX-2.3-GGUF --audio --seconds 3 "a cat purring"      # a 48 kHz stereo wav
```

So text to image, text to video with an audio track, and text to audio all go through one entry point, with a file dependency named in a comment for the video case. Run without a prompt opens an interactive loop with a prompt marker, where a slash command adjusts width, height, steps, seed, cfg, negative, seconds and media, which is the parameter set a diffusion workflow actually needs. The same model answers the image, video and audio speech endpoints when the daemon is serving, so an application can reach it over HTTP rather than only from the terminal. A separate path exists for Diffusers-layout repositories, recognised by a root model index file: Qwen-Image 2.1 runs from its safetensors directly, and anything else in that layout is served by vLLM-Omni, installed next to vllm or run from a container image. That is the one place in the runtime list where the tool is not enough on its own, because the omni runtime is a separate package, and the README says so rather than hiding it. The backends documentation carries the details for both this path and the container one.

## llmman serve exposes an Ollama, OpenAI or Anthropic compatible endpoint

Three commands cover most of what the tool does, and the third is the one that turns it into infrastructure. Launch points an agent at a local model, run is just chatting with a model, and serve is described as an Ollama, OpenAI and Anthropic compatible endpoint. The container invocation shows the shape and the port:

```sh
docker run -p 127.0.0.1:17434:17434 -e LLMMAN_API_KEYS=<key> -v llmman:/root/.local/share/llmman --entrypoint llmman ai/llmman serve
```

Three details there are deliberate. The port mapping binds to loopback rather than all interfaces, so the endpoint is not exposed by default. An API key list is passed as an environment variable, which is how the endpoint is protected. And a named volume holds the local share directory, so the store survives the container. On top of that endpoint sits the aggregation feature: several serve daemons can be grouped so that a request to any of them runs on whichever node has the model loaded or the most room for it, and the README's illustration is a laptop, a workstation and a Spark appearing as one endpoint. That is the feature that makes a heterogeneous set of machines usable as a single inference service, and it is also the one that needs a scheduler behind it. The placement rule is stated simply: name the other nodes, and a request to any of them runs on whichever node has the model loaded or the most room for it. There is no consensus protocol described and no consistency claim, so this should be read as request routing rather than a replicated service, and anyone depending on it should check the behaviour when two nodes disagree about where a model lives. The endpoint contract itself stays the same across all of them, which is the property that lets a laptop, a workstation and a Spark answer as one address.

## The manifest says 0.1.0 and the tags are at 0.1.507

The versioning scheme explains itself in a comment in the manifest and is worth reading before you pin anything. The manifest carries 0.1.0 with the major and minor components only, because continuous integration replaces the patch component with the commit count before building and publishing, so every commit on the main branch ships as a distinct patch version. To start a new series you bump major or minor and leave patch at zero. In practice that means the release list is a wall of consecutive patch numbers, and the three most recent in this repository are 0.1.505, 0.1.506 and 0.1.507, with the last of those published on the day this was checked. So there is no release cadence to speak of and no stable channel; there is a commit count. The rest of the tree explains the scope of the build. Four files at the root name the upstream releases the tool fetches, covering llama.cpp, mlx-lm, Python and uv, which is how a tool that downloads binaries keeps track of what it considers tested. There are also an Android directory for the on-device daemon, a web user interface embedded into the binary, a Go shim directory that the build needs as a CGO archive, a vLLM plugin directory, a fuzz directory that is its own crate, and a packaging directory holding the version script. The exclusion list in the manifest explains why some of those are not published: the build script needs the Go shim archive and the embedded web interface bytes, and the comment notes that an exclude list is used rather than an include list precisely so that a new build input is never dropped silently. Documentation lives in its own directory with separate pages for Android, backends and verification, and the examples directory holds a compose setup rather than a single script.

## Conclusion

Adopt llmman if the friction you actually have is moving models rather than running them, since air-gapped and compliance-bound environments are the case its transfer and signing commands are built for, and if you want an agent to talk to a local model without assembling a llama.cpp install by hand. The registry reuse argument is the real one: mirroring, access control, retention and signing infrastructure you already run for containers starts working for models the same day. Do not adopt it if you want a curated model library or a stable version cadence, because there is no library, no gatekeeper, and the manifest version is 0.1.0 while the tags are in the five hundreds. Four things to check first. Which runtimes you need, since upstream llama.cpp, vllm, sglang and mlx-lm are all supported as they are and vLLM-Omni is a separate install. How you install, since the shell one-liner, Homebrew, Scoop, cargo and a container image all exist and pick different supply chains. Whether you trust a registry at all, because verification is off until you name someone. And whether you need media generation, which works for diffusion repositories but requires ffmpeg for video.

## FAQ

### What is llmman?

A Rust command line tool for running agents against local or hosted models in one command, with models stored as standard OCI images in any compatible registry. It fetches a tested llama.cpp build for your GPU, serves an Ollama, OpenAI and Anthropic compatible endpoint, and can transfer a model from Hugging Face into your own registry without a copy landing locally.

### How do I install llmman?

A shell one-liner for Linux and macOS, an invoke-rest-method one-liner for Windows PowerShell, a Homebrew tap, a Scoop bucket, two cargo routes (a prebuilt binary via cargo binstall or a source build needing Go 1.25 and Rust, plus LLVM on Windows), an Android package for the on-device daemon, and a container image that runs serve behind a loopback port binding.

### How does llmman verify a model?

Because there is no gatekeeper, it verifies the artifact rather than trusting the hub, using cosign-format signatures so cosign verify reads what it writes. A trust policy written as a verify directive turns that into an automatic check on every pull, warning or refusing outright per repository, and it is off by default until you say whom you trust.

### Can llmman generate images, video and audio?

Yes. Running a diffusion repository pulls the transformer, its VAEs, the text projection and the text encoder, then produces an image, an mp4 with an audio track for the video case, or a 48 kHz stereo wav for the audio case. The video path needs ffmpeg, and the same model answers the images, videos and audio speech endpoints when the daemon is serving.

### How does llmman differ from Ollama?

The comparison is drawn on storage: Ollama keeps models in its own blob layout behind its own registry, while llmman stores them as standard OCI artifacts in any compatible registry, which is what makes transfer, mirroring and air-gapped copies work without a new platform. Both serve unmodified model files, and llmman uses upstream runtimes rather than a fork.

## Sources

- [License: Apache-2.0](https://github.com/llmmanorg/llmman/blob/main/LICENSE)
- [llmmanorg/llmman on GitHub](https://github.com/llmmanorg/llmman)
- [Project website](https://llmmanorg.github.io/)
- [README](https://github.com/llmmanorg/llmman/blob/main/README.md)
- [Releases](https://github.com/llmmanorg/llmman/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/llmmanorg-llmman
