Shimmy: a pure-Rust GGUF inference server with an OpenAI-compatible API
Pure-Rust WebGPU inference engine, OpenAI-API compatible, GGUF native, runs on any GPU. No Python. No llama.cpp. Single binary.
At a glance
- What is it?
- Shimmy is a single-binary inference server for GGUF models that speaks the OpenAI chat completions API and runs on WebGPU through its Airframe engine. The hard part is not installing it; it is knowing which model and quantization combinations are actually certified.
- Who is it for?
- Adopt Shimmy if you want a GGUF inference server that ships as one Rust binary, exposes chat completions, text completions, streaming and model endpoints on the OpenAI surface, and avoids a Python runtime. Do not adopt it if you need a model family outside the twelve listed, if you need MOE support today, or if SafeTensors inference is a requirement rather than a roadmap item.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 30 days ago.
- What is it written in?
- Mainly Rust, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Shimmy is for, and who it is not for
Shimmy exists to remove two dependencies from local inference: a Python runtime and a C++ toolchain. The README frames it as a single-binary OpenAI-compatible server for GGUF models, and the pitch is that existing AI tools pointed at it keep working. That framing matters more than the marketing line about size. If your application already talks to the OpenAI chat completions endpoint, the migration is a base URL change rather than a client rewrite.
The intended user is someone running local models on a workstation or a small server who does not want to manage a Python environment or compile llama.cpp. The README states the project is independently maintained and free forever, with sponsorship funding certification, compatibility work and releases. The Cargo.toml declares the package under the MIT license, while the repository description and the GitHub license badge point to Apache-2.0. That discrepancy is worth checking in the LICENSE file before you depend on either identifier.
It is the wrong tool if your workload depends on a model family the project has not listed. Twelve families are named, from Llama and Qwen through Phi, Gemma, DeepSeek-R1, Ministral and StarCoder2. Anything outside that set is unverified territory.
Shimmy is the server, Airframe is the engine
The architecture is split in two. Shimmy handles the HTTP surface, model discovery, configuration and the OpenAI-compatible endpoints. Airframe, pinned at v0.4.0 in the README, is the transformer engine that executes the model through WebGPU compute shaders written in WGSL. The Cargo.toml makes this explicit: the default feature set is airframe, and the comment next to it says to use --no-default-features for a CPU-only build.
Model loading is GGUF native. The README states that model specifications are auto-derived from GGUF metadata rather than hardcoded per-model constants, which is why the project claims GGUF files load as-is with no recompilation. The dependency list supports that reading: memmap2 for memory-mapped file access, safetensors for the alternate format, and shimmyjinja plus minijinja for chat template rendering.
Two engine-level claims are worth separating. The README says the engine uses F32 accumulation precision with deterministic output, meaning the same model, seed and parameters produce the same output. It also describes TurboShimmy, an INT4 KV cache mode that the README says cuts KV-cache memory by roughly seven times in tested configurations, enough to run Llama-3.2-3B on a 4 GB GPU. TurboShimmy is off by default; docker-compose.yml ships the SHIMMY_KV_QUANT=int4 line commented out.
How to install Shimmy and send a first request
The README's Quick Start installs from crates.io and starts the server against an absolute path to a GGUF file. The bind address in the example is 127.0.0.1:11435, which is not the same port the Docker image uses, so pick one and stay consistent.
cargo install shimmy
shimmy serve --model-path /absolute/path/to/model.gguf --bind 127.0.0.1:11435With the server running, the README lists models and then issues a chat completion with curl. The model field in the request body is the short name, not the file path.
shimmy list --short
curl -s http://127.0.0.1:11435/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model":"tinyllama-1.1b","messages":[{"role":"user","content":"Say hi in 5 words."}],"max_tokens":32}'You should get a JSON chat completion back on the OpenAI response shape. If you would rather not install Rust at all, the repository ships a Dockerfile and a docker-compose.yml. The compose file maps port 11434, mounts ./models into /app/models, and sets SHIMMY_BASE_GGUF to that directory.
services:
shimmy:
image: ghcr.io/michael-a-kuykendall/shimmy:latest
ports:
- "11434:11434"
volumes:
- ./models:/app/models
environment:
- SHIMMY_BASE_GGUF=/app/models
- SHIMMY_PORT=11434
- SHIMMY_HOST=0.0.0.0The README points to docs/quickstart.md for GPU setup, VRAM sizing and platform-specific builds. It does not document a rollback procedure for a failed upgrade, so plan your own before moving versions.
Certification is narrower than model support
The README is unusually careful here, and the distinction is the most useful thing on the page. Twenty-six model and quantization combinations across twelve families are certified, and certification applies to the named combination only. The README says plainly that architecture recognition does not automatically mean certification. In other words, a Llama-3.1-8B file in a quantization that is not on the list may load and may produce output, but it has not been through the project's three-box regimen of MATH, INFERENCE and DETERMINISM tests.
The table shows the pattern. Llama-3.2-1B-Instruct is certified at Q4_K_M and Q6_K. Llama-3.2-3B-Instruct is certified at Q4_K_M only. TinyLlama-1.1B-Chat appears at Q4_0, Q5_K_M and Q6_K. Most Qwen entries are Q4_K_M alone. One row, Gemma-2-9B-it, is marked as supported with certification still pending and a pointer to docs/v2-roadmap.md.
That is a real constraint, not a footnote. If you are choosing a model for a production path, the certified list is your shopping list. If you are experimenting, the wider GGUF ecosystem will probably load, but you are the one running the correctness check.
Where Shimmy runs into limits
The clearest limitation is MoE. The README lists Mixture-of-Experts CPU offloading as Airframe roadmap work, which means MoE models are not a supported path today. If your plan depends on a sparse expert model, Shimmy is the wrong layer.
SafeTensors is the second boundary. The README says .safetensors loading is supported through safetensors_native, but full Airframe-native inference for that format remains roadmap work. Loading and running are different promises, and the README keeps them separate.
Backend history is a third consideration. Version 2.0 removed the llama.cpp, MLX, HuggingFace and RustChain backends, and the Cargo.toml still carries llama, llama-cuda, llama-vulkan and llama-opencl as deprecated empty feature stubs. Anyone upgrading from v1.x should read docs/MIGRATION_v2.md rather than assume the old backend flags still select anything. The Makefile also still references a llama feature in its test target, which is a sign of how recently that removal landed.
Finally, the certification regimen is the project's own test harness. It is not an independent evaluation, and the README does not publish comparative throughput numbers against other servers.
Shimmy compared with llama.cpp and Ollama
The README positions Shimmy against Ollama directly, calling itself the 5MB alternative and describing Ollama compatibility in the crate description. The architectural difference is what matters. Ollama bundles a runtime and a model store behind its own CLI and API surface; Shimmy exposes the OpenAI endpoints and expects you to supply a GGUF path or a models directory through SHIMMY_BASE_GGUF.
Against llama.cpp, the split is about what you compile and what you ship. llama.cpp is the reference GGUF implementation in C++ and is the origin of the quantization formats Shimmy reads. Shimmy replaces that execution path with WGSL compute shaders dispatched through WebGPU, and the README's claim is that this removes the C++ toolchain from the build. The trade is maturity: llama.cpp has a far wider set of supported architectures, while Shimmy certifies twenty-six combinations.
If you already run llama.cpp and it works, Shimmy's value is packaging and the OpenAI surface, not a different model format. If you are starting fresh and want a single Rust binary with no Python, the install path is shorter.
Maintenance, upgrades and licence questions
The repository is not archived and the last push was on 2026-08-29, which is recent. Releases v2.6.1, v2.6.2 and v2.6.3 all landed on 2026-08-28 and 2026-08-29, while Cargo.toml already reads version 2.6.4, so the crate and the tagged releases can drift by a patch. Pin the version you install rather than tracking latest.
Upgrade cost is dominated by the v2.0 backend removal. The README states that v2.0 and later are a pure Airframe product, so any configuration that selected a llama.cpp, MLX, HuggingFace or RustChain backend needs rewriting. The deprecated feature stubs exist so old build commands do not fail immediately, which can hide the fact that they no longer do anything.
On licensing, the two identifiers in the repository do not agree: the repository metadata says Apache-2.0, while Cargo.toml declares MIT and the README carries an MIT badge. Both are permissive, but they carry different patent and notice obligations. Check the LICENSE and NOTICE files in the repository yourself before shipping a product that embeds the binary, and treat the crate metadata as unverified until you do.
Editorial conclusion
Adopt Shimmy if you want a GGUF inference server that ships as one Rust binary, exposes chat completions, text completions, streaming and model endpoints on the OpenAI surface, and avoids a Python runtime. Do not adopt it if you need a model family outside the twelve listed, if you need MOE support today, or if SafeTensors inference is a requirement rather than a roadmap item. Verify first that your exact model and quantization appear in the certification table, because the README states that architecture recognition does not imply certification.
Frequently asked questions
What is Shimmy?
Shimmy is a single-binary OpenAI-compatible inference server for GGUF models, written in Rust and running on the Airframe WebGPU engine. It is the server; Airframe is the engine that executes the transformer.
How to use Shimmy?
Install it with cargo install shimmy, then run shimmy serve --model-path /absolute/path/to/model.gguf --bind 127.0.0.1:11435. Point an OpenAI client at that base URL and send chat completions requests.
What does shimmy mean in slang?
The README does not define the word as slang. In this project, Shimmy is the name of a Rust inference server, and the documentation gives no other meaning.
Is it shimmy or shimmie?
The repository, crate and binary all use the spelling shimmy, including the cargo install shimmy command and the crates.io package name.
What is another word for shimmy?
The README offers no synonym. The project name is used throughout as the name of the inference server, and the crate is published as shimmy.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/michael-a-kuykendall-shimmy)