RWKV-Runner: A Desktop Launcher and OpenAI-Compatible Server for RWKV Models
A RWKV management and startup tool, full automation, only 8MB. And provides an interface compatible with the OpenAI API. RWKV is a large language model that is fully open source and available for commercial use.
At a glance
- What is it?
- RWKV-Runner wraps RWKV model management, a chat UI and an OpenAI-compatible HTTP API in a small Wails desktop app. It is a reasonable fit if you want a local RWKV endpoint without assembling the Python stack yourself, and a poor fit if you need a supported Android client or a fully documented production server.
- Who is it for?
- Adopt RWKV-Runner if you want a local RWKV endpoint that existing OpenAI clients can point at, and you are willing to run the backend yourself on the hardware you have. Skip it if you need an Android client (the repository ships none) or if you expect the README to tell you how to roll back a model switch or a version upgrade; it does not.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 26 days ago.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 26, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The gap RWKV-Runner fills between a model file and a usable endpoint
Running an RWKV model by hand means picking a runtime, matching CUDA or WebGPU support to your card, converting weights, and keeping a Python process alive behind something that speaks HTTP. RWKV-Runner collapses that into one executable measured in a few megabytes, plus a backend it starts for you. The README states the goal plainly: automate everything so that "every ChatGPT client is an RWKV client".
The audience is narrow but real. You are a developer or hobbyist who wants an RWKV model answering requests on localhost, and you would rather click a strategy preset than edit a launch script. You are not the target if you already run a tuned inference server and only need a thin client, though the project does allow that split: the README says you can deploy backend-python on a server and use the desktop program purely as a client by filling the server address into the Settings API URL field.
How the desktop shell, the Python backend and the OpenAI-shaped API fit together
The repository is a Wails application. The Go layer (main.go, go.mod, wails.json) builds the shell and its bindings; the frontend directory holds a TypeScript UI compiled by npm; backend-python holds the FastAPI inference service that actually loads weights. Two further backends exist in the tree, backend-golang and backend-rust, and the Dockerfile shows the shipped container path uses backend-python plus a compiled librwkv.so from rwkv.cpp.
Data flow is one-directional and simple. The UI or an external client sends an HTTP request to the backend. The backend loads the model named by the /switch-model API, runs inference, and returns a response in the OpenAI chat shape. The README points at http://127.0.0.1:8000/docs once a model is started, which is the FastAPI schema page for the endpoints. Because the surface is the OpenAI API, the client side is whatever you already use: the feature list names ChatGPT, GPT-Playground, Ollama and llmman clients, configured by filling in an API URL and API key on the Settings page.
The deployment modes are the interesting part of the design. You can run the frontend service alone, the inference service alone, or the inference service with a WebUI attached. The README's feature list calls this front-end and back-end separation and links to a deploy-examples directory with server configurations. That is more flexibility than most single-binary launchers offer, and it is the reason the project can serve both a laptop and a shared GPU box.
Installing RWKV-Runner and getting a first completion
The README routes desktop users to prebuilt installers: it links separate Windows, macOS and Linux install instructions under build/windows, build/darwin and build/linux, and the download link points at the GitHub releases page. The current release line is v1.9.12, published 2026-07-07. There are no command-line install steps for the desktop build in the README; you download the artifact for your platform and run it.
The source path is documented and is what you want on a server. Clone the repository, then start the backend inference service:
git clone https://github.com/josStorer/RWKV-Runner
cd RWKV-Runner
python ./backend-python/main.pyAfter that the README says the backend inference service has started, and that you request the /switch-model API to load a model. The API documentation is served at http://127.0.0.1:8000/docs. If you want the web UI served alongside the backend instead of the desktop shell, the README gives a single flag:
python ./backend-python/main.py --webuiFor a container deployment, the repository ships a docker-compose.yml that builds the image, exposes port 27777, mounts /mnt, and reserves one NVIDIA GPU. The compose file notes that you can append --rwkv.cpp to the backend command to use rwkv.cpp, and that the GPU reservation block should be commented out in that case:
docker compose upThe Dockerfile's final stage starts the backend with --port 27777 --host 0.0.0.0 --webui, so a container run is meant to be reachable from outside the host. Note that the compose file spells the service rmkv_runner, not rwkv_runner; use the name exactly as written when you target the service directly.
Where RWKV-Runner gets in your way
The default configuration enables a custom CUDA kernel for speed and lower VRAM use. The README is explicit that this can produce garbled output on some setups, and the fix is to open the Configs page and turn off Use Custom CUDA kernel to Accelerate, or to upgrade the GPU driver. That is a compatibility trade-off baked into the defaults, not an edge case you can ignore.
Resource exhaustion is the second sharp edge. The README warns that if you expose the service publicly you should cap request size at an API gateway, and separately cap max_tokens, because the backend default is le=102400 and a single long response can consume significant resources. The README points at backend-python/utils/rwkv.py line 567 for that limit. A local single-user setup will not notice; a shared endpoint will.
Platform coverage has a hole worth naming. The feature list marks one-click LoRA finetune as Windows only, and nothing in the repository listing or README describes an Android client, despite that being a common search phrase around the project. The build directory contains Windows, macOS and Linux instructions only. If your plan depends on training on Linux or on a phone client, this project does not cover it.
Finally, the README documents how to start services but not how to roll back a model switch or revert an automatic update. Automatic updates are listed as a feature, and the update mechanism is present in the Go dependencies, but the README is silent on downgrade procedure. Treat that as unknown rather than safe.
RWKV-Runner against running Ollama or a raw FastAPI server
The closest comparison is Ollama, because both aim at a local HTTP endpoint for a model. The difference is what sits behind the endpoint. Ollama is built around its own model registry and its own API shape, with OpenAI compatibility layered on. RWKV-Runner is built around RWKV specifically: it ships a model conversion tool, remote model inspection, download management and multi-level VRAM presets, and it speaks the OpenAI API directly rather than as an adapter. If your models are RWKV and you want the runner to fetch and convert them, RWKV-Runner is doing work Ollama does not attempt.
The other alternative is skipping the wrapper and running backend-python yourself, or running rwkv.cpp directly. The Dockerfile shows exactly what that costs: you build the frontend with Node, install Python 3.10 with the deadsnakes PPA, compile rwkv.cpp with CMake and Ninja, and copy librwkv.so into backend-python/rwkv_pip/cpp. RWKV-Runner's value is that this is already scripted. If you already have that build pipeline, the wrapper adds a Go desktop shell you may not want.
Maintenance, licensing and what an upgrade actually costs
The repository is not archived, and the last push was on 2026-09-04, so it is being touched. The release cadence visible in the tags is slow: v1.9.12 on 2026-07-07, v1.9.11 on 2026-05-08, v1.9.10 on 2026-02-01. Roughly two to three months between tagged releases. Plan upgrades around that rhythm rather than expecting continuous drops.
Upgrade cost is mostly in the model and embeddings layer, not the shell. The README carries a concrete warning: v1.4.0 improved the quality of the embeddings API, and results generated before that are not compatible with later versions. Anyone who built a knowledge base on earlier embeddings has to regenerate it. That is the pattern to expect from this project: the API surface is stable enough to point clients at, but derived data is not guaranteed across versions.
The licence is MIT, per the LICENSE file and the badge in the README. MIT is permissive and places few conditions on redistribution or commercial use of the software itself. It does not settle the licence of the RWKV model weights you download, which is a separate question the README does not address. Check the model's own terms before commercial deployment; this is not legal advice.
Editorial conclusion
Adopt RWKV-Runner if you want a local RWKV endpoint that existing OpenAI clients can point at, and you are willing to run the backend yourself on the hardware you have. Skip it if you need an Android client (the repository ships none) or if you expect the README to tell you how to roll back a model switch or a version upgrade; it does not. Before committing, check three things: that your GPU driver works with the default custom CUDA kernel, that you have set an API gateway limit in front of the request endpoint, and that you are comfortable with the backend-python deployment path rather than the desktop executable.
Frequently asked questions
What is RWKV-Runner used for?
It manages and starts RWKV models from a small desktop executable and exposes an OpenAI-compatible HTTP API, so existing ChatGPT-style clients can talk to a local RWKV model. The README also lists chat, completion and composition interfaces, model conversion, and one-click LoRA finetune on Windows.
How do I use the RWKV-Runner API?
Start the model, then open http://127.0.0.1:8000/docs to see the endpoint schema. The README notes that the API is OpenAI compatible, so clients such as ChatGPT, GPT-Playground, Ollama and llmman can be pointed at it by filling in the API URL and API key on the Settings page.
Can I download RWKV-Runner for Windows, macOS or Linux?
Yes. The README links separate install instructions under build/windows, build/darwin and build/linux, and directs downloads to the GitHub releases page. The current tagged release is v1.9.12 from 2026-07-07.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/josstorer-rwkv-runner)