AI00 RWKV Server: a Vulkan-based OpenAI-compatible inference box
The all-in-one RWKV runtime box with embed, RAG, AI agents, and more.
At a glance
- What is it?
- AI00 RWKV Server runs RWKV language models on any Vulkan-capable GPU and exposes an OpenAI-shaped HTTP API. It is a narrow tool with a real niche, and the README is thinner than the feature list suggests.
- Who is it for?
- Adopt AI00 RWKV Server if you already run RWKV weights and your GPU is AMD, Intel or integrated, since the Vulkan path avoids CUDA entirely. Do not adopt it if you need a broad model catalogue, a documented rollback path or a support contract; the README covers neither.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 113 days ago.
- What is it written in?
- Mainly Rust, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The gap AI00 RWKV Server fills: RWKV weights without a CUDA stack
Most local inference servers assume an Nvidia card and a PyTorch install. AI00 RWKV Server does not. The README describes it as an inference API server for the RWKV language model built on the web-rwkv engine, and the distinguishing claim is that it supports Vulkan parallel and concurrent batched inference on all GPUs that support Vulkan, explicitly including AMD cards and integrated graphics. The second claim is that no bulky pytorch, CUDA or other runtime environment is required.
That combination defines the audience. If you have RWKV weights and a machine whose GPU is not Nvidia, the usual path is painful: CUDA-only runtimes will not start, and CPU inference on a multi-billion parameter model is slow enough to change what you can build. AI00 targets exactly that machine. It is not a general model runner. It runs RWKV, and the model download links in the README point to V5, V6 and V7 checkpoints hosted on Hugging Face rather than to a catalogue of third-party models.
The compatibility surface is the other half of the pitch. The API follows the OpenAI specification, so a client already written against OpenAI's SDK can be pointed at a local base URL instead. The README lists chat, completions, embeddings and models endpoints, and notes that chat and completions accept additional optional fields for advanced functionality.
How the Rust workspace is put together
The repository is a Cargo workspace with two members, crates/ai00-core and crates/ai00-server, and the default member is the server crate. That split is the architecture in miniature: core holds the model and inference logic, the server crate wraps it in an HTTP layer.
The web server dependency is salvo, and the async runtime is tokio with the full feature set. The inference engine is web-rwkv at version 0.10.20, pulled with default features off and the native feature on. Model weights are read through safetensors 0.6 and memmap2 0.9, which is consistent with the README's statement that only safetensors models with the .st extension are supported. Weights are memory-mapped rather than copied into an owned buffer, which is what you would expect for multi-gigabyte checkpoints.
Two details in Cargo.toml are worth flagging. The workspace patches crates.io so that hf-hub comes from a fork at github.com/cgisky1980/hf-hub on the main branch, and ort is replaced by a local path at crates/ort-patch. Depending on a git branch for a transitive dependency means a build can change without any change in this repository. The release profile also sets lto = false, which keeps link times down at the cost of some runtime performance. Both are ordinary engineering choices, but they affect how reproducible a given build is.
The declared rust-version is 1.76 in Cargo.toml while the README badge says Rust 1.78.0+, and the workspace version field still reads 0.6.2 while the most recent release is tagged v0.7.1. Those are minor inconsistencies, not blockers, but they mean you should trust the release tag over the manifest when you are tracking versions.
Installing AI00 RWKV Server and making a first call
The README gives two routes: a pre-built executable from the releases page, or a source build. The pre-built route is four steps. Download the release, put a model file under assets/models/, optionally edit assets/configs/Config.toml, then run the binary.
./ai00_rwkv_serverAfter that the README says to open a browser at http://localhost:65530, or https://localhost:65530 if tls is enabled. The same port serves the WebUI and the API.
Building from source needs Rust installed, plus a clone and a release build:
git clone https://github.com/cgisky1980/ai00_rwkv_server.git
cd ai00_rwkv_server
cargo build --release
cargo run --releaseThe model file is the step people skip. The README names the expected layout, for example assets/models/RWKV-x060-World-3B-v2-20240228-ctx4096.st, and says the model path is set in assets/configs/Config.toml. If your weights are .pth files saved by torch, they must be converted first, either with the Python script or with the converter binary shipped in the release:
python assets/scripts/convert_safetensors.py --input /path/to/model.pth --output /path/to/model.stThe script needs Python with torch and safetensors installed. If you would rather not install Python, the release includes an executable called converter that takes the same --input and --output flags, and from a source checkout you can run it as cargo run --release --package converter -- --input /path/to/model.pth --output /path/to/model.st.
The server accepts three command-line arguments: --config for the configuration file path (default assets/configs/Config.toml), --ip for the bound address, and --port. Once it is up, the README's Python example points the OpenAI SDK at the local base URL and uses a placeholder key:
import openai
openai.api_base = "http://127.0.0.1:65530/api/oai"
openai.api_key = "JUSTSECRET_KEY"The README's sample client sets parameters including max_tokens, top_p, temperature, presence_penalty, frequency_penalty, half_life and stop, and passes half_life through to the completion call. half_life is not part of the OpenAI schema; it is one of the extra optional fields the README mentions. The API schema itself is served at http://localhost:65530/api-docs, which is the place to check when a field is not documented in the README.
Where AI00 RWKV Server stops being the right tool
The model constraint is the first hard edge. The README states that only safetensors models with the .st extension are supported. Anything else has to go through the converter, and the converter is a separate binary or a Python script with its own dependencies. If your workflow already produces .st files, this costs nothing. If it does not, conversion is a mandatory step before the server will load anything.
The second edge is the absence of a rollback story. The README documents installation, compilation, conversion and the API surface, and stops there. There is no documented downgrade procedure, no compatibility matrix between server versions and model versions, and no statement about whether a config file written for one release will load in another. For a component that sits in front of your application's inference calls, that is a gap you have to close yourself by pinning a release and keeping the config under version control.
The third is the API key. The README's own example uses the literal string JUSTSECRET_KEY as the key, which reads as a placeholder rather than a security mechanism. Nothing in the README describes authentication, rate limiting or request isolation. Treat the server as a localhost service. If you expose port 65530 to a network, you are doing so without any documented access control.
Finally, this is not the tool for someone who wants to try many models. It runs RWKV. A user who wants to compare a dozen architectures will spend more time converting weights than evaluating them.
AI00 RWKV Server versus RWKV-Runner and the web-rwkv path
The most direct comparison in the related searches is RWKV-Runner, which is a desktop application for running RWKV models with a graphical interface. The difference is the shape of the deliverable. RWKV-Runner is something a person launches and clicks through. AI00 RWKV Server is a headless process that exposes HTTP endpoints and is meant to be called by another program. If your goal is to chat with a model on your own machine, a desktop app is the shorter path. If your goal is to have an application call a local model over an OpenAI-compatible API, the server is the right shape and the desktop app is not.
The second comparison is web-rwkv itself. AI00 is built on top of it, and the README says so plainly. Using web-rwkv directly means writing your own HTTP layer, your own batching, your own request handling and your own OpenAI-compatible translation. AI00 packages that work. The trade-off is the same one that applies to any wrapper: you inherit its dependency pins, including the hf-hub fork and the local ort patch, and you take its release cadence rather than the engine's. The last push to the repository was on 2026-06-09, the same day v0.7.1 was tagged, so the wrapper is not lagging its engine by any visible margin. The previous release, v0.6.2, was tagged on 2025-10-20, which puts roughly eight months between those two releases. Expect a slow cadence.
Licence position and what upgrading actually costs
The README and the release badge say MIT. The Cargo.toml workspace package field says MIT OR Apache-2.0, and the package.json for the documentation site says MIT. The repository's own LICENSE file is the authority here, and the README does not reproduce its contents. If you need certainty on which terms apply, read that file rather than the badges. Both MIT and Apache-2.0 are permissive and permit commercial use, and the README states that the project is 100% open source and commercially usable under the MIT license. That is a statement about the project's own code. It says nothing about the RWKV model weights you download separately, which carry their own terms on Hugging Face.
Upgrade cost is mostly the model and the config. Moving between server releases means re-checking assets/configs/Config.toml against whatever keys the new version expects, since the README documents the file but does not version it. Moving between model generations is heavier: V5, V6 and V7 checkpoints are separate downloads, and a V6 file is not a V7 file. The converter step is a one-time cost per model, not per server upgrade.
The dependency patches are the part that ages worst. hf-hub is pinned to a git branch rather than a released version. A branch can move. If you build from source for a production deployment, vendor the dependencies or build from a tagged commit rather than from main.
Editorial conclusion
Adopt AI00 RWKV Server if you already run RWKV weights and your GPU is AMD, Intel or integrated, since the Vulkan path avoids CUDA entirely. Do not adopt it if you need a broad model catalogue, a documented rollback path or a support contract; the README covers neither. Before committing, verify three things on your own hardware: that your GPU exposes a Vulkan driver the build can use, that your weights are in .st safetensors form (the converter handles .pth), and that the endpoints your client calls appear under /api/oai/v1/ at http://localhost:65530.
Frequently asked questions
What is an AI server?
In this context it is a process that hosts a language model and answers requests over HTTP. AI00 RWKV Server is one: it loads an RWKV model and exposes chat, completion and embedding endpoints on port 65530 in the OpenAI API format.
What is the best AI server?
There is no single answer, and the README makes no comparison to other servers. AI00 RWKV Server is aimed at a specific case: running RWKV models on Vulkan-capable GPUs, including AMD cards and integrated graphics, without a CUDA or PyTorch runtime.
What do you need for an AI server?
For AI00 RWKV Server the README requires a GPU with Vulkan support and an RWKV model in .st safetensors format placed under assets/models/, with the path set in assets/configs/Config.toml. You run the pre-built binary or build it with cargo, and the WebUI and API listen on port 65530.
Can I run my own AI server?
Yes, on your own hardware. AI00 RWKV Server ships pre-built executables and can also be built from source with cargo build --release, so no hosted service is involved. The README gives http://localhost:65530 as the address for both the WebUI and the API.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/ai00-x-ai00-server)