Library / SDK
BlinkDL/ChatRWKV avatar
BlinkDL/ChatRWKV

ChatRWKV: Running an RNN Language Model as a Local Chatbot

ChatRWKV is like ChatGPT but powered by RWKV (100% RNN) language model, and open source.

9,497 stars685 forksPythonApache-2.0

At a glance

What is it?
ChatRWKV is the chat-facing part of the RWKV project: a Python client for a 100% RNN language model that keeps conversation state instead of re-reading the whole context. It is a developer toolkit, not a packaged app, and the README is written for people who already know their way around PyTorch checkpoints.
Who is it for?
Adopt ChatRWKV if you want to build a chatbot on top of an RWKV checkpoint and you are comfortable managing state objects, model strategies and tokenizer files yourself. Do not adopt it if you want a finished desktop or mobile chat app: the README points to RWKV-Runner and RWKV_APP for that, and to ai00_rwkv_server or rwkv.cpp when you need a serving layer or CPU inference.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 72 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What ChatRWKV actually is, and who it is not for

ChatRWKV is the chat application layer for RWKV, a language model architecture that is a pure RNN rather than a transformer. The README describes it as "like ChatGPT but powered by my RWKV (100% RNN) language model", and the repository is the Python code that turns an RWKV checkpoint into something you can talk to. It is not a product. There is no installer, no desktop binary and no hosted service in this repository. What you get is chat.py, a set of API demo scripts, a small src/ package, a tokenizer directory, and a v2/ directory holding the strategy-based implementation with streaming, splitting and INT8 support.

The intended audience is developers who already work with PyTorch checkpoints and want a conversational front end they can modify. The README says outright that if you are building your own inference engine you should start with src/model_run.py, which it calls easier to understand and which chat.py itself uses. That is a strong signal about the project's self-image: it is a reference implementation you are expected to read, not a black box you are expected to configure.

If you want a graphical chat application, this is the wrong repository. The README lists RWKV-Runner as a community GUI, RWKV_APP for local inference on Android and iOS, and ai00_rwkv_server as the fastest GPU inference API with Vulkan. Those are separate projects with separate maintenance. ChatRWKV is where the model semantics are defined; the other repositories are where the packaging lives.

The RNN state model that shapes every design decision

The mechanism that separates RWKV from a transformer chatbot is state. A transformer re-processes the entire conversation on every turn. An RNN carries a fixed-size state forward, and the README's example makes this concrete: feeding tokens [187, 510, 1563, 310, 247] in one call and then feeding [187, 510], [1563], [310, 247] across successive calls with the returned state produces the same logits. The model does not need to see the earlier tokens again.

That property is why the README claims RWKV is faster and saves VRAM, and why the v2 notes say 3G VRAM is enough to run a 14B model. It is also the source of the project's main operational hazard. The state is an opaque tensor, and if it drifts out of sync with the text it represents, the model produces nonsense that looks like a bug in the model rather than a bug in your bookkeeping. The README addresses this directly with an instruction in bold: never call raw forward() directly, and instead wrap it in a function that records the text corresponding to the state.

The chat format is part of the same contract. The README specifies Bob/Alice for v4-raven models and User/Assistant for v4, v5 and v6 world models, with the pattern ending in a bare "Alice:" and no trailing space. It also warns against double newlines inside a turn and suggests normalizing with a strip and replace chain. These are not stylistic preferences. They are the format the checkpoints were trained on, and deviating from it degrades output in ways that are hard to diagnose.

Installing the rwkv package and running a first forward pass

The repository's requirements.txt lists only tokenizers>=0.13.2 and prompt_toolkit. The model itself comes from the separate rwkv package on PyPI, which the README links and instructs you to keep upgraded. There is no pinned version in the README, so treat the installed version as something to check rather than assume.

bash
pip install rwkv
pip install -r requirements.txt

With the package installed, the README's minimal example sets two environment variables before importing. RWKV_JIT_ON controls JIT compilation, and RWKV_CUDA_ON controls a CUDA kernel that the README says is much faster and saves VRAM. The strategy string is where you declare device and precision; the example uses 'cuda fp16'.

python
import os
os.environ["RWKV_JIT_ON"] = '1'
os.environ["RWKV_CUDA_ON"] = '0'
from rwkv.model import RWKV
model = RWKV(model='/path/to/RWKV-4-Pile-1B5-20220903-8040', strategy='cuda fp16')
out, state = model.forward([187, 510, 1563, 310, 247], None)

The model path points at a checkpoint you supply; the README uses a local filesystem path in its example. What you should see is a logits tensor from out, and a state object you must carry into the next call. For interactive chat, the README points at chat.py as the entry point, and at v2/chat.py when you want the CUDA kernel build, which requires ninja and, on Linux, CUDA paths exported into PATH and LD_LIBRARY_PATH.

bash
pip install ninja
export PATH=/usr/local/cuda/bin:$PATH
export LD_LIBRARY_PATH=/usr/local/cuda/lib64:$LD_LIBRARY_PATH
python v2/chat.py

The README also recommends converting a model for a chosen strategy with v2/convert_model.py, which it says gives faster loading and saves CPU RAM. That step is optional but is presented as the normal path for repeated runs.

Where ChatRWKV breaks, and the cost of getting state wrong

The failure mode the README spends the most words on is state and text falling out of alignment. Because the state is a tensor and the text is a string, nothing in the API enforces that they correspond. A copied state, a truncated history, or a message appended without updating the state will produce a chatbot that answers coherently for a few turns and then degrades. The README's advice to always check the text corresponding to the state is the only guardrail offered.

The chat format is the second failure surface. Models trained with Bob/Alice will not behave the same way when fed User/Assistant, and the README is explicit about which family uses which. The final "Alice:" with no trailing space, and the ban on double newlines, are easy to violate when you build prompts from user input that contains blank lines. The README's replace chain for normalizing newlines is a workaround, not a validation layer.

The third limitation is scope. This repository does not include a training pipeline; the README directs you to RWKV-LM for explanation, fine-tuning and training. It does not include a serving API; ai00_rwkv_server and rwkv.cpp cover that. And it does not include quantized inference for CPU beyond what community projects provide. If your deployment target is a phone, a CPU-only server, or a service with an HTTP contract, ChatRWKV is the wrong layer to build on.

How ChatRWKV differs from llama.cpp and other GGUF-based chat stacks

The obvious comparison is with GGUF-based runners such as llama.cpp, and the difference is architectural rather than cosmetic. A transformer runner must manage a KV cache that grows with conversation length, so memory use and per-turn latency climb as the chat gets longer. An RWKV model keeps a constant-size state, so the cost per turn is flat. That is the trade-off the README is advertising when it says RWKV is faster and saves VRAM, and it is why the v2 notes can claim a 14B model fits in 3G VRAM.

The cost of that trade is ecosystem. GGUF tooling is built around a file format with wide loader support, and the README itself links to a GGUF collection for RWKV-7 weights, plus rwkv.cpp for int4/int8/fp16/fp32 CPU inference via ggml. So the honest framing is not ChatRWKV versus llama.cpp. It is ChatRWKV as the PyTorch reference path, and rwkv.cpp or ai00_rwkv_server as the deployment paths for the same model family. If you want to understand or modify how RWKV chat works, read this repository. If you want to ship inference, the README sends you elsewhere.

Maintenance, licensing and what upgrading costs you

The repository is not archived, and the last push was on 2026-07-19. That is recent enough that the code is being touched, but the README offers no release notes, no changelog and no versioned releases in the retrieved material, so there is no documented upgrade path. The README's own instruction is to check for the latest version of the rwkv pip package and upgrade, which means the model-loading code and the package can drift apart without a compatibility statement.

There is a real upgrade cost hidden in the model families. The README distinguishes v4-raven from v4/v5/v6 world models by chat format, and points to RWKV-7 as the latest version with preview weights. Moving between these is not a version bump; it is a prompt-format change plus a different checkpoint plus, potentially, a different tokenizer. The repository ships 20B_tokenizer.json at the top level and a tokenizer/ directory, and the README's forward example notes to use 20B_tokenizer.json, so tokenizer selection is part of the migration.

The licence is Apache-2.0, which is permissive and includes an explicit patent grant. That covers the code in this repository. It does not automatically settle the terms attached to individual model weights, which are hosted on Hugging Face under separate repositories. Check the licence stated on the specific checkpoint you download; the repository licence and the weights licence are different documents.

Editorial conclusion

Adopt ChatRWKV if you want to build a chatbot on top of an RWKV checkpoint and you are comfortable managing state objects, model strategies and tokenizer files yourself. Do not adopt it if you want a finished desktop or mobile chat app: the README points to RWKV-Runner and RWKV_APP for that, and to ai00_rwkv_server or rwkv.cpp when you need a serving layer or CPU inference. Before committing, verify the exact checkpoint you intend to use, run v2/convert_model.py against it, and confirm that the state-to-text bookkeeping described in the README matches the format your chosen model family expects.

Frequently asked questions

How do I install ChatRWKV and run it for the first time?

Install the rwkv pip package and the repository's requirements.txt, which lists tokenizers and prompt_toolkit. Then run chat.py for interactive chat, or v2/chat.py if you want the CUDA kernel path, after setting RWKV_JIT_ON and RWKV_CUDA_ON and supplying a checkpoint path.

Which chat format should I use with ChatRWKV?

The README specifies Bob/Alice for v4-raven models and User/Assistant for v4, v5 and v6 world models. The prompt must end with a bare "Alice:" or "Assistant:" with no trailing space, and you should avoid double newlines inside a turn.

Why does my ChatRWKV conversation degrade after a few turns?

The README warns that the model state and the text it represents must stay in sync, and that you should never call raw forward() directly. Wrap it in a function that records the text corresponding to the state, and check that text when output looks wrong.

Official sources

  1. BlinkDL/ChatRWKV on GitHub
  2. Issues
  3. License: Apache-2.0
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/blinkdl-chatrwkv.svg)](https://hysenlabs.com/projects/blinkdl-chatrwkv)