# RWKV-LM: Training an Attention-Free RNN at Transformer Scale

> RWKV-LM is the reference training repository for the RWKV architecture, an RNN that runs in linear time with constant memory and trains with full parallelism. At RWKV-7, the project is a Linux Foundation AI initiative with production inference already integrated into Windows and Office.

**BlinkDL/RWKV-LM** — RWKV (pronounced RwaKuv) is an RNN with great LLM performance, which can also be directly trained like a GPT transformer (parallelizable). We are at RWKV-7 "Goose". So it's combining the best of RNN and transformer - great performance, linear time, constant space (no kv-cache), fast training, infinite ctx_len, and free sentence embedding.

- Repository: https://github.com/BlinkDL/RWKV-LM
- Stars: 14,710 · Forks: 1,018
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/blinkdl-rwkv-lm

## What RWKV solves and who it is for

Standard transformer-based LLMs store a key-value cache that grows linearly with context length, making very long sequences expensive in memory and increasingly slow to decode. RWKV replaces attention with a recurrent state that stays constant in size no matter how many tokens have been processed. According to the repository, RWKV-7 at 1.6 million state parameters compares to a similarly sized Qwen 3.5 model whose state grows from 5 million parameters and then adds another 6 million for every 1,000 tokens of context.

The target audience is researchers and engineers who want to train or fine-tune large language models on hardware that cannot hold a growing KV cache, or who need predictable, constant-speed decoding throughput in production. RWKV is also a Linux Foundation AI project, and the README notes that the runtime is already shipped inside Windows and Microsoft Office.

## Architecture: linear time with a meta-in-context state

RWKV is a pure RNN whose four core parameters give it its name: R, W, K, and V. Unlike earlier RNNs, it can be trained with full parallelism across sequence positions, matching transformer training throughput. The README describes RWKV-7 as a "meta-in-context learner" that performs test-time training of its recurrent state via in-context gradient descent at every token, rather than attending over past tokens.

At inference time, RWKV-7 runs in RNN mode: each new token updates a fixed-size state and is processed in constant time and space. At training time, the same model runs in GPT-like parallel mode over the full sequence. The repository documents a benchmark showing RWKV-7 7.2B trained on 4x8 H100 GPUs at context length 25,600 with DeepSpeed zero2 reaching 277,000 tokens per second and 38 percent MFU.

The repository also documents the state size calculation explicitly:

```
RWKV-7 L24-D1024 #state_params = L*(2*D+64*D) = 1.622016 M
Qwen3.5 L24-D1024 #state_params = L*3/4*(3*6*D+2*128*D) + L/4*(2*2*256*T)
  = 5.050368 + 6.144*(T/1000) M
```

This shows the structural difference: RWKV's state is a fixed number regardless of sequence length T, while attention-based alternatives grow with T.

## Running a first RWKV-7 training job

The README points to the RWKV-v7/train_temp directory as the official reference implementation for RWKV-7. The default configuration requires one GPU with 7 GB of VRAM, and batch size can be reduced further for smaller cards. The training demo file is rwkv7_train_simplified.py.

The README also provides demo scripts for running an already-trained model. The GPT-like (parallel) mode, the RNN (token-by-token) mode, and a combined fastest mode each have their own file:

```bash
# GPT-like parallel mode
python RWKV-v7/rwkv_v7_demo.py

# RNN step-by-step mode
python RWKV-v7/rwkv_v7_demo_rnn.py

# Fastest mode (both)
python RWKV-v7/rwkv_v7_demo_fast.py
```

Pre-trained weights are hosted on Hugging Face at the BlinkDL account, and GGUF-quantized versions are available from separate community collections. The README also lists the official rwkv pip package at pypi.org/project/rwkv for inference without cloning this repository.

The README calls out three training requirements that are easy to get wrong. First, use PreLN LayerNorm rather than RMSNorm. Second, apply weight decay only to large matrix parameters such as projection weights, not to all parameters. Third, use the initialization scheme from the reference code. The README states that getting these wrong leads to training instability.

## Limitations and where RWKV is the wrong tool

The biggest practical constraint is that RWKV-7 cannot be treated as a modular layer. The README explicitly states: "there is no good simple 'RWKV-7 layer' because a pytorch layer can't make sure itself is using correct init and hyperparameters." Using RWKV-7 for a new task requires reading several hundred lines of train_temp code and adapting it carefully, not importing a package.

The README also flags that the FLA (Flash Linear Attention) implementation of RWKV-7 is not yet aligned with the reference implementation and performs noticeably worse. Anyone relying on third-party framework integrations should check current alignment status before training.

Another limit involves the recurrent state itself. Because the state is fixed-size and summarizes the entire context, RWKV cannot retrieve specific tokens from far back in the context the way an attention model can. For tasks that require precise long-range recall of exact token positions (such as needle-in-a-haystack retrieval across very long documents), the README provides a needle-in-a-haystack evaluation image but does not quantify failure modes, so teams should run their own evals before committing.

The repository has releases only up to RWKV-v5 from December 2023 on GitHub. RWKV-7 and RWKV-8 work is tracked on the main branch, not through GitHub releases.

## Ecosystem: inference, fine-tuning, and community tools

The README documents the Albatross project as the recommended efficient inference backend. Albatross benchmarks on an RTX 5090 show 145 tokens per second at batch size 1 for the 7.2B fp16 model, and over 10,000 tokens per second at batch size 960, all at constant memory. The runtime is available from a separate repository rather than from this one.

For fine-tuning, the README points to RWKV-PEFT for LoRA and similar methods, and to RWKV-LM-RLHF for reinforcement learning from human feedback. Mobile inference is handled by a separate rwkv-mobile library.

The GUI entry point for non-researchers is RWKV-Runner, which provides a desktop application with releases. The Ai00 Server project offers an API server for deployment.

Compared to Mamba (another linear-recurrence architecture), RWKV differs in that it uses a gating structure based on the R, W, K, V formulation rather than selective state spaces. Both avoid the quadratic attention cost, but RWKV-7 adds the in-context gradient descent mechanism described as meta-in-context learning, which Mamba does not have.

## Licence and maintenance

The code is released under Apache-2.0, which permits use in commercial products without requiring source disclosure. The last push to the repository was on 2026-09-03, meaning the project sees regular commits.

The RWKV-8 direction is tracked in the RWKV-8.md file and a related ROSA document in the repository root. The changelog on the wiki covers the evolution from v1 through v7, though the README notes that the wiki is AI-written and may contain errors.

The repository carries a CITATION.cff file, which means citation tooling in academic environments can pick up a structured reference automatically.

## Conclusion

RWKV-LM is worth evaluating if you need an LLM that runs at constant memory regardless of context length, or if you want to train a model on a single GPU with 7 GB of VRAM. It is the wrong choice if you want a drop-in replacement for a standard transformer layer: the README specifically warns that there is no well-isolated RWKV-7 layer because correct initialization and per-parameter weight decay settings must be preserved for the architecture to scale stably. Before adopting it, read the train_temp code and verify that your use case can tolerate the integration work required to port RWKV-7 to a new task.

## FAQ

### What is RWKV used for?

RWKV is used for training and running large language models that require constant memory regardless of context length. The README lists LLM and multimodal applications as primary targets, and the runtime is already integrated into Windows and Microsoft Office.

### How does an LLM work internally?

The README explains that RWKV-7 replaces the attention mechanism with a recurrent state that is updated at each token via in-context gradient descent. During training the model runs in parallel across sequence positions like a transformer; at inference time it runs as a pure RNN, processing one token at a time with a fixed-size state.

### What is the minimum GPU requirement to train RWKV-7?

According to the README, the default train_temp configuration requires one GPU with 7 GB of VRAM, and the batch size can be reduced further for cards with less memory.

## Sources

- [BlinkDL/RWKV-LM on GitHub](https://github.com/BlinkDL/RWKV-LM)
- [Issues](https://github.com/BlinkDL/RWKV-LM/issues)
- [License: Apache-2.0](https://github.com/BlinkDL/RWKV-LM/blob/main/LICENSE)
- [README](https://github.com/BlinkDL/RWKV-LM/blob/main/README.md)
- [Releases](https://github.com/BlinkDL/RWKV-LM/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/blinkdl-rwkv-lm
