# rotary-embedding-torch: RoPE for PyTorch Transformers

> A standalone PyTorch library for rotary position embeddings (RoPE), based on the RoFormer paper, with support for axial embeddings, XPos length extrapolation, positional interpolation for context extension, and a fused Flash Attention kernel. It is the go-to package for adding relative positional encoding to custom transformer attention layers with a single pip install.

**lucidrains/rotary-embedding-torch** — Implementation of Rotary Embeddings, from the Roformer paper, in Pytorch

- Repository: https://github.com/lucidrains/rotary-embedding-torch
- Stars: 827 · Forks: 68
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/lucidrains-rotary-embedding-torch

## What RoPE Does and Why Transformers Benefit From It

Standard transformer attention layers need a way to encode the position of each token in a sequence. The original approach from 'Attention Is All You Need' adds a fixed sinusoidal embedding to each token before attention. Rotary position embeddings take a different approach: they encode relative position by rotating the query and key vectors in the attention computation. The rotation is applied in the frequency domain, which has a useful property: the dot product of a rotated query and a rotated key depends only on their relative offset, not their absolute position. This makes RoPE a form of relative positional encoding without the computational overhead of computing pairwise relative positions explicitly. The RoFormer paper (Su et al. 2021) introduced RoPE and showed improved results on language modeling tasks. The technique is now used in many production language models. The rotary-embedding-torch library provides a standalone, importable implementation for PyTorch that covers not just the basic 1D sequential case but also axial embeddings for video and image transformers, length extrapolation, and an optional fused Flash Attention path.

## Installation and Basic Attention Layer Usage

The package is available on PyPI:

```bash
pip install rotary-embedding-torch
```

The library requires einops>=0.8 and torch>=2.4. The basic usage pattern instantiates a RotaryEmbedding object and calls it on queries and keys after splitting attention heads but before the dot product:

```python
import torch
from rotary_embedding_torch import RotaryEmbedding

rotary_emb = RotaryEmbedding(dim = 32)

q = torch.randn(1, 8, 1024, 64)
k = torch.randn(1, 8, 1024, 64)
```

The dim argument controls the size of the frequency bands, not the full head dimension. A common pattern is to set dim to half the head dimension and rotate only that portion, leaving the other half unchanged (partial rotary). The README notes that applying the rotations correctly should produce a noticeable improvement during training.

## Inference Key-Value Cache and Offset Handling

Autoregressive inference typically caches key and value tensors from previous steps and appends new ones at each decoding step. This breaks naive rotary embedding application because the single new query token sits at a different absolute position than its offset into the cached key sequence. The library provides a dedicated method for this case:

```python
q = torch.randn(1, 8, 1, 64)
k = torch.randn(1, 8, 1024, 64)

q, k = rotary_emb.rotate_queries_with_cached_keys(q, k)
```

This method computes the correct offset automatically. The equivalent manual form passes the offset argument directly to rotate_queries_or_keys. Without this correction, a model that works correctly during training will produce degraded output at inference when using KV caching, because the positional signal for the latest query token will be computed at position 0 rather than at its actual sequence offset.

## Axial Embeddings, XPos, Positional Interpolation, and Flash Attention

Beyond the basic sequential case, the library covers four additional use patterns. Axial embeddings extend RoPE to n-dimensional inputs such as video frames, where each position has separate temporal, height, and width coordinates. The RotaryEmbedding class accepts freqs_for='pixel' and a max_freq argument for this mode, and the apply_rotary_emb function handles the resulting frequency tensor. XPos, from the 'A Length-Extrapolatable Transformer' paper, adds an ALiBi-style decay to the rotations to help models generalize to sequence lengths longer than those seen during training. It is enabled with use_xpos=True and works only for autoregressive transformers. Positional interpolation, from the MetaAI paper on context window extension, allows a pretrained model to handle longer sequences by setting interpolate_factor on initialization. The README includes a community report that this method does not work well in practice. The fused Flash Attention kernel combines attention with rotary embedding in a single pass, falling back to a reference PyTorch implementation when Triton is unavailable. It handles non-rotary tokens such as CLS or register tokens by accepting explicit rotary_pos_emb_indices.

## Constraints: XPos Is Causal-Only and Interpolation Has Known Issues

The XPos option cannot be used with bidirectional (encoder-style) transformers. The implementation requires causal masking because the decay is directional. Engineers building BERT-style models or encoders with full attention cannot use use_xpos=True. The positional interpolation feature (interpolate_factor) has an explicitly flagged issue: the README records that a community member reported it does not work well, and the author asks readers to report positive or negative results via email. The pyproject.toml lists the current version as 0.9.1 and the dependency as torch>=2.4. The README links to a potential successor project, PoPE-pytorch, suggesting the author views the design space as not yet settled. The README documentation for some methods is presented through code examples rather than prose, which can leave edge cases undocumented.

## Compared to Learned Absolute Positional Embeddings

The main alternative is the learned absolute positional embedding from the original Transformer paper. A learned embedding assigns each position a trainable vector that is added to the token embedding before the attention layers. This approach is simple to implement and works well within the training distribution, but it does not generalize to sequence lengths longer than those seen during training. RoPE encodes position through rotation rather than addition, which preserves the inner product structure of attention and allows the model to measure relative distances directly from the dot product. The practical consequence is that RoPE-equipped models tend to extrapolate more gracefully to moderately longer sequences without explicit length-extension techniques. The trade-off is that RoPE requires applying the rotation at each attention layer rather than once at the embedding stage, and the inference KV cache offset handling adds implementation complexity that a learned absolute embedding does not require.

## Maintenance History and Licence

The library was first built to implement the RoFormer paper and has since accumulated support for XPos, positional interpolation, axial embeddings, and the fused Flash Attention path. The repository lists releases 0.8.7, 0.8.8, and 0.8.9 from 2025, with the current pyproject.toml showing version 0.9.1. The last push to the main branch was on 2026-06-20. The package is MIT-licensed, which permits use and modification in commercial projects and in open-source derivatives without copyleft restrictions. The build system uses Hatchling with a wheel target that includes the rotary_embedding_torch directory and the README and licence files.

## Conclusion

rotary-embedding-torch is the right choice for engineers building custom transformer architectures in PyTorch who want a maintained, pip-installable implementation of RoPE without writing the rotation logic themselves. The library requires torch>=2.4 and Python>=3.9. Before adopting XPos, note that it works only with autoregressive (causal) transformers, not bidirectional ones. Before adopting the positional interpolation approach, note that the README itself records a community report that it does not work well, and the README author asks for confirmation in either direction. The potential successor project linked in the README is PoPE-pytorch, worth reviewing if this library does not cover your use case. The library is MIT-licensed. The last push was on 2026-06-20.

## FAQ

### How to install rotary_embedding_torch?

Run pip install rotary-embedding-torch. The package requires Python 3.9 or later, torch 2.4 or later, and einops 0.8 or later.

### Does rotary-embedding-torch work with bidirectional transformers?

The core RoPE and axial embedding functionality works with both causal and bidirectional attention. However, the XPos option (use_xpos=True) works only with autoregressive (causal) transformers and cannot be used with bidirectional encoder models.

### Can rotary-embedding-torch extend the context window of a pretrained model?

The library includes a positional interpolation feature controlled by the interpolate_factor parameter. The README notes a community report that this method does not work well in practice, and the author asks readers to report results. XPos is a separate length-extrapolation mechanism that works during training.

## Sources

- [Issues](https://github.com/lucidrains/rotary-embedding-torch/issues)
- [License: MIT](https://github.com/lucidrains/rotary-embedding-torch/blob/main/LICENSE)
- [lucidrains/rotary-embedding-torch on GitHub](https://github.com/lucidrains/rotary-embedding-torch)
- [README](https://github.com/lucidrains/rotary-embedding-torch/blob/main/README.md)
- [Releases](https://github.com/lucidrains/rotary-embedding-torch/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/lucidrains-rotary-embedding-torch
