titans-pytorch: Unofficial PyTorch Implementation of Titans Memory Architecture
Unofficial implementation of Titans, SOTA memory for transformers, in Pytorch
At a glance
- What is it?
- titans-pytorch is Phil Wang's (lucidrains) unofficial PyTorch implementation of the Titans architecture from Google DeepMind, which introduces a learnable neural memory module that updates at test time. The library provides the NeuralMemory module, a MemoryAsContextTransformer, and includes experiment scripts. It requires PyTorch 2.8 or newer.
- Who is it for?
- titans-pytorch is appropriate for researchers experimenting with memory-augmented transformer architectures who want a working PyTorch implementation of the Titans paper without building the neural memory module from scratch.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 79 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Titans Is and the Problem It Addresses
Standard transformers use self-attention with fixed context windows. Extending context is expensive: attention cost scales quadratically with sequence length. Recurrent models like LSTMs and Mamba handle long sequences with bounded memory, but they compress past information into a fixed-size state with no mechanism to decide what is worth remembering.
The Titans paper from 2024 proposes a different approach: a learnable neural memory module that updates its own weights at test time based on the input sequence. The idea is that the memory module learns to memorize relevant information by gradient descent during inference, not just during training. Titans introduces three variants for combining this memory with attention: MAC (Memory as Context), MAG (Memory as Gated), and MAL (Memory as Layer).
lucidrains' implementation focuses on the MAC configuration, where the neural memory's retrieved content is concatenated with the current input tokens before local attention, giving the model access to both recent context and the long-term memory retrieval. The README notes this is an unofficial implementation and that it may also explore architectures beyond the simple 1 to 4 layer MLP that the original paper uses for the memory module.
Installation and Basic Usage
Install from PyPI:
$ pip install titans-pytorchThe minimal NeuralMemory usage:
import torch
from titans_pytorch import NeuralMemory
mem = NeuralMemory(
dim = 384,
chunk_size = 64
).cuda()
seq = torch.randn(2, 1024, 384).cuda()
retrieved, mem_state = mem(seq)The `chunk_size` parameter controls the granularity of memory updates. Smaller values give better performance on short sequences but use more memory. The MAC transformer configuration:
from titans_pytorch import MemoryAsContextTransformer
transformer = MemoryAsContextTransformer(
num_tokens = 256,
dim = 256,
depth = 2,
segment_len = 128,
num_persist_mem_tokens = 4,
num_longterm_mem_tokens = 16,
)The `segment_len` parameter sets the local attention window, `num_persist_mem_tokens` controls persistent memory tokens, and `num_longterm_mem_tokens` sets how many tokens from the neural memory are prepended to each segment.
Repository Layout and Experiment Scripts
The repository follows lucidrains' typical structure. The `titans_pytorch/` directory contains the library code. The `tests/` directory contains the test suite. Experiment data lives in `data/` and training scripts are at the top level.
The primary experiment script is `train_mac.py`, which trains a MAC transformer. The recommended way to run experiments is:
$ pip install uvThen modify `train_mac.py` and run it:
$ uv run train_mac.pyA Colab notebook is available for a quick test run without local setup. The `fig1.png` and `fig2.png` files in the repository root contain architecture diagrams from the paper.
The library also includes `train_implicit_mlp_attn.py` for a different architectural variant. The pyproject.toml lists wandb as an optional dependency in the `examples` group, suggesting that experiment tracking is supported but not required.
Key Technical Details: Memory Update and MAC Architecture
The neural memory module in Titans learns to memorize by treating its own weights as a key-value store. During a forward pass on a sequence chunk, the module computes a surprise-weighted gradient update to its MLP weights, allowing it to incorporate the current input. This is described in the paper as 'learning to memorize at test time.'
In the MAC configuration, the retrieved memory content is prepended to the local attention window as context tokens. The `num_persist_mem_tokens` parameter adds tokens that persist across all segments without being updated by the memory mechanism, analogous to learned positional or task tokens. The `num_longterm_mem_tokens` parameter controls how many memory retrievals are added per segment.
The `chunk_size` passed to `NeuralMemory` determines how many tokens are processed together in each memory update step. The README notes that smaller chunk sizes improve performance on shorter sequences but increase memory usage, since more update steps run within a single forward pass.
The library's citations in the README acknowledge several related works including TTT (Test-Time Training), Gated Delta Networks, and fast-weight product key memory, indicating that the implementation draws from multiple concurrent research directions on recurrent-style memory for transformers.
Dependencies and Hardware Requirements
The package requires Python 3.9 or newer and PyTorch 2.8 or newer. This is a recent version requirement: PyTorch 2.8 was released in mid-2025, and some environments pinned to earlier versions will need to upgrade before installing titans-pytorch.
The dependency list in pyproject.toml includes several lucidrains libraries: `x-transformers`, `rotary-embedding-torch`, `axial_positional_embedding`, `hyper-connections`, `einops` and `einx`. The `tensordict` library is also required. The optional examples group adds `adam-atan2-pytorch` and wandb.
Running `train_mac.py` requires a CUDA GPU. The usage example calls `.cuda()` explicitly. The library's performance on CPU or MPS (Apple Silicon) is not documented. The Ninja build tool is listed as a required dependency, which suggests that some C extensions may be compiled at import time or when specific attention kernels are triggered.
The `assoc-scan>=0.0.4` dependency provides an associative scan primitive, which is needed for the parallel recurrence computation in the memory update step.
Limitations and Comparison with Mamba
Titans is an architectural research experiment, not a production-ready building block. The implementation is unofficial and the README explicitly says it may explore beyond the original paper. The most recent release is 0.5.3 from 2026-02-09; the dev branch has received updates since then but no new release has been tagged. This means users who install from PyPI get version 0.5.3, while the source on GitHub may be at 0.5.5 per pyproject.toml.
The requirement for PyTorch 2.8 is a real constraint for projects that need to maintain compatibility with older environments or that deploy on platforms where the latest PyTorch version is not available. Some cloud providers or containerized environments default to earlier PyTorch versions, and upgrading in those contexts requires testing the full pipeline.
Mamba (the Mamba state space model architecture) is a direct comparison point for long-context sequence modeling. Mamba achieves sub-quadratic scaling through a selective state space mechanism, while Titans achieves it through explicit neural memory with gradient-based updates at test time. Mamba has more mature implementations, including the official mamba-ssm repository and integrations in Hugging Face Transformers, whereas Titans is newer and less integrated into standard toolchains. Teams looking for a production-ready long-context architecture would currently find more infrastructure support around Mamba than around Titans. The titans-pytorch library itself also has no benchmarks or evaluation results in the README, so comparing effectiveness to Mamba requires running the experiment scripts independently.
Maintenance Status and Licence
The repository is MIT licensed and maintained by Phil Wang (lucidrains). Version 0.5.5 is shown in pyproject.toml. The most recent GitHub release is 0.5.3, tagged on 2026-02-09, followed by 0.5.1 on 2026-01-27 and 0.5.0 on 2026-01-07. The last push was on 2026-07-13.
lucidrains maintains a large collection of ML research implementations. The repository style is consistent with his other projects: clean Python, heavy use of einops for tensor manipulation, and citation blocks in the README for all referenced works. The citations cover the original Titans paper, the TTT paper, Gated Delta Networks, min-p sampling, Hyper-Connections, Value Residual Learning, Accelerated Scan, test-time regression, ATLAS and fast-weight product key memory.
The package is published to PyPI at pypi.org/project/titans-pytorch/. New versions are released through PyPI rather than only through GitHub releases.
Editorial conclusion
titans-pytorch is appropriate for researchers experimenting with memory-augmented transformer architectures who want a working PyTorch implementation of the Titans paper without building the neural memory module from scratch. The implementation is unofficial and exploratory: the README notes that lucidrains may extend it beyond the original paper's architecture, and the library requires PyTorch 2.8 or newer, which is a recent version requirement that may not align with existing project environments. The most recent release is 0.5.3 from 2026-02-09. The last push to the repository was on 2026-07-13, more than two months before today's date, and no release has been tagged since then. Before building on this library for a longer project, verify that the specific configuration and chunk size you need produce correct gradients with `train_mac.py`.
Frequently asked questions
What is the Titans AI model?
Titans is a transformer architecture from a 2024 Google DeepMind paper that introduces a learnable neural memory module. The memory module updates its own weights at test time based on the input sequence, allowing the model to memorize relevant information during inference. titans-pytorch is an unofficial PyTorch implementation of this architecture.
How do I install titans-pytorch?
Run `pip install titans-pytorch`. The library requires Python 3.9 or newer and PyTorch 2.8 or newer. For running the included experiment scripts, `pip install uv` is recommended, then run `uv run train_mac.py` to train the MAC configuration.
How is Titans different from Mamba for long-context modeling?
Mamba uses a selective state space mechanism to achieve sub-quadratic scaling. Titans uses a neural memory module that performs gradient-based weight updates at test time to store and retrieve information over long horizons. Mamba has more mature toolchain integration; Titans is a newer research direction with fewer production deployments documented.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/lucidrains-titans-pytorch)