# xLSTM: the NX-AI repository behind the extended LSTM architecture

> The official xlstm package ships two generations of code in one repository: the NeurIPS sLSTM and mLSTM blocks, and the standalone xLSTM Large 7B model. Here is what installs cleanly, what needs Triton, and where the documentation stops.

**NX-AI/xlstm** — Official repository of the xLSTM.

- Repository: https://github.com/NX-AI/xlstm
- Website: https://www.nx-ai.com/
- Stars: 2,207 · Forks: 186
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/nx-ai-xlstm

## What xlstm actually ships, and who the package is for

The repository is not one model. It holds two code paths that share a name and little else. The first is the NeurIPS paper implementation, built around xLSTMBlockStack for non-language sequence work and xLSTMLMModel for token-based language modeling. The second is xLSTM Large, the optimized architecture behind the 7B parameter model trained on 2.3T tokens, living under xlstm/xlstm_large and described in a separate paper. The README states plainly that the Large implementation has no dependency on the NeurIPS architecture code beyond the shared mlstm_kernels package.

That split matters for anyone deciding whether to adopt it. If you are researching recurrent sequence models for time series or signal data, the block stack is your entry point. If you want a pretrained recurrent language model to run inference against, xLSTM Large is the relevant path and the weights sit on Hugging Face under NX-AI/xLSTM-7b. The two audiences have different install requirements and different hardware constraints.

The package targets Python 3.10 and above per pyproject.toml, and it depends on torch>=2.0, PyYAML, safetensors, packaging and ninja. The README says the package was tested for PyTorch versions >=1.8, which sits awkwardly against the pyproject floor of 2.0. Trust the pyproject file for resolution and treat the README line as historical.

## Exponential gating, matrix memory, and why the kernels decide everything

The architectural claim in the README is that exponential gating with normalization and stabilization techniques, plus a new matrix memory, overcomes the limits of the original LSTM. Those two mechanisms are the whole story. Exponential gating replaces the sigmoid gates of a standard LSTM with an exponentially scaled input gate, and the matrix memory replaces the scalar cell state with a matrix-valued one. The practical consequence is that the recurrent state is larger and the update is a matrix operation, which is why the project ships custom kernels instead of relying on stock PyTorch ops.

The xLSTMLargeConfig object exposes exactly that seam. It takes chunkwise_kernel, sequence_kernel and step_kernel as separate strings, and the README's own example sets them to "chunkwise--triton_xl_chunk", "native_sequence__triton" and "triton". Those names map to TFLA Triton kernels. The same config can be switched to "chunkwise--native_autograd", "native_sequence__native" and "native" to drop Triton entirely.

This is the design decision worth pausing on. Kernel selection is a string in a config dataclass, not a runtime capability check. Nothing in the README suggests the library probes your hardware and falls back on its own. You pick the kernel family, and if you pick wrong the failure surfaces at execution time. The README acknowledges this indirectly by recommending the native PyTorch implementations for platforms other than NVIDIA, and by pointing Apple MLX users to a separate community port, xLSTM-metal, rather than claiming first-party support.

## Installing xlstm with pip or conda, and a first forward pass

The README offers two install routes. The conda route builds a tested environment from environment_pt240cu124.yaml, and the repository also carries environment_pt220cu121.yaml and environment_pt260cu126.yaml for other CUDA and PyTorch combinations. The pip route installs the module alone. Note the ordering in the README: for the xLSTM Large 7B model, mlstm_kernels goes in first.

Create and activate the conda environment from the pinned file:

```bash
conda env create -n xlstm -f environment_pt240cu124.yaml
conda activate xlstm
```

If you only want the module, install mlstm_kernels first and then the package:

```bash
pip install mlstm_kernels
pip install xlstm
```

Cloning and installing in editable mode is the alternative the README gives:

```bash
git clone https://github.com/NX-AI/xlstm.git
cd xlstm
pip install -e .
```

The README's quick start for xLSTM Large is a notebook at notebooks/xlstm_large/demo.ipynb. The core of it builds a small random model and runs a forward pass on CUDA:

```python
import torch
from xlstm.xlstm_large.model import xLSTMLargeConfig, xLSTMLarge

xlstm_config = xLSTMLargeConfig(
    embedding_dim=512,
    num_heads=4,
    num_blocks=6,
    vocab_size=2048,
    return_last_states=True,
    mode="inference",
    chunkwise_kernel="chunkwise--triton_xl_chunk",
    sequence_kernel="native_sequence__triton",
    step_kernel="triton",
)
xlstm = xLSTMLarge(xlstm_config)
xlstm = xlstm.to("cuda")
input = torch.randint(0, 2048, (3, 256)).to("cuda")
out = xlstm(input)
```

What you should see is a tensor whose trailing dimensions are (256, 2048), matching the sequence length and vocabulary size in the config. If the Triton kernels are unavailable, that is where the failure appears. On non-NVIDIA hardware, swap the three kernel strings for the native variants the README lists.

## The sLSTM CUDA path and its compile-time constraints

The older sLSTM code compiles CUDA kernels at install time, and the README is explicit that this requires Compute Capability 8.0 or higher. That rules out older cards outright. If compilation fails, the README suggests pinning the target architectures:

```bash
export TORCH_CUDA_ARCH_LIST="8.0;8.6;9.0"
```

There is a second escape hatch for include-path mismatches, the XLSTM_EXTRA_INCLUDE_PATHS environment variable, which the README shows both as a shell export and as an os.environ assignment inside Python. The README's own framing is that torch and CUDA versions have to match, which is a polite way of saying the build is sensitive to your toolchain. The pyproject build-system requires scikit-build-core>=0.11, pybind11>=3.0 and torch>=2.0, so a source install pulls a compiled extension into your environment rather than staying pure Python.

For faster sLSTM kernels the README points to the separate FlashRNN library rather than offering an in-repo option. That is a real boundary: the fastest path for the sLSTM block is not in this repository.

## Where xlstm is the wrong tool

The clearest limitation is hardware. The README says the model was tested mostly on NVIDIA GPUs and that Triton kernels should also run on AMD, but "should" is doing a lot of work in that sentence. There is no compatibility matrix, no CI badge for AMD, and no statement about which ROCm versions were exercised. For Apple Silicon the README recommends the native PyTorch kernels and points elsewhere for MLX, so if your target is a Mac laptop, this package gives you the slow path by design.

A second limitation is documentation depth. The README covers installation, a forward pass, and kernel selection. It does not document training loops, checkpoint formats, or how to convert the Hugging Face weights for use with the standalone model class. The repository has an experiments/ directory and a tests/ directory, but the README does not describe what either contains, so you are reading source to learn the intended usage.

A third issue is versioning. The jump from v2.0.4 in May 2025 to v2.0.6 in September 2026 spans more than a year with no intermediate releases listed. Nothing in the README explains what changed between them or whether configs written for v2.0.4 still load. If you are pinning this as a dependency, pin the exact version and read the diff yourself.

Finally, if your task is a small tabular or short-sequence problem, the matrix memory and custom kernels buy you nothing. A standard LSTM or a gradient-boosted model will train faster and be easier to debug.

## xLSTM against Mamba and RWKV, and against the plain LSTM

The comparison people reach for is against state space models like Mamba and against RWKV. All three are attempts to escape the quadratic cost of attention while keeping competitive language modeling quality, and all three end up shipping custom CUDA kernels to make the recurrent formulation fast. The difference is in the state. Mamba carries a selective state space with input-dependent parameters; xLSTM carries a matrix-valued memory cell updated through exponential gating. The README's own framing is that xLSTM shows promising performance on language modeling when compared to Transformers or State Space Models, which is a claim about the paper rather than a benchmark table in the repository.

Against the original LSTM the difference is more concrete and easier to state. The README says the architecture overcomes the limitations of the original LSTM through exponential gating with normalization and stabilization, and through the matrix memory. A standard LSTM has a scalar cell state per unit and sigmoid gates; xLSTM has a matrix memory and exponential gates. That is a larger state per layer, which is why the kernels exist and why the Compute Capability floor is where it is.

The honest position is that the repository does not contain a head-to-head benchmark against Mamba or RWKV. If that comparison is what you need before adopting, you will have to run it yourself on your own workload.

## Licence, maintenance and what upgrading costs

The repository is Apache-2.0, and pyproject.toml declares license = {file="LICENSE"}. That is permissive and unsurprising for a research codebase. The complication is the 7B model. The README's badge for xLSTM 7B links to a license labelled nxai_community rather than Apache-2.0, and the badge URL points at a path inside a repository named tirex-internal. So the code and the weights are not obviously under the same terms. If you intend to use the 7B weights commercially, read the licence file attached to the Hugging Face model page rather than assuming the repository licence covers it. This is not legal advice; it is a pointer to the two different licence strings the README itself displays.

On maintenance, the last push to the default branch was on 2026-09-07, the same day v2.0.6 was tagged. The repository is not archived. The gap between v2.0.4 and v2.0.6 is the main cost signal: if you adopt, expect to verify upgrades manually rather than following a changelog. The uv.lock file and the three pinned environment_pt*.yaml files give you reproducible starting points, but they are snapshots, and the README does not describe a migration path between them. Budget for reading the diff between tags before you move a pinned version.

## Conclusion

Adopt xlstm if you are evaluating recurrent alternatives to attention and you have an NVIDIA GPU with Compute Capability 8.0 or higher for the sLSTM CUDA path, or you want to run the xLSTM Large 7B weights from Hugging Face. Do not adopt it if you need CPU-only training, Apple Metal as a first-class target, or a documented rollback and versioning policy, because the README covers none of those. Before committing, check that your installed torch version matches one of the three environment_pt*.yaml files, and confirm the mlstm_kernels package resolves on your Python version, since pyproject.toml only pulls it in for Python 3.11 and above.

## FAQ

### What is a LSTM in simple terms?

The README describes xLSTM as a recurrent neural network architecture based on ideas of the original LSTM, which uses gating to control what a recurrent cell keeps and discards. It does not give a general tutorial definition of an LSTM beyond that.

### Is LSTM still relevant?

The README's position is that xLSTM overcomes the limitations of the original LSTM through exponential gating with normalization and stabilization plus a new matrix memory, and shows promising performance on language modeling compared to Transformers or State Space Models. That is the project's claim about the architecture, not a statement about older LSTM models.

### Is LSTM a CNN or RNN?

The repository lists rnn among its topics and describes xLSTM as a recurrent neural network architecture, so the project places itself in the recurrent family rather than the convolutional one.

### What is xlstm?

It is a recurrent neural network architecture based on the original LSTM, using exponential gating with normalization and stabilization plus a new matrix memory. The NX-AI repository is the official implementation and also contains the optimized xLSTM Large architecture behind the 7B model.

### How does xLSTM compare to Mamba?

Both aim to avoid attention's cost while keeping language modeling quality, but the README does not present a head-to-head benchmark against Mamba. It only states that xLSTM shows promising performance compared to Transformers or State Space Models, citing the paper.

### How does xLSTM compare to RWKV?

The repository does not contain a comparison against RWKV. The README's performance claim is limited to Transformers and State Space Models generally, and evaluation against RWKV would have to be run on your own workload.

## Sources

- [License: Apache-2.0](https://github.com/NX-AI/xlstm/blob/main/LICENSE)
- [NX-AI/xlstm on GitHub](https://github.com/NX-AI/xlstm)
- [Project website](https://www.nx-ai.com/)
- [README](https://github.com/NX-AI/xlstm/blob/main/README.md)
- [Releases](https://github.com/NX-AI/xlstm/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/nx-ai-xlstm
