# EverMind-AI/MSA: Memory Sparse Attention for 100M-Token Contexts

> MSA is an end-to-end trainable latent-memory framework that replaces full attention with sparse document retrieval during decoding. The README claims near-linear complexity and under 9% degradation from 16K to 100M tokens, but the repository ships research code and a paper, not a packaged inference server.

**EverMind-AI/MSA** — Memory Sparse Attention -  A scalable, end-to-end trainable latent-memory framework for 100M-token contexts.

- Repository: https://github.com/EverMind-AI/MSA
- Website: https://arxiv.org/abs/2603.23516
- Stars: 3,515 · Forks: 228
- Language: Python
- License: not declared
- Published: 2026-09-23 · Updated: 2026-09-23 · Language: en
- Canonical page: https://hysenlabs.com/projects/evermind-ai-msa

## The context-length wall MSA is aimed at

Full attention scales quadratically with sequence length, which is why the README states that most LLMs are effectively capped between 128K and 1M tokens. MSA targets that ceiling with a different structure: instead of attending over every token, it stores a compressed latent memory of the corpus and retrieves a small number of documents during decoding. The README frames the alternatives it is reacting to: hybrid linear attention, fixed-size state memory such as RNNs, and external retrieval stacks like RAG or agents. Each, in the project's account, either decays in precision at extreme scale, lacks end-to-end differentiability, or requires a complicated pipeline. MSA's pitch is that retrieval and generation sit in one differentiable loop, so the router can be trained rather than bolted on. The intended audience is narrow. This is a research framework for teams studying long-context memory, not an application library. The repository layout confirms that: src/, scripts/, paper/, and a QUICK_START.md, with no packaging metadata at the top level.

## How MSA compresses and routes documents

The mechanism described in the README has three parts. First, document latent states, written as K/V/Kr, are chunk-mean pooled into compressed vectors. Second, a router projector scores relevance by cosine similarity, mean-pooled across heads and then maxed over tokens, and picks the Top-k documents. Third, the compressed K/V of those documents are concatenated with the query's local K/V for autoregressive decoding. Routing only happens in the upper layers; lower layers process documents independently so that hierarchical alignment is preserved.

The positional scheme is the part that makes the scaling claim possible. With parallel document-wise RoPE, each document resets its positions to zero, so training at 64k does not drift when inference runs at 100M. Global RoPE handles the active context by offsetting the query's starting index by k, the number of retrieved blocks, keeping the order background, then query, then generation. The inference pipeline has three stages: offline global memory encoding that caches the pooled triples, online routing that projects the query to Qr and matches it against Kr, and sparse generation over the assembled context. Memory Parallel shards Kr across GPUs, with query broadcast, local scoring and a global reduce, while content K/V stays in host DRAM and is fetched asynchronously when selected. That split between VRAM and host memory is what the README says makes 100M-token deployment feasible on 2xA800 GPUs.

## Installing MSA and running the quick start

The repository pins its dependencies in requirements.txt at the top level, including torch==2.6, transformers==4.51.3, liger_kernel==0.5.10 and accelerate==1.0.1. Install them into a virtual environment before touching src/. The README does not document a pip package, so treat this as a source checkout.

```bash
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
```

The remaining entry points are QUICK_START.md and the scripts/ directory. The README does not reproduce the quick start steps, so open QUICK_START.md for the actual commands and check what scripts/ expects before running anything.

```bash
ls scripts/
cat QUICK_START.md
```

There is also a Hugging Face model referenced by the README badge as MSA-4B at EverMind-AI/MSA-4B. Note that the requirements include openai, which suggests the evaluation path calls an external API for LLM judging rather than running everything locally. The README does not state which scripts need that key or where it is read from.

## What the evaluation numbers do and do not show

The README reports MSA at an average of 3.760 on nine QA datasets against same-backbone RAG, RAG with reranking, and HippoRAG2, with improvements of 16.0%, 11.5% and 14.8% respectively. The backbone is Qwen3-4B-Instruct-2507, memory banks range from 277K to 10M tokens, and the metric is an LLM judge on a 0 to 5 scale. NIAH results come from RULER, eight subtasks between 32K and 1M tokens.

Two things deserve scrutiny. The judge is an LLM, not an exact-match metric, so the absolute scores are not comparable to F1 or accuracy numbers from other papers. And MSA does not lead everywhere: on NarrativeQA the table shows RAG with reranking at 3.638 against MSA at 3.395, a case the README acknowledges by saying MSA leads on all but NarrativeQA within the same-backbone group. The scaling claim itself, under 9% degradation from 16K to 100M tokens, is stated for MS MARCO specifically, and the README notes that some baseline curves end early because of their own context limits. That is a fair caveat, but it means the comparison at the far end of the range is thinner than the headline suggests.

## The hardware floor and the missing licence file

The 100M-token throughput claim is tied to 2xA800 GPUs in the README. There is no documented single-GPU path, no CPU fallback, and no statement about how much host DRAM the content K/V tier needs at that scale. If you do not have that class of hardware, the headline capability is not something you can evaluate locally; you can still read src/ and paper/, but the pipeline described in the README assumes a multi-GPU setup with a host-memory tier.

The licence situation is worse. The README badge links to MIT, but the repository has no licence file among its top-level entries. The badge is a claim, not a grant. Until a LICENSE file appears, the terms you are actually operating under are unclear, and that matters more for a project you might build on than for one you merely read. Maintenance is also a consideration: the last push to main was on 2026-05-06, and the repository has no releases. There is no changelog and no versioning scheme, so any upgrade means diffing src/ yourself.

## MSA against a conventional RAG stack

The natural comparison is a standard RAG pipeline over the same corpus. The difference is architectural rather than incremental. A conventional RAG stack keeps the language model frozen and retrieves text chunks at inference time; the retriever is trained separately, if at all, and the generator never sees the retrieval decision as part of its own objective. MSA compresses documents into latent K/V, routes with a learned projector, and trains the whole loop end to end. Retrieval becomes a differentiable part of the model rather than a preprocessing step.

That buys the scaling behaviour the README describes, and it costs flexibility. A RAG stack lets you swap the embedding model, the vector store, or the generator independently, and it runs on a single GPU. MSA couples all of those into one checkpoint, and its deployment story is bound to the Memory Parallel engine. The README also positions MSA against HippoRAG2, which it reports beating by 14.8% on average, but HippoRAG2 is itself a graph-based retrieval method rather than a latent-memory model, so the comparison is between two different answers to the same question, not between two implementations of one.

## Where MSA is the wrong tool

MSA is not a drop-in replacement for a vector database, and it is not a serving framework. If your problem is retrieving a few relevant passages from a 500K-token corpus, a conventional retriever will be cheaper to operate and easier to debug, and the README's own NarrativeQA result shows that plain RAG with reranking can beat MSA on some datasets. If you need stable APIs, semantic versioning, or a support channel, the absence of releases and the single-author organization structure are disqualifying. The repository is a research artifact: a paper, a src/ tree, scripts, and a quick start document. Treating it as production infrastructure means accepting that you are the maintainer of your fork.

## Conclusion

Adopt MSA if you are researching extremely long-context memory, you can read the paper alongside the src/ tree, and you have multi-GPU hardware to reproduce the inference pipeline. Do not adopt it if you need a maintained library with releases, a stable API or a published licence file. Before anything else, verify the licence and the actual contents of QUICK_START.md, because the README badge says MIT while the repository carries no licence file.

## FAQ

### What is an LLM with persistent memory, and how does MSA relate to it?

The README describes MSA as an end-to-end trainable latent-state memory framework that keeps a compressed memory of a corpus and retrieves Top-k documents during decoding, which is how it extends effective context to 100M tokens without full attention.

### How does long-term memory work in AI, according to the MSA documentation?

In MSA, document latent states K/V/Kr are chunk-mean pooled offline, a router projector scores them by cosine similarity to select Top-k, and the compressed K/V of those documents are concatenated with the query's local K/V for generation. The README also describes Memory Interleave for multi-hop reasoning across scattered segments.

### How do I install EverMind-AI/MSA?

There is no pip package documented. The repository pins dependencies in requirements.txt, including torch==2.6 and transformers==4.51.3, and points to QUICK_START.md for the actual usage steps.

### What hardware does MSA need for 100M-token inference?

The README states that the Memory Parallel engine delivers 100M-token throughput on 2xA800 GPUs, with routing keys sharded across GPUs and content K/V held in host DRAM. No single-GPU path is documented.

### Is MSA licensed for commercial use?

The README badge links to MIT, but the repository has no licence file among its top-level entries, so the actual terms are not stated in the repository itself. Confirm this before building on the code.

## Sources

- [EverMind-AI/MSA on GitHub](https://github.com/EverMind-AI/MSA)
- [Issues](https://github.com/EverMind-AI/MSA/issues)
- [Project website](https://arxiv.org/abs/2603.23516)
- [README](https://github.com/EverMind-AI/MSA/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/evermind-ai-msa
