# Doc-to-LoRA: A Hypernetwork That Encodes Documents Into LLM Weights

> Doc-to-LoRA (D2L) is a research library from SakanaAI that trains a hypernetwork to generate LoRA weight updates from a document, letting a language model recall specific facts without carrying the document in its context window. It publishes a pretrained checkpoint for Gemma and requires Python 3.10 and uv.

**SakanaAI/doc-to-lora** — Hypernetworks that update LLMs to remember factual information

- Repository: https://github.com/SakanaAI/doc-to-lora
- Website: https://arxiv.org/abs/2602.15902
- Stars: 828 · Forks: 105
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/sakanaai-doc-to-lora

## Factual Memory Without Consuming Context Window Tokens

Large language models answer factual questions either by recalling training-time knowledge or by reading a document placed in the context window. Both paths have limits. Training-time facts become stale. Context windows have a token ceiling, and long documents slow generation.

Doc-to-LoRA takes a different path. A hypernetwork reads a document and produces a set of LoRA weight deltas. Those deltas are applied to a frozen base LLM, which then answers questions about the document from its modified weights rather than from tokens in its prompt. The paper, presented at the Forty-third International Conference on Machine Learning, is available at arxiv.org/abs/2602.15902.

The technique is useful when the same document needs to be queried many times and the token cost of re-reading it on every call is prohibitive. It is less suited to one-off lookups or cases where the document changes frequently, since each update requires a new hypernetwork pass.

## The Hypernetwork Architecture: From Text to LoRA Deltas

The central class is ModulatedPretrainedModel, defined in src/ctx_to_lora/modeling/hypernet.py. It wraps a pretrained base model and adds the hypernetwork that generates the LoRA weight updates. A call to model.reset() clears any previously applied LoRA deltas before encoding a new document.

The training pipeline uses self-generated synthetic data. The repository includes a viewer for inspecting these samples:

```bash
uv run webui/self_gen_viewer.py
```

Training runs through accelerate with DeepSpeed for distributed execution. The configs/ directory holds the accelerate configuration. The experiment scripts follow a three-stage pattern: data download or generation, training, then evaluation. The README's experiment table lists the exact script paths for the main experiment and the Needle in a Haystack (NIAH) evaluation.

The pyproject.toml names the installable package ctx-to-lora, not doc-to-lora, which matters when importing in Python: the import path is ctx_to_lora, as the usage example shows.

## Installing D2L and Loading the Pretrained Gemma Checkpoint

The repository uses uv for environment management. Install uv and then run the project's install script:

```bash
curl -LsSf https://astral.sh/uv/install.sh | sh
./install.sh
```

The pretrained checkpoint targets Gemma. Download it from HuggingFace after logging in:

```bash
uv run huggingface-cli login
uv run huggingface-cli download SakanaAI/doc-to-lora --local-dir trained_d2l --include "*/"
```

Load the checkpoint and prepare the model in Python:

```python
import torch

from ctx_to_lora.model_loading import get_tokenizer
from ctx_to_lora.modeling.hypernet import ModulatedPretrainedModel

checkpoint_path = "trained_d2l/gemma_demo/checkpoint-80000/pytorch_model.bin"
state_dict = torch.load(checkpoint_path, weights_only=False)
model = ModulatedPretrainedModel.from_state_dict(
    state_dict, train=False, use_sequence_packing=False
)
model.reset()
tokenizer = get_tokenizer(model.base_model.name
```

The code comment notes that this high-level interface supports only non-batched inputs. For batched inference, the README points to src/ctx_to_lora/modeling/hypernet.py directly. Python 3.10 or later is required.

## Interactive Demo and Experimental Evaluation Scripts

A Gradio-based interactive demo is included. Run it with:

```bash
uv run demo/app.py
```

This starts a local web interface where you can provide a document and ask questions about it to see the internalized-fact recall in action.

The experimental scripts reproduce the results from the paper. The main experiment follows four files: 0-download_data.sh for data acquisition, 1-train.sh for training, and then the evaluation scripts under scripts/main_exp/eval/. The README notes that downloading data is fastest and that regenerating synthetic data is only necessary if fresh samples are needed.

The NIAH (Needle in a Haystack) evaluation, which probes how reliably the model recalls a planted fact inside a long document, has its own dedicated scripts under scripts/niah/. These run in order: data generation, training, then evaluation. All scripts are run with uv run from the repository root.

## Where Doc-to-LoRA Falls Short

The high-level Python API is documented as non-batched only. This is explicitly noted in the source comment: batched inference requires using hypernet.py directly rather than the top-level loading functions. For applications that need to encode many documents concurrently, this is a real integration cost.

Generalization to documents outside the training distribution is not documented in the README. The technique requires a trained hypernetwork, and how well that hypernetwork transfers to document types or styles not present in its training data is left to the practitioner to evaluate.

The training infrastructure requires DeepSpeed, accelerate, and vLLM (version 0.8.5.post1 as pinned in pyproject.toml). These are large dependencies with strict version constraints. A project that already pins different versions of transformers (the repo pins 4.51.3) or vLLM will face dependency conflicts.

The repository has no GitHub releases, and the last push was on 2026-06-15.

## Doc-to-LoRA Compared With Retrieval-Augmented Generation

Retrieval-Augmented Generation is the most common approach to grounding LLM answers in specific documents. A RAG system indexes document chunks, retrieves the most relevant ones at query time, and injects them into the prompt. Any document can be added to the index without modifying model weights, and the technique generalizes broadly.

Doc-to-LoRA takes the opposite approach. Rather than retrieving text and placing it in context, it encodes the document into parameter space before any query arrives. The advantage is zero token overhead per query once encoding is done. The disadvantage is that encoding requires a forward pass through the hypernetwork, and the technique depends on that hypernetwork having generalized to the document's content and style during training.

For a single document queried many times, D2L can reduce per-query cost. For a large or frequently changing corpus, RAG remains the more practical choice because adding a document requires only indexing, not a new hypernetwork pass.

## Dependencies, Maintenance, and MIT Licensing

The pyproject.toml pins specific versions of core dependencies: transformers 4.51.3, deepspeed 0.17.1, accelerate 1.6.0, and vllm 0.8.5.post1. These pins ensure reproducibility for the paper's experiments but create friction for projects with their own transitive dependencies.

The full dependency list is long, including wandb for experiment tracking, Gradio 4.40.0 or later for the demo, google-cloud-storage, and llmlingua among others. Most of these are needed only for training or evaluation; a project that only loads the pretrained checkpoint and runs inference faces a lighter subset, but pyproject.toml does not separate optional dependencies from required ones.

The repository is MIT licensed. The pretrained checkpoint is hosted on HuggingFace under the SakanaAI organization; its license is not stated in the README, so verifying terms on the HuggingFace model page is necessary before redistribution.

The last push was on 2026-06-15. The repository is not archived.

## Conclusion

Doc-to-LoRA is a research codebase aimed at teams exploring alternatives to retrieval-augmented generation and long-context prompting for factual recall. The pretrained Gemma checkpoint lets practitioners run the technique without retraining. The current Python API does not support batched inference through its high-level interface; production use requires working with hypernet.py directly. Before committing to D2L, verify that your document distribution is close enough to the training data to expect useful generalization, since the README does not document the distribution limits of the released checkpoint.

## FAQ

### How does Doc-to-LoRA differ from in-context learning with a long document?

Doc-to-LoRA encodes the document into LoRA weight deltas applied to the base model before any query, so no document tokens occupy the context window during inference. In-context learning places the document directly in the prompt, consuming tokens and slowing generation for every query.

### What base model does the released Doc-to-LoRA checkpoint use?

The pretrained checkpoint available on HuggingFace at SakanaAI/doc-to-lora targets Gemma. The checkpoint is at trained_d2l/gemma_demo/checkpoint-80000/pytorch_model.bin after download.

### Does Doc-to-LoRA support batched document encoding?

The high-level Python API documented in the README supports only non-batched inputs. The README notes that batched inference requires using src/ctx_to_lora/modeling/hypernet.py directly.

## Sources

- [Issues](https://github.com/SakanaAI/doc-to-lora/issues)
- [License: MIT](https://github.com/SakanaAI/doc-to-lora/blob/main/LICENSE)
- [Project website](https://arxiv.org/abs/2602.15902)
- [README](https://github.com/SakanaAI/doc-to-lora/blob/main/README.md)
- [SakanaAI/doc-to-lora on GitHub](https://github.com/SakanaAI/doc-to-lora)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/sakanaai-doc-to-lora
