Doc-to-LoRA: A Hypernetwork That Encodes Documents Into LLM Weights
Hypernetworks that update LLMs to remember factual information
At a glance
- What is it?
- Doc-to-LoRA (D2L) is a research library from SakanaAI that trains a hypernetwork to generate LoRA weight updates from a document, letting a language model recall specific facts without carrying the document in its context window. It publishes a pretrained checkpoint for Gemma and requires Python 3.10 and uv.
- Who is it for?
- Doc-to-LoRA is a research codebase aimed at teams exploring alternatives to retrieval-augmented generation and long-context prompting for factual recall. The pretrained Gemma checkpoint lets practitioners run the technique without retraining.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 107 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
DEEP OPEN-SOURCE ANALYSIS
Factual Memory Without Consuming Context Window Tokens
Large language models answer factual questions either by recalling training-time knowledge or by reading a document placed in the context window. Both paths have limits. Training-time facts become stale. Context windows have a token ceiling, and long documents slow generation.
Doc-to-LoRA takes a different path. A hypernetwork reads a document and produces a set of LoRA weight deltas. Those deltas are applied to a frozen base LLM, which then answers questions about the document from its modified weights rather than from tokens in its prompt. The paper, presented at the Forty-third International Conference on Machine Learning, is available at arxiv.org/abs/2602.15902.
The technique is useful when the same document needs to be queried many times and the token cost of re-reading it on every call is prohibitive. It is less suited to one-off lookups or cases where the document changes frequently, since each update requires a new hypernetwork pass.
The Hypernetwork Architecture: From Text to LoRA Deltas
The central class is ModulatedPretrainedModel, defined in src/ctx_to_lora/modeling/hypernet.py. It wraps a pretrained base model and adds the hypernetwork that generates the LoRA weight updates. A call to model.reset() clears any previously applied LoRA deltas before encoding a new document.
The training pipeline uses self-generated synthetic data. The repository includes a viewer for inspecting these samples:
uv run webui/self_gen_viewer.pyTraining runs through accelerate with DeepSpeed for distributed execution. The configs/ directory holds the accelerate configuration. The experiment scripts follow a three-stage pattern: data download or generation, training, then evaluation. The README's experiment table lists the exact script paths for the main experiment and the Needle in a Haystack (NIAH) evaluation.
The pyproject.toml names the installable package ctx-to-lora, not doc-to-lora, which matters when importing in Python: the import path is ctx_to_lora, as the usage example shows.
Installing D2L and Loading the Pretrained Gemma Checkpoint
The repository uses uv for environment management. Install uv and then run the project's install script:
curl -LsSf https://astral.sh/uv/install.sh | sh
./install.shThe pretrained checkpoint targets Gemma. Download it from HuggingFace after logging in:
uv run huggingface-cli login
uv run huggingface-cli download SakanaAI/doc-to-lora --local-dir trained_d2l --include "*/"Load the checkpoint and prepare the model in Python:
import torch
from ctx_to_lora.model_loading import get_tokenizer
from ctx_to_lora.modeling.hypernet import ModulatedPretrainedModel
checkpoint_path = "trained_d2l/gemma_demo/checkpoint-80000/pytorch_model.bin"
state_dict = torch.load(checkpoint_path, weights_only=False)
model = ModulatedPretrainedModel.from_state_dict(
state_dict, train=False, use_sequence_packing=False
)
model.reset()
tokenizer = get_tokenizer(model.base_model.nameThe code comment notes that this high-level interface supports only non-batched inputs. For batched inference, the README points to src/ctx_to_lora/modeling/hypernet.py directly. Python 3.10 or later is required.
Interactive Demo and Experimental Evaluation Scripts
A Gradio-based interactive demo is included. Run it with:
uv run demo/app.pyThis starts a local web interface where you can provide a document and ask questions about it to see the internalized-fact recall in action.
The experimental scripts reproduce the results from the paper. The main experiment follows four files: 0-download_data.sh for data acquisition, 1-train.sh for training, and then the evaluation scripts under scripts/main_exp/eval/. The README notes that downloading data is fastest and that regenerating synthetic data is only necessary if fresh samples are needed.
The NIAH (Needle in a Haystack) evaluation, which probes how reliably the model recalls a planted fact inside a long document, has its own dedicated scripts under scripts/niah/. These run in order: data generation, training, then evaluation. All scripts are run with uv run from the repository root.
Where Doc-to-LoRA Falls Short
The high-level Python API is documented as non-batched only. This is explicitly noted in the source comment: batched inference requires using hypernet.py directly rather than the top-level loading functions. For applications that need to encode many documents concurrently, this is a real integration cost.
Generalization to documents outside the training distribution is not documented in the README. The technique requires a trained hypernetwork, and how well that hypernetwork transfers to document types or styles not present in its training data is left to the practitioner to evaluate.
The training infrastructure requires DeepSpeed, accelerate, and vLLM (version 0.8.5.post1 as pinned in pyproject.toml). These are large dependencies with strict version constraints. A project that already pins different versions of transformers (the repo pins 4.51.3) or vLLM will face dependency conflicts.
The repository has no GitHub releases, and the last push was on 2026-06-15.
Doc-to-LoRA Compared With Retrieval-Augmented Generation
Retrieval-Augmented Generation is the most common approach to grounding LLM answers in specific documents. A RAG system indexes document chunks, retrieves the most relevant ones at query time, and injects them into the prompt. Any document can be added to the index without modifying model weights, and the technique generalizes broadly.
Doc-to-LoRA takes the opposite approach. Rather than retrieving text and placing it in context, it encodes the document into parameter space before any query arrives. The advantage is zero token overhead per query once encoding is done. The disadvantage is that encoding requires a forward pass through the hypernetwork, and the technique depends on that hypernetwork having generalized to the document's content and style during training.
For a single document queried many times, D2L can reduce per-query cost. For a large or frequently changing corpus, RAG remains the more practical choice because adding a document requires only indexing, not a new hypernetwork pass.
Dependencies, Maintenance, and MIT Licensing
The pyproject.toml pins specific versions of core dependencies: transformers 4.51.3, deepspeed 0.17.1, accelerate 1.6.0, and vllm 0.8.5.post1. These pins ensure reproducibility for the paper's experiments but create friction for projects with their own transitive dependencies.
The full dependency list is long, including wandb for experiment tracking, Gradio 4.40.0 or later for the demo, google-cloud-storage, and llmlingua among others. Most of these are needed only for training or evaluation; a project that only loads the pretrained checkpoint and runs inference faces a lighter subset, but pyproject.toml does not separate optional dependencies from required ones.
The repository is MIT licensed. The pretrained checkpoint is hosted on HuggingFace under the SakanaAI organization; its license is not stated in the README, so verifying terms on the HuggingFace model page is necessary before redistribution.
The last push was on 2026-06-15. The repository is not archived.
Editorial conclusion
Doc-to-LoRA is a research codebase aimed at teams exploring alternatives to retrieval-augmented generation and long-context prompting for factual recall. The pretrained Gemma checkpoint lets practitioners run the technique without retraining. The current Python API does not support batched inference through its high-level interface; production use requires working with hypernet.py directly. Before committing to D2L, verify that your document distribution is close enough to the training data to expect useful generalization, since the README does not document the distribution limits of the released checkpoint.
Frequently asked questions
How does Doc-to-LoRA differ from in-context learning with a long document?
Doc-to-LoRA encodes the document into LoRA weight deltas applied to the base model before any query, so no document tokens occupy the context window during inference. In-context learning places the document directly in the prompt, consuming tokens and slowing generation for every query.
What base model does the released Doc-to-LoRA checkpoint use?
The pretrained checkpoint available on HuggingFace at SakanaAI/doc-to-lora targets Gemma. The checkpoint is at trained_d2l/gemma_demo/checkpoint-80000/pytorch_model.bin after download.
Does Doc-to-LoRA support batched document encoding?
The high-level Python API documented in the README supports only non-batched inputs. The README notes that batched inference requires using src/ctx_to_lora/modeling/hypernet.py directly.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/sakanaai-doc-to-lora)
Community notes