Open-source project
yifanzhang-pro/deep-delta-learning avatar
yifanzhang-pro/deep-delta-learning

Deep Delta Learning says out loud that it does not enlarge the function class

Official Project Page for Deep Delta Learning (https://arxiv.org/abs/2601.00417)

364 stars26 forksPythonApache-2.0

At a glance

What is it?
A January 2026 paper and its PyTorch and Triton implementations, replacing the residual stream with a gated rank-1 delta update applied over depth rather than over time. The best loss and the best one-shot accuracy belong to different variants at different scales, and every run was a single unreplicated point estimate.
Who is it for?
Read this project if you are working on residual stream design and want a delta-rule parameterisation that has both a clean mathematical characterisation and runnable code, at two small scales, with the throughput and memory cost stated next to the accuracy gain. Do not treat the results table as evidence of a settled improvement.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 14 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 9, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The update is still additive, by the authors' own statement

The most useful sentence on the project page is a concession. A sufficiently expressive residual block can already represent content replacement, but standard architectures do not parameterise reading, comparison and replacement as an explicit residual operation. Deep Delta Learning's answer is to keep the identity path and add target-seeking edits.

What it is not is more expressive. The write up says plainly that the update is still additive, so DDL does not enlarge the function class; it makes the edit target-seeking. That is a claim about parameterisation and inductive bias rather than about capacity, and it is unusual to see written this plainly on a project page.

The mechanic is rank-1. Each layer reads the residual state along a learned direction, compares that readout with a learned target, and writes back a gated rank-1 correction along the same direction. The read/write direction comes from the normalised output of the attention or MLP sublayer, the target comes from a lightweight branch reading the sublayer input, and the gate comes from the normalised context.

The gate has three regimes and a hard ceiling at 2

The shortcut matrix for a given direction and gate is a rank-one projector, and its spectrum gives three local regimes, which is the cleanest part of the treatment.

The shortcut has eigenvalue 1 on the perpendicular complement of the direction, with multiplicity d minus one, and eigenvalue 1 minus beta along the direction itself. That single eigenvalue therefore controls what the layer does to the readout. Near zero, the update approaches the identity and the layer skips. At exactly one, the readout error is annihilated and the target is written in exactly. Above one and below two, the readout overshoots and crosses the target.

The ceiling is the interesting part. As beta approaches 2, the shortcut, and specifically the shortcut rather than the full update, approaches the Householder reflector. That is the standard orthogonal reflection matrix, and it is the natural boundary of this gate parameterisation: the design cannot express an orthogonal reflection of the residual state, only approach it. The gate is bounded in the open interval by construction, since it is twice a sigmoid.

One run per configuration, and the caveat is in the table

The results are honest in a way that is worth quoting rather than paraphrasing. Decoder-only models of roughly 124 million and roughly 353 million parameters were trained on FineWeb-Edu for 49.15 billion tokens, and validation loss and average one-shot accuracy were measured over eight benchmarks.

Then the caveat, stated in the same place: each configuration was trained once with the same token budget, so these are point estimates rather than compute-matched comparisons, and the expanded-state gains are not separated from the added residual capacity.

Both halves of that sentence matter. A single run cannot distinguish a real improvement from seed variance, and a table with one column of numbers per cell should not be read as a mean over seeds. And if the expanded state carries more channels, then some of any gain over the scalar version is just that extra capacity rather than the delta rule. The project says the paper discusses these limits in detail.

The gains themselves are small in absolute terms: the small-scale baseline validation loss of 2.8543 improves to 2.8299 at best, and the medium-scale baseline of 2.6053 improves to 2.5758.

The best variant at small scale is not the best at medium scale

There are six rows and two axes, and reading them as one ranking gets the wrong answer.

At the small scale the best validation loss belongs to the token-compressed variant at 2.8299, and it also takes the best one-shot accuracy at 49.47 percent. The channel-mixed variant is second on both at 2.8329 and 49.29 percent. At the medium scale the ordering flips: the channel-mixed variant has the best loss at 2.5758 and the best one-shot accuracy at 55.14 percent, while the token-compressed variant is second on loss at 2.5905.

So the token-compressed variant wins small and the channel-mixed variant wins medium, which is exactly the kind of result that gets flattened into a claim about a method when the summary drops the scale.

The two axes are also doing different jobs. TC compresses the expanded residual state with a causal convolution across tokens, while CC mixes the value channels at each token without moving information along the sequence. EC is a separate initialisation of the expanded state, and the rows marked without EC are the ablation that removes it.

Expanded states buy capacity by spending throughput

The cost side is given as precisely as the benefit side, and it is the number to look at before you plan a training run.

At the small scale the channel-mixed expanded-state variant trains at 1158.0 thousand tokens per second with 3.08 gigabytes of peak memory, against 1509.6 thousand tokens per second and 2.94 gigabytes for the baseline. That is roughly a quarter less throughput for about five percent more peak memory.

The abstract states the same shape as a general finding rather than one measurement: DDL improves validation loss and one-shot accuracy over additive residual baselines, with lower throughput in every measured configuration and higher peak memory for expanded states.

So the throughput penalty is not specific to one variant at one scale, and the memory penalty is specific to the expanded states, which is consistent with the scalar variant keeping the ordinary residual vector.

The design intent behind the memory cost is stated too: the expanded state stores several persistent value channels while attention and MLP computation stay at the original model width, so residual-state capacity can grow without widening the backbone.

Ten model files, one for each variant and its accelerated twin

The `model/` directory holds PyTorch implementations, and each file defines a `GPTConfig` and a `GPT` class that is a Hugging Face `PreTrainedModel`. The reported runs use each file's own default configuration, which is worth noting because it means the defaults are the experiment rather than a simplified reproduction of it.

The defaults give the roughly 124 million model. The roughly 353 million model is the one that sets `num_hidden_layers=24`, `num_attention_heads=8` and `hidden_size=1024`.

Five variants are shipped, and each has two files: a reference implementation and a Triton accelerated implementation. The scalar variant with dimension one, the token-compressed variant with and without the expanded-state initialisation, and the channel-mixed variant with and without it. The naming encodes the whole grid, so reading a filename tells you which cell of the ablation it is.

The dependency file pins exact versions for torch and transformers, allows a bounded range for pydantic, and installs triton only where the platform is Linux. That last marker is the practical boundary of the accelerated path.

The depth-wise framing is a change of axis, not a new rule

The positioning against prior work is careful. DeltaNet applies the delta rule over time to update a memory matrix. DDL applies the same erase and write update over network depth, using it as the residual interface between Transformer sublayers.

The project then says something that would be easy to omit: the rule itself is prior work, and DDL's contribution is its depth-wise use and analysis.

That framing matters for how you read the rest. The erase-then-write structure is inherited, not proposed. What is proposed is that the sequence position is the right place to apply it, and that the residual stream rather than an explicit memory matrix is the right thing to update.

The spectral treatment is likewise scoped explicitly. The analysis describes the operator for a given direction and a given gate, and it does not show that the learned directions correspond to human-readable features. So the mathematics tells you what the operator can do, not what any particular trained model has learned.

Editorial conclusion

Read this project if you are working on residual stream design and want a delta-rule parameterisation that has both a clean mathematical characterisation and runnable code, at two small scales, with the throughput and memory cost stated next to the accuracy gain. Do not treat the results table as evidence of a settled improvement. Four things to check before you build on it. That the numbers are single runs rather than averages, since the project says so explicitly, and that the expanded-state variants are not separated from their extra capacity. That the spectral argument is about the operator for a given direction and gate, not a claim that learned directions are interpretable, which the project also says. Which variant you want, because the winner changes between scales: the token-compressed variant has the best small-scale loss and the channel-mixed variant the best medium-scale loss. And that the Triton implementations only install on Linux, since the dependency is gated on that platform. Licence is Apache-2.0, there are no GitHub releases, and the last push to master is dated 25 September 2026.

Frequently asked questions

What is Deep Delta Learning?

It is a structured residual update for Transformers, from a January 2026 paper by Yifan Zhang, Yifeng Liu, Mengdi Wang and Quanquan Gu at Princeton and UCLA. Each layer reads the residual state along a learned direction, compares the readout with a learned target, and writes back a gated rank-1 correction along that same direction.

Does Deep Delta Learning make Transformers more expressive?

No, and the project says so directly. The update is still additive, so DDL does not enlarge the function class. What changes is that the edit becomes target-seeking, with reading, comparison and replacement parameterised explicitly as a residual operation rather than left implicit.

How were the Deep Delta Learning models trained and evaluated?

Decoder-only models of roughly 124 million and roughly 353 million parameters, trained on FineWeb-Edu for 49.15 billion tokens, then measured on validation loss and average one-shot accuracy over eight benchmarks. Each configuration was trained once, so the numbers are point estimates rather than compute-matched comparisons.

What is the cost of the expanded residual state in Deep Delta Learning?

At the small scale the channel-mixed expanded variant trains at 1158.0K tokens per second with 3.08 GB peak memory, against 1509.6K tokens per second and 2.94 GB for the baseline. The abstract notes lower throughput in every measured configuration and higher peak memory for expanded states.

What code does the deep-delta-learning repository ship?

PyTorch implementations under model/, where each file defines a GPTConfig and a GPT model that is a Hugging Face PreTrainedModel. Five variants are provided, each as a reference file and a Triton accelerated file, and the reported runs use each file's default configuration.

Official sources

  1. Issues
  2. License: CC-BY-4.0
  3. Project website
  4. README
  5. yifanzhang-pro/deep-delta-learning on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/yifanzhang-pro-deep-delta-learning.svg)](https://hysenlabs.com/projects/yifanzhang-pro-deep-delta-learning)