The Recurrent Looped Transformer's best task has a 46 point deviation
Official Project Page for Recurrent Looped Transformer (RLT)
At a glance
- What is it?
- A project page reporting 108 training runs of a recurrent decoder that feeds each token's final hidden state forward. Three seeds for one split score 17.58, 19.53 and 98.96, and the page ships plots rather than code.
- Who is it for?
- This is worth reading as an early report rather than a settled claim, and the page helps you do that because it states its own statistics plainly. The parity result is the strong one: an RLT-1 split at six and two reaches 99.44 percent parity accuracy at step 500 where the eight-layer Transformer baseline sits at 48.48, and all three seeds are at 100 percent by step 600.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 2 days ago.
- What is it written in?
- Mainly HTML, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 9, 2026, and from our analysis. They are not legal advice.
Editorial analysis
One row of the table has a 46 point deviation
The depth-eight snapshot compares six configurations, RLT-1 at splits 4+4, 5+3, 6+2, 7+1 and 8+0 plus an eight-layer Transformer, on six algorithmic tasks. All 108 runs completed 2,000 optimizer steps using initialization seeds 42, 43 and 44 for every task, with training examples and held-out sets fixed across initializations, and every aggregate reported as mean plus or minus sample standard deviation with n of 3 and a degrees-of-freedom correction of 1. Now look at the flat mod-5 row, the one without brackets:
| Task | RLT-1 4+4 | RLT-1 5+3 | RLT-1 6+2 | RLT-1 7+1 | RLT-1 8+0 | Transformer 8 | | --- | ---: | ---: | ---: | ---: | ---: | ---: | | Mod 5, no brackets | 45.36±46.43 | 94.18±7.29 | 69.62±33.01 | 70.01±43.15 | 60.33±35.70 | 64.02±37.64 |
Four of the six entries have deviations between 33 and 46 points on means between 45 and 70. A standard deviation larger than half the mean means the three runs disagreed about whether the task was solved at all. This is a small experiment by construction, and the deviation is the arithmetic consequence of three seeds rather than a defect in any one of them.
The three seeds of one split are 17.58, 19.53 and 98.96
The page spells out the per-seed numbers for the configuration with the largest deviation, and they are the most informative line in the whole report. The three 4+4 seeds score 17.58%, 19.53% and 98.96%. Two runs failed to learn the task and one learned it almost completely, which is why the aggregate reads 45.36 with a deviation of 46.43 rather than something near 45. Read that way, the mean is not a measurement of the method at all; it is a measure of how often an initialization escapes a bad basin. The configuration that reports the tightest number, 5+3 at 94.18 plus or minus 7.29, is the one where all three seeds converged. So the table is really two tables: one for tasks where every seed converges and one for a task where the outcome is bimodal, and the mean plus or minus form hides that difference completely. The page is candid about it in the sentence introducing the number, which says flat mod-5 varies strongly with initialization. It would be more useful if the per-seed figures appeared in the table itself.
The bracketed variant is a different task and the gap disappears
The second mod-5 row is bracketed rather than flat, and it tells the opposite story about the method. Every RLT-1 split lands between 70.53 and 75.17 percent, and the eight-layer Transformer reaches 73.87 plus or minus 9.07. So in the bracketed form the baseline is no longer behind, and the best RLT-1 split beats it by well under a point. The page explains why the two rows cannot be read as the same experiment: the generators differ in operator structure and label distribution, and parentheses also occupy token positions. That last clause is the mechanical detail that matters most, since a token-level model has to spend capacity on the bracket characters before it can spend it on the arithmetic. The honest summary is that the architecture's advantage on mod-5 appears in one formulation of the task and not in the other, and the page's own framing, that bracketed model means are closer, acknowledges it without explaining the cause.
S5 standard sits near zero on an axis the page admits is narrower
The S5 swaps row is close to solved for everything, with four configurations at exactly 100.00 plus or minus 0.00 and the other two at 99.35 and 99.61 against a baseline at 99.09. The S5 standard row is the opposite. The six entries are 0.78, 2.47, 2.21, 2.08, 1.82 and 0.52 percent. Nothing there is a result; every model is at or near the floor. What makes this worth a paragraph is the footnote attached to the table, which says the bands show sample standard deviation clipped to the accuracy range and that standard S5 uses a narrower vertical scale. A chart with a narrower axis makes a near-zero number look like a meaningful one, and the page flags that choice itself rather than leaving the reader to notice. It is the kind of disclosure that costs a little credibility and buys a lot, and it is the reason the rest of the table can be taken at face value. Parity and mod-5 score the final label; S5 scores the final state; addition measures teacher-forced answer-token accuracy including answer formatting and the end-of-sequence token, excluding prompt and padding positions.
Addition is saturated for every model including the baseline
The addition row is 100.00 plus or minus 0.00 in all six columns, baseline included. That row carries no information about the architecture, and it is worth knowing that before reading the rest of the table as a comparison. The same is close to true of S5 swaps, where the spread between the best RLT-1 split and the baseline is under half a point. So of the six tasks, two are at ceiling for everything and one is at the floor for everything. The task that actually separates the models is parity, and the page reports it as a learning-speed result rather than a final-accuracy result: at step 500, RLT-1 6+2 reaches 99.44 plus or minus 0.98 percent while the eight-layer Transformer sits at 48.48 plus or minus 0.53, and all three 6+2 seeds reach 100 percent by step 600. At step 2,000 the splits from 4+4 through 7+1 are at 100 in all three seeds while the Transformer ends at 94.84 plus or minus 3.43. Both panels use the same 768 validation examples and correspond to 256,000 and 1,024,000 training examples per seed, so the comparison is at matched data rather than matched steps alone.
One table is a fixed step, the other is a selected checkpoint
The accuracy table is titled as validation accuracy after 2,000 steps, with every run having consumed 1,024,000 training examples and all curves extending through step 2,000. The length generalization section works differently. Each run selects its lowest in-distribution validation-loss checkpoint over the full training history, taking the earliest step on ties, and the page states that test results do not enter selection. So one table reports a fixed point on the training curve and the other reports the best point the curve ever reached on a validation criterion. That is the right way to do length generalization and the wrong way to read an accuracy comparison, and the two sets of numbers should not be placed next to each other. The evaluation protocol for the generalization runs is described as giving every model and seed the same 256 addition pairs or 1,024 formal-task sequences at every task and length. Whether that is the same protocol as the accuracy table is not stated on the visible page.
Three variants, and chunk size one recovers the base case
The architecture section describes three points in a design space rather than three unrelated models. RLT-1 is the full version: the decoder's final hidden state goes to the next token together with that token's causal encoder representation, and the state pair at step t is written as a recurrent output plus a layerwise sliding-window key-value cache, with the initial state carrying an empty cache. The decoder reads encoder-derived global key-value memory, and with one memory group every decoder layer reads the same projected memory using its own queries. Local memory is per layer, with a window that includes the current token and retains up to W minus 1 past entries. RLT-0 removes the feedback path, the learned initial state, state normalization, the gated merge and the feedback projection, which together cost 787,968 parameters per split; known tokens then run in parallel within each decoder layer during training and prefill. RLT-2 holds the feedback state fixed within a chunk and updates it at chunk boundaries anchored at BOS and continuing across prompt, response and message boundaries, and at chunk size one the equations recover RLT-1 at the same weights. So the three form a clean interpolation, and the mod-5 experiments evaluate RLT-0 across splits 4+4 to 8+0 and RLT-2 at chunk sizes four and eight with measured CPU training-step times.
The repository is a page, two PDFs and a directory of plots
The top level holds a LICENSE file, the README, an English PDF of the paper, a Chinese PDF of the paper, an assets directory, a figure1.png and an index.html. There is no source tree, no model definition, no training script and no data. The primary language is recorded as HTML, which matches: the deliverable is a project page under a GitHub Pages site, with the default branch called master and no GitHub releases at all. The mechanics of the architecture are therefore explained in binary form rather than in code, through an architecture detail PDF, a control architecture PDF for the no-feedback variant, a chunk architecture diagram and a chunk schedule PDF, with the validation curves, the parity error bars and the depth-sixteen results also linked as PDFs inside assets. A reader who wants to reproduce any of this has to rebuild it from the equations and the prose. The bookkeeping is otherwise careful: the report is dated September 12, 2026 and marked updated September 20, the depth-eight snapshot is September 17, and the last recorded push to the branch is 2026-09-22, with Apache-2.0 in the metadata matching the LICENSE file.
Editorial conclusion
This is worth reading as an early report rather than a settled claim, and the page helps you do that because it states its own statistics plainly. The parity result is the strong one: an RLT-1 split at six and two reaches 99.44 percent parity accuracy at step 500 where the eight-layer Transformer baseline sits at 48.48, and all three seeds are at 100 percent by step 600. The mod-5 result is not strong, and the reason is visible in the per-seed numbers rather than hidden by the mean: one split's three seeds land at 17.58, 19.53 and 98.96, so the method either learns the task or does not, and with three seeds the aggregate is close to meaningless for that row. Three things to check before building on it. Whether the two mod-5 rows are the same task, since the page says the generators differ in operator structure and label distribution and the gap disappears in the bracketed version. Whether the length-generalization table, which uses a validation-selected checkpoint, is comparable to the accuracy table, which is a fixed step count. And whether there is anything to run, because there is not: the repository is a page, two PDFs and a directory of plots.
Frequently asked questions
What is the Recurrent Looped Transformer?
RLT-1 passes the decoder's final hidden state to the next token together with that token's causal encoder representation. The decoder reads encoder-derived global key-value memory and keeps a sliding-window attention cache at every layer, and the same update runs over prompt and response tokens.
How large are the reported experiments?
Six algorithmic tasks, 108 runs, 2,000 optimizer steps each, three initialization seeds, width 512, FFN width 1,365, four attention heads, global batch 512 and microbatch 32. RLT-1 has 26.10 to 28.73M parameters against 25.31M for the eight-layer Transformer.
Does the recurrent-looped-tranformer repository contain code?
No. The top level holds a LICENSE, the README, an English and a Chinese PDF of the paper, an assets directory, a figure and an index.html. It is a project page with plots and architecture diagrams, and the primary language is recorded as HTML.
What do the split names like 4+4 and 8+0 mean?
The page uses a two-number split notation for the untied eight-layer layouts and describes the 8+0 variant as having no decoder blocks while still applying the gated recurrent merge. The depth-sixteen section continues the same notation.
What is TBPTT set to in the reported runs?
128, alongside a sliding-window attention window of eight, one shared encoder-memory group and a feedback scale of 0.1. The page notes that 128 covers every training sequence in these runs and cuts no gradients.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/yifanzhang-pro-recurrent-looped-tranformer)