# Dream-RSI: an official repository with no code in it yet

> Seventeen authors across Google, Google DeepMind and two universities argue that a finished self-improvement run can be replayed as a simulator, so a different exploration policy can be scored without executing anything again. The idea is well specified. The repository is a README, a citation file, a folder of images and a paper.

**zhengkid/Dream-RSI** — The offical repo for "Dream-RSI: Recursive Self-Improvement through Evolving Worlds"

- Repository: https://github.com/zhengkid/Dream-RSI
- Stars: 1,336 · Forks: 119
- Language: Unknown
- License: not declared
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/zhengkid-dream-rsi

## An official repository with five entries and no code

The top level is .gitignore, CITATION.cff, README.md, assets/ and papers/.

That is the whole thing. There is no source tree, no package manifest, no lock file, no test directory and no data. The repository's declared primary language is unknown and so is its licence, which is unusual for a project of this size and worth resolving before anyone builds on it.

The README says so itself. A marked note near the top reads that code is being prepared for release and points at a release plan, and the plan table lists the paper as available, the project page and interactive demo as available, the arXiv posting as in progress, and the discovered programs, the full codebase and the reproduction scripts all as being prepared.

So what a reader can take from this repository today is the paper, the project page, a citation file and the images embedded in the README. The description even calls it the official repo, with a misspelling of official in the original.

## The citation block carries an arXiv number the release table says is still in progress

The README ends with a BibTeX block, and that block is complete: title, all seventeen authors in order, a year, and a journal field.

The journal field reads arXiv preprint arXiv:2609.14858.

Four lines earlier, the release plan table gives the arXiv posting a status of in progress, marked with a single arrow.

So the citation a reader is told to paste into their own paper already carries an identifier for a posting that the same document describes as not yet made. Either the preprint went live after the table was written, or the identifier was reserved in advance, and the document does not say which.

It is a small thing, and it is the kind of thing that matters when the artifact is a citation. Anyone copying that block today gets an identifier whose status the repository itself calls unconfirmed.

The machine-readable citation at the repository root, CITATION.cff, is the version a tool would read, and the two files have to agree for the entry to be worth anything.

## Every number in the README is a picture

The document's quantitative content is carried by images rather than text.

There is a statistics graphic near the top, a figure for the overview diagram, the project tagline, and three institutional logos. All of them are dark-theme variants referenced from an assets directory, which means the numbers a reader is most likely to notice are not in text that can be copied, quoted or checked.

What survives as prose is one empirical sentence: across algorithm engineering, mathematical optimisation and GPU kernel engineering, the method improves both discovery effectiveness and efficiency in several settings.

Several settings is doing a lot of work there. No baseline is named, no metric is named, no figure appears, and there is no ablation in the text. The three domains are named, which is more than a paper link would give you, but a reader cannot tell from this document whether the improvement is a few percent or an order of magnitude, or whether it holds on all three domains or one.

For that, the paper and the interactive walkthrough on the project page are the only sources available, and both are outside the repository.

## The evaluator replays recorded trees instead of running anything

The core insight is stated plainly: accumulated discovery history can serve as a replay simulator over the realised search space.

The mechanism has four parts. A completed discovery process already records a structured tree of past exploration decisions together with their realised code-execution outcomes. An alternative policy can then traverse that same tree differently, choosing different subsets of recorded branches, in different orders, with different parallel groupings and different stopping decisions. Because every outcome is already saved, evaluating the policy requires only reading past records, with no rerunning of the discovery agent and no rerunning of the evaluator.

The analogy is to model-based reinforcement learning and world models, and the phrase the document lands on is that the history becomes a world the agent can dream in.

That analogy is also where the tension sits, and a reader has to resolve it. A world model predicts what would happen in states not yet seen. This simulator replays states that have already happened. The pool is finite and grows only as new online exploration runs complete, so a policy that would exploit branches nobody has taken yet gets no credit for it. The README does not claim otherwise, and the cost of the shortcut is the price of the immediacy.

## The orchestration layer is explicit, and the coding agent underneath is untouched

The engineering claim is narrow and specific, which is the most useful thing in the overview.

A lightweight orchestration layer makes exploration explicit and programmable. Branching, parallel exploration and stopping are named as the three things the layer owns, and the underlying coding agent is left unchanged.

That matters for anyone trying this on their own agent loop. The contribution is not a better coder and not a better evaluator; it is that the search structure becomes something you can write down, schedule and instrument.

The loop itself is three numbered stages, given with inline numerals in the text. The current policy drives online discovery and logs its traces. The recorded trees are converted into a reusable simulator pool. Candidate policies are evaluated and refined by dreaming over that pool, which returns immediate and cheap off-policy feedback, and the improved policy is redeployed online, continuously expanding the pool.

So the system is bootstrapped by its own search, and the pool is both the training signal and the constraint.

## Three domains, one claim, no comparison

The highlights section has three entries and they are not the same kind of statement.

The first is conceptual, naming the treatment of completed discovery histories as replay simulators and the claim that this makes delayed exploration feedback reusable for meta-exploration policy evaluation.

The second is architectural, describing the loop in the same terms as the overview: collect histories online, construct simulators from them, refine the exploration strategy by dreaming, redeploy.

The third is empirical, and it is the short one: across algorithm engineering, mathematical optimisation and GPU kernel engineering, the method improves both discovery effectiveness and efficiency in several settings.

Two of the three highlights restate the idea, and the one that would let a reader judge the work is the one with no numbers in it.

That is a defensible choice for a preprint page whose job is to route people to the paper, and it is also why the release plan matters: the discovered programs, when they arrive, are the artefact that would let anyone see what the improved policies actually produced.

## Seventeen authors, four institutions, two corresponding

The author block is a good map of where the work sits.

Seventeen names, with superscript affiliations resolved at the foot: Google, the University of Maryland College Park, Google DeepMind and the University of Virginia. The first author carries two affiliations, Google and Maryland, and the senior names cluster at Google.

Two authors are marked as corresponding, both at Google, and the mark is repeated in the caption line rather than shown once.

The paper itself lives in the repository under a papers directory, linked twice at the top of the README, once through a paper badge and once through the references, and there is a project page at a separate domain with an interactive walkthrough.

At the root, CITATION.cff sits alongside the README rather than in a folder, which is the one piece of tooling this minimal repository does have, alongside a .gitignore. For a repository whose entire deliverable is a citation, that is the file that matters most.

## Conclusion

Dream-RSI is worth reading as a proposal rather than as a tool, because the mechanism is described precisely enough to argue with: a replay simulator evaluates a policy over a finite pool of recorded trees, which is cheaper than rerunning and narrower than a world model, and the paper is where that trade is defended. What is not available is any code, any discovered programs and any reproduction script, so the empirical claims cannot be checked from the repository. Wait for the release, and read ATTRIBUTION and licence terms carefully once it lands, since both are currently unknown for this repository.

## FAQ

### What is Dream-RSI?

A research project on recursive self-improvement that treats completed discovery histories as replay simulators. An alternative exploration policy traverses a recorded decision tree differently, reading saved outcomes rather than rerunning the discovery agent or the evaluator, which the document compares to dreaming inside a world model.

### Is the Dream-RSI code available?

Not yet. A note in the README says code is being prepared for release, and the release plan table lists the arXiv posting as in progress with the discovered programs, the full codebase and the reproduction scripts all marked as being prepared. The repository holds only a gitignore, a citation file, the README, an assets directory and a papers directory.

### How does Dream-RSI evaluate an exploration policy?

By replaying recorded traces. A finished discovery process records a tree of exploration decisions and their code-execution outcomes, and a candidate policy traverses that tree differently, choosing different subsets of branches in different orders with different parallel groupings and stopping decisions, with no rerunning of the discovery agent or the evaluator.

### What results does Dream-RSI report?

Across algorithm engineering, mathematical optimisation and GPU kernel engineering, an improvement in both discovery effectiveness and efficiency in several settings. No baseline, metric or per-setting figure appears in the text of the README, and the statistics graphic is an image, so the numbers have to come from the paper.

### Who publishes Dream-RSI?

Seventeen authors across Google, Google DeepMind, the University of Maryland College Park and the University of Virginia, with two corresponding authors. The repository carries a CITATION.cff file and a BibTeX block in the README, and the latter already cites arXiv:2609.14858 while the release plan still marks the arXiv posting as in progress.

## Sources

- [Issues](https://github.com/zhengkid/Dream-RSI/issues)
- [README](https://github.com/zhengkid/Dream-RSI/blob/main/README.md)
- [zhengkid/Dream-RSI on GitHub](https://github.com/zhengkid/Dream-RSI)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/zhengkid-dream-rsi
