autoresearch gives an agent one file and a five-minute budget to improve it
GitHub describes it as AI agents running research on single-GPU nanochat training automatically. The repository metadata lists Python as its primary language. This article stays within the project description and details documented in the GitHub repository README.
At a glance
- What is it?
- autoresearch is a deliberately tiny LLM training setup where an agent edits a single file, runs a fixed five-minute experiment, and keeps the change only if validation bits per byte improved. The smallness is the design, and the five-minute clock is what makes a hundred overnight runs comparable to each other.
- Who is it for?
- autoresearch suits someone with one spare NVIDIA GPU who wants to watch what an agent does to a real training loop rather than read about it, and who is comfortable that the search is bounded by design. It does not suit a machine without a CUDA GPU, and its numbers do not transfer to other hardware, so treat a val_bpb figure as a property of your own card and your own five minutes.
- Can I use it commercially?
- Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
- Is it still maintained?
- Activity is slowing. The repository last received commits 6 months ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
DEEP OPEN-SOURCE ANALYSIS
Three files carry the project, and only train.py is the agent's to edit
The repository is small on purpose, and the division of labour is explicit. `prepare.py` holds fixed constants, the one-time data prep that downloads training data and trains a BPE tokenizer, plus the runtime utilities for the dataloader and evaluation. It is marked as not modified. `train.py` is the single file the agent edits: it contains the full GPT model, the optimizer combining Muon and AdamW, and the training loop, and everything in it is fair game, from architecture to hyperparameters to batch size. `program.md` holds the baseline instructions for one agent, and that is the file a human edits.
The point of the split is scope. Restricting the agent to one file keeps the change reviewable, because a diff against `train.py` is a diff you can read in a minute rather than a changeset spread across a training library. The README makes the same point from the other side: as a researcher you are not touching the Python files at all, you are programming the `program.md` Markdown that supplies context to the agent and constitutes the research organisation.
The training code itself is a simplified single-GPU implementation of nanochat, and `pyproject.toml` describes the package as an autonomous pretraining research swarm at version 0.1.0.
The five-minute clock is what makes a hundred overnight runs comparable
Training runs for a fixed five-minute wall-clock budget regardless of what is inside the loop, and startup and compilation are excluded from that clock. The single score is `val_bpb`, validation bits per byte, where lower is better. The choice of that metric is the load-bearing decision: it is vocab-size independent, so a change to the tokenizer or the embedding width does not quietly move the number and make two runs look different when only the reporting changed.
With a fixed budget you get roughly 12 experiments an hour and roughly 100 while you sleep, and the agent keeps or discards each edit based on whether the metric improved. The README gives two reasons for the fixed clock. First, experiments stay directly comparable whatever the agent changed, whether that is model size, batch size or architecture. Second, the search converges on the best model that fits the budget, which for a fixed hardware is a real and useful target.
The stated cost is that results stop being portable. Runs and numbers from your card are not comparable to results from someone else's compute, so a val_bpb figure is a measurement of your hardware under your clock, not a property of the architecture.
A single NVIDIA GPU is the whole platform story, on purpose
The requirements are one NVIDIA GPU, tested on H100, Python 3.10 or newer, and uv. There is no CPU or MPS path, and the reason given is that supporting them would bloat the code, with the author explicitly unsure about taking on that work personally. The escape hatch is the parent nanochat repository, which has wider platform support and shows the solutions that would be needed, such as a Flash Attention 3 kernels fallback, generic device support and autodetection.
For a reader the practical effect is that the interesting advice in the README is addressed to people who do not have an H100. Trying this on a Macbook or any smaller machine means following a list of manual retunes, and the list is specific: use a lower-entropy dataset such as TinyStories, drop `vocab_size` from 8192 towards 4096, 2048, 1024 or a byte-level tokenizer of 256 bytes, lower `MAX_SEQ_LEN` in `prepare.py` as far as 256, raise `DEVICE_BATCH_SIZE` in `train.py` to compensate since tokens per forward and backward pass are the product of the two, and decrease `EVAL_TOKENS` so validation loss is measured on less data.
Note the tension in that list. Three of those five knobs live in `prepare.py`, the file the design says must not be modified. Running on small hardware means editing the one file the agent is not supposed to touch.
The install pulls torch from an explicit CUDA 12.8 index
The quick start is four commands, and the third one is a one-time cost:
# 1. Install uv project manager (if you don't already have it)
curl -LsSf https://astral.sh/uv/install.sh | sh
# 2. Install dependencies
uv sync
# 3. Download data and train tokenizer (one-time, ~2 min)
uv run prepare.py
# 4. Manually run a single training experiment (~5 min)
uv run train.pyRun the last command by hand before involving an agent, because the project uses that manual run as the check that the setup works at all. Dependencies are pinned more tightly than most projects would dare: `torch==2.9.1` is an exact pin, and `pyproject.toml` routes it through an explicit index at the PyTorch CUDA 12.8 wheel URL rather than letting the resolver choose a build.
The consequence is that the failure mode arrives at install time, not at run time. A machine whose driver or CUDA setup does not match that index fails during `uv sync`, before any GPU work begins, and the error is a wheel resolution problem rather than anything about the experiment. The repository also carries a `.python-version` file, so the interpreter is pinned as well as the floor of 3.10.
program.md is the research org's source, and the shipped one is a stub
The README is blunt about what you are actually authoring: the default `program.md` is intentionally kept as a bare bones baseline, and the real work is iterating on it over time to find the version that achieves the fastest research progress, and adding more agents to the mix. In the author's framing, `program.md` is essentially a super lightweight skill.
That reframing has a direct consequence for expectations. Out of the box the agent receives very little guidance, so the quality of an overnight run is a function of how much you wrote into that one Markdown file. A thin baseline does not produce a weak agent so much as an unguided one, and the metric will still move, because 100 edits at five minutes each will find something. What it will not do is encode your judgement about which changes are worth trying.
The surrounding material is a research sketch rather than documentation. Longer term plans live in a parent repository and two posts, and `analysis.ipynb` sits at the top level for looking at results. The last push to this repository was on 2026-03-26, it has no GitHub releases, and no license is recorded for it, which is worth knowing before you build anything on top.
The instructions say to disable all permissions before starting the agent
To run the agent you spin up Claude, Codex or whatever you prefer inside the repository, and the README tells you to disable all permissions, then prompt something like this:
Hi have a look at program.md and let's kick off a new experiment! let's do the setup first.Take that instruction seriously rather than as a convenience, because the combination is what defines the risk. The agent holds unrestricted write access to the repository and rewrites `train.py` unsupervised, repeatedly, for as long as you leave it running. Nothing in the loop reviews a diff, and nothing rejects a change on grounds other than the metric. A change that makes training worse is discarded automatically; a change that makes training marginally better while quietly removing a guard, a data check or a checkpoint is kept.
So the boundary here is the agent harness, not the repository. There is no sandbox to configure, no allowlist of paths, and no human in the loop between an edit and a five-minute run. The only rollback available is version control, which makes committing before the first run the single most useful thing you can do, and the reason the single-file restriction is worth respecting even if it is only a convenience for diffing.
Editorial conclusion
autoresearch suits someone with one spare NVIDIA GPU who wants to watch what an agent does to a real training loop rather than read about it, and who is comfortable that the search is bounded by design. It does not suit a machine without a CUDA GPU, and its numbers do not transfer to other hardware, so treat a val_bpb figure as a property of your own card and your own five minutes. Before starting, commit the repository so you have a diff to roll back to, and confirm that `torch==2.9.1` installs from the explicit CUDA 12.8 index on your machine.
Frequently asked questions
What is autoresearch?
autoresearch is a small Python LLM training setup in which an AI agent modifies one file, trains for a fixed five minutes, checks whether validation bits per byte improved, then keeps or discards the change and repeats. The training code is a simplified single-GPU implementation of nanochat.
How do I install autoresearch?
You need a single NVIDIA GPU, Python 3.10 or newer, and uv. Install uv, run uv sync to install dependencies, then run uv run prepare.py once to download training data and train a BPE tokenizer, which takes about two minutes. Dependencies resolve torch 2.9.1 from an explicit CUDA 12.8 index.
How do I use autoresearch with Claude Code?
Point the agent at program.md inside the repository with permissions disabled and ask it to start an experiment. The project's own suggestion is to spin up Claude or Codex in the repo, disable all permissions, and prompt it to look at program.md and kick off a new experiment.
How does autoresearch work?
Three files carry the project. prepare.py holds fixed constants, one-time data prep and runtime utilities and is not modified, train.py holds the model, the Muon and AdamW optimizer and the training loop and is the file the agent edits, and program.md holds the instructions that a human edits and the agent reads.
What is karpathy's autoresearch?
It is described as giving an AI agent a small but real LLM training setup to experiment with autonomously overnight, where the human programs the program.md Markdown rather than the Python files, iterating on it to find the research org code that makes the fastest progress. The last push was on 2026-03-26 and the repository has no GitHub releases.
Community notes