Model or dataset
trevin-creator/autoresearch-mlx avatar
trevin-creator/autoresearch-mlx

autoresearch-mlx: a fixed five minute research loop, and why it needs a referee

Apple Silicon (MLX) port of Karpathy's autoresearch — autonomous AI research loops on Mac, no PyTorch required.

1,840 stars366 forksPythonMIT

At a glance

What is it?
trevin-creator/autoresearch-mlx ports Karpathy's autonomous experiment loop to Apple Silicon through MLX, with one mutable train.py, one metric and a git keep-or-revert rule. The interesting contribution is rigor.py, which admits the loop's own numbers are optimistically biased.
Who is it for?
The thing to take from autoresearch-mlx is not the numbers in its tables but the split between the loop and the judge. The loop is a thin, readable protocol that an agent can drive, and rigor.py is a separate process that refuses to be fooled by a 0.03 wobble, never edits train.py, never touches git and never re-scores a file it has already scored.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 97 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 20, 2026, and from our analysis. They are not legal advice.

Editorial analysis

One mutable file, one metric, one wall clock

The protocol is deliberately narrow. One mutable `train.py`, one metric called `val_bpb`, a fixed five minute training budget, and a keep-or-revert decision made through git. The README credits Andrej Karpathy for the core idea of fixed-time autonomous research loops controlled through `program.md`, and this port keeps the same basic rules rather than extending them.

That constraint set is what makes the thing automatable. A coding agent such as Claude Code is pointed at `program.md`, and the loop it runs is: edit `train.py`, run a fixed-budget experiment, read `val_bpb`, keep the change if it wins, revert it if it loses, and repeat. There is no hyperparameter search space to specify and no schedule to tune, because the schedule is the point.

The file tree explains how small this is. Nine tracked entries: `README.md`, `LICENSE`, `.gitignore`, `pyproject.toml`, `uv.lock`, `program.md`, `prepare.py`, `train.py`, `rigor.py`, plus `results.tsv`. Of those, the README assigns clear roles. `prepare.py` handles data prep, the tokenizer, the dataloader and evaluation, and is to be treated as fixed. `train.py` holds the model, the optimizer and the training loop, and is the only file the agent edits. `program.md` is the experiment protocol. `results.tsv` is the logged experiment history.

So a large fraction of the design is in what is forbidden. Keeping `prepare.py` fixed means evaluation cannot drift to flatter a change, which is the failure mode that quietly ruins most automated tuning setups.

What the committed results file actually shows

The public `results.tsv` covers the initial hardware-local walk, and it is short enough to read in full. Four commits, all kept, running from the default baseline down to 1.807902:

| Commit | val_bpb | Description | |---|---:|---| | `383abb4` | 2.667000 | baseline (AdamW, default config) | | `909dd59` | 2.588904 | halve total batch size to `2^16` | | `4161af3` | 2.533728 | increase matrix LR to `0.04` | | `5efc7aa` | 1.807902 | reduce depth from `8` to `4` |

The distribution of the gains is the interesting part. The first two steps, batch size and learning rate, together buy about 0.13 bpb. The last step, cutting depth from 8 to 4, buys 0.73, which is more than five times as much. The README draws the conclusion plainly: with a fixed five minute wall clock, smaller faster-training models can beat larger ones simply by fitting more optimizer steps into the budget.

That reframes what the metric means. `val_bpb` is not measuring model quality in the usual sense, where a better model wins. It measures quality reachable inside five minutes, so the winning move is usually to buy steps rather than to buy parameters. On a unified-memory Apple Silicon machine with no separate accelerator, that bias is amplified, because per-step cost scales with the work you ask for and the budget is fixed.

The README also quantifies the iteration cost honestly, at roughly 6 to 7 minutes per experiment, being 5 minutes of training plus compile and eval overhead. So the loop performs perhaps eight experiments an hour, and the five minute figure in the protocol is not the wall clock cost of a hypothesis test.

The loop flatters itself, and the repository says so

The most valuable section in this README is the one explaining why the headline loop is statistically broken.

A single five minute run is noisy, and re-running the same unmodified `train.py` moves `val_bpb` by about 0.03. Deciding keep or discard on one run below that threshold is chasing noise. Worse, the loop only keeps a run when it dips below the current best, so the recorded curve is an optimistic running minimum, and it regresses when you re-evaluate honestly.

`rigor.py` is the fix, and the design choice is that it stays out of the way. It runs a few seeds and keeps a change only if it beats the current best with high confidence under a bootstrap. It fails a clear loser fast, after one run. And it never re-scores an identical `train.py`, which prevents the specific waste of burning seeds to re-confirm something already measured.

The interface is four subcommands:

bash
uv run rigor.py run "halve the batch size"   # score train.py vs best (3 seeds)
uv run rigor.py run "..." --seeds 5 --confidence 0.9
uv run rigor.py best                          # current best
uv run rigor.py log                           # every scored config

The constraints it accepts are worth reading as policy. Seeds and a confidence threshold are flags rather than constants, so the rigor level is a per-experiment decision, and a 0.9 threshold is reachable at five seeds.

Most importantly, the README states what `rigor.py` never does: it never edits `train.py`, never touches git, and never changes `evaluate_bpb`. It only decides, and it writes samples to `rigor_ledger.jsonl`. So the judge cannot contaminate the thing it judges, and the git keep-or-revert step remains the only writer of `train.py`. If you take one design idea from this repository, take this one.

Three machines and three different answers

The longer runs section is where the project stops being a demo. Overnight runs on the working MLX port pushed much lower, and the table reports current best, starting point and the repeated wins for each machine:

| Machine | Current best | Starting point | |---|---:|---:| | M4 Max #1 | 1.294526 | 1.596971 | | M4 Max #2 | 1.330509 | 1.807902 | | Mac Mini (long run) | 1.353329 | 1.922472 |

The two M4 Max machines converge on a similar stack: AdamW-only, low matrix learning rate, a 3x MLP, no logit cap and moderate weight decay on the first, and a leaner batch, long anneal, SiLU, lower regularization and again no logit cap on the second. The common thread is the absence of the logit cap, which both runs found worth removing.

The Mac Mini is included precisely because it disagreed. Its winner stack leans the other way: Muon as the optimizer, sharper attention, a smaller MLP and a lower scalar learning rate. The README's reading is that on smaller Apple Silicon hardware the strongest changes leaned toward more aggressive step-efficiency wins, and that later transfer tests showed some of those Mac Mini findings did not carry cleanly onto the Max baseline.

That negative result is the most valuable line in the table. It means the loop's answers are conditioned on the hardware, and a recipe that wins on one machine can be worthless on another. This is a property of fixed-wall-clock search in general, not a flaw in the port, but it does mean the committed `results.tsv` should be read as a record of one machine rather than a set of portable hyperparameters.

Setup, dependencies and the one-time preparation cost

Requirements are narrow: an Apple Silicon Mac, Python 3.10 or newer, and `uv`. The quick start is four steps.

bash
# install uv if needed
curl -LsSf https://astral.sh/uv/install.sh | sh

# install dependencies
uv sync

# one-time data + tokenizer prep
uv run prepare.py

# run one 5-minute training experiment
uv run train.py

`uv sync` against a committed `uv.lock` is the right choice here, and the dependency set explains why. From `pyproject.toml`, version 0.1.0, requiring Python `>=3.10,<3.14`, the direct dependencies are `mlx>=0.30.0`, `numpy>=2.2.6`, `pyarrow>=21.0.0`, `regex>=2025.7.34`, `requests>=2.32.0`, `rustbpe>=0.1.0` and `tiktoken>=0.11.0`.

Two details stand out. `rustbpe` alongside `tiktoken` means the tokenizer path has a Rust implementation available, which matters when tokenization is inside your iteration loop. And the absence of `torch` is the entire point of the port, since native MLX training on unified memory removes the PyTorch and CUDA requirement. The upper bound on Python is `<3.14`, so the project will refuse to install on the newest interpreter rather than fail obscurely.

`prepare.py` is the step people skip and then regret. It does the data prep, builds the tokenizer, sets up the dataloader and defines `evaluate_bpb`, and it needs a network connection given `requests` in the dependency list. Because the evaluation budget is smaller than upstream, the README flags this as a deliberate trade: reduced for faster iteration on Apple Silicon while keeping the same `evaluate_bpb` interface. Any score you record is therefore comparable to other runs in this repository, and not directly comparable to upstream numbers.

Where the port diverges from upstream, and what is not shipped

The README lists five differences from upstream, and one of them is a caveat rather than a feature.

The first is MLX instead of PyTorch and CUDA, giving native Apple Silicon training with unified memory. The second is an AdamW-only public path: the public `train.py` keeps the default simple, and the Muon variant that won on the Mac Mini was explored in the working port but is deliberately not exposed as a public default here. That is a defensible choice for a reference repository, and it means the Muon result in the table is not something you can reproduce from a clean clone.

The third is the smaller evaluation token budget, already covered. The fourth is the 6 to 7 minute wall clock per experiment against the nominal 5. The fifth is the one to read twice: MFU reporting is a placeholder, because there is no Apple Silicon equivalent to the H100 FLOPs reference used upstream.

That last item is the real gap. Model Flops Utilization is how you tell a compute-bound run from a badly configured one, and without it you cannot tell whether the loop is finding genuine architectural wins or simply filling the budget with more steps. It is the first thing you would want to add.

On shipping, GitHub reports no releases for this repository, and the project sits at version 0.1.0 in `pyproject.toml` with the default branch as the only thing to install. GitHub also reports no recognised topics, so there is no ecosystem discovery path. The last push landed on 2026-07-02, so the branch is being worked on, and anyone following the loop by cloning should expect the file set to shift under them.

Editorial conclusion

The thing to take from autoresearch-mlx is not the numbers in its tables but the split between the loop and the judge. The loop is a thin, readable protocol that an agent can drive, and rigor.py is a separate process that refuses to be fooled by a 0.03 wobble, never edits train.py, never touches git and never re-scores a file it has already scored. That separation is what keeps the experiment history honest, and it is reusable well beyond this repository. Two practical cautions. The committed results.tsv stops at 1.807902 while the far better figures in the README come from longer runs on hardware you may not own, and the Mac Mini winners are explicitly reported as not transferring to the Max baseline. And the project has no releases at all, so you install a moving main branch at version 0.1.0.

Frequently asked questions

What does autoresearch-mlx do and what does it run on?

It is an Apple Silicon port of Karpathy's autoresearch that performs fixed-time autonomous research loops on a Mac with no PyTorch dependency. The protocol keeps one mutable train.py, one metric called val_bpb, a fixed five minute training budget and a keep-or-revert rule applied through git. It requires an Apple Silicon Mac, Python 3.10 or newer, and uv, and trains natively through MLX on unified memory.

How do I set up the project and run a first experiment?

Install uv, then run `uv sync` to install dependencies from the committed uv.lock, run `uv run prepare.py` once for data prep, tokenizer and evaluation setup, and then run `uv run train.py` for one five minute training experiment. Expect roughly 6 to 7 minutes of wall clock per experiment once compile and evaluation overhead are counted, since the five minute figure covers training alone.

What is rigor.py and why does it exist?

rigor.py is the statistical gate on the keep or discard decision. Re-running the same unmodified train.py moves val_bpb by about 0.03, and because the plain loop only keeps runs that dip below the current best, its recorded curve is an optimistic running minimum. rigor.py instead runs several seeds, keeps a change only when a bootstrap gives high confidence that it beats the best, and drops clear losers after a single run. It never edits train.py, never touches git and never changes evaluate_bpb.

Do the hyperparameters in the results table transfer between machines?

Not reliably. The two M4 Max machines converged on similar winner stacks, but the Mac Mini run found a meaningfully different one built around Muon, sharper attention, a smaller MLP and a lower scalar learning rate. The README states that later transfer tests showed some Mac Mini findings did not carry cleanly onto the Max baseline, so a recipe found on one Apple Silicon machine should be re-established rather than copied.

Official sources

  1. Issues
  2. License: MIT
  3. README
  4. trevin-creator/autoresearch-mlx on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/trevin-creator-autoresearch-mlx.svg)](https://hysenlabs.com/projects/trevin-creator-autoresearch-mlx)