# openai/parameter-golf: the 16MB language model challenge and its leaderboard

> Parameter Golf is OpenAI's Model Craft Challenge: train the best language model that fits in a 16MB artifact and trains in under 10 minutes on 8xH100s, scored by compression on FineWeb validation. The repository is mostly a scoring harness plus a long history of merged submissions.

**openai/parameter-golf** — Train the smallest LM you can that fits in 16MB. Best model wins!

- Repository: https://github.com/openai/parameter-golf
- Stars: 5,199 · Forks: 3,256
- Language: Python
- License: MIT
- Published: 2026-09-23 · Updated: 2026-09-23 · Language: en
- Canonical page: https://hysenlabs.com/projects/openai-parameter-golf

## What Parameter Golf actually asks you to optimize

The objective is not loss alone. It is loss under two hard constraints: the trained model must fit into a 16MB artifact, and a leaderboard submission must train in under 10 minutes on 8xH100s. Evaluation is compression on the FineWeb validation set, measured in bits per byte and described in the README as tokenizer-agnostic. That last detail matters more than it looks. Because the score is bytes rather than tokens, a submission is free to invent its own tokenizer, and the README explicitly lists novel tokenizers among the directions it expects people to explore.

The README frames this as L(N) optimization: push loss down with N, the parameter count, fixed, while data, compute, steps and architecture stay unconstrained in principle. That places it in the same family as NanoGPT Speedrunning, which optimizes for time at a fixed loss, and NanoGPT Slowrun, which optimizes for loss at a fixed dataset size. The audience is people who find that framing motivating: the README says the challenge is inspired by elite mathematics and programming competitions and that a small cohort of early-career researchers is planned, targeting current undergraduates and recent graduates. Completing the participant form is stated as optional.

## How scoring works, and why the leaderboard reads like a git log

There is no package to import. The repository is a working record: train_gpt.py, a separate train_gpt_mlx.py for Apple silicon, a data/ directory, scripts/, results/, records/, and measure_variance_ratio.py. The leaderboard table in the README is the primary artifact, and every row is a pull request. Read the top row and you get the whole culture: a score of 1.0565, an author handle, a summary that names the parent PR, the exact configuration change, and a statistical claim against the previous record.

The summaries are unusually specific. They reference named techniques such as SmearGate, LQER, SparseAttnGate, PolarNS, FusedCE, CaseOps, AsymLogit, AWQ-lite GPTQ, quant gates, depth recurrence and test-time training, and they carry hyperparameter values like QK_GAIN_INIT=5.25, MIN_LR=0.1 and GPTQ_CALIBRATION_BATCHES=32. Most rows report a three-seed or five-seed mean plus a p-value against the prior record. That is the mechanism that keeps the board honest: a single lucky seed is not enough, and the README's own table shows submissions being rerun for compliance and reproduction. The cost is that the frontier is a chain of small deltas, and understanding any one entry means reading the PR it builds on.

## Installing Parameter Golf and running a first training pass

The README does not document an installation procedure, a supported Python version, or a packaging story. What the repository does provide is requirements.txt at the top level, which is the only concrete dependency list available:

```bash
pip install -r requirements.txt
```

That file pins typing-extensions==4.15.0 and otherwise leaves numpy, tqdm, torch, huggingface-hub, kernels, setuptools, datasets, tiktoken and sentencepiece unpinned. Unpinned torch in a repository whose scores depend on numerical reproducibility is a real friction point: a different torch build can change kernel selection, and the leaderboard rows are explicitly seed-sensitive.

The training entry point is train_gpt.py at the repository root. The README does not list its flags, so the honest first step is inspection rather than a copied command:

```bash
python train_gpt.py
```

There is also train_gpt_mlx.py for MLX, which the README does not describe. If you are on Apple silicon, that file is the one to read. Beyond this, the README points at the data/ directory for the FineWeb validation setup and at scripts/ and results/ for the surrounding tooling. It does not state how to download the dataset, so plan on reading data/ and the evaluation script before assuming a command exists. The README does note that OpenAI is sponsoring $1,000,000 in compute credits and links a credit request form, with the instruction to pick an appropriate level and submit using an email tied to an OpenAI or ChatGPT account.

## The 16MB artifact and the ten-minute cap are the whole difficulty

Two constraints do the work here, and both are enforced outside the model code. The 16MB limit is a property of the serialized artifact, so quantization, parameter tying, low-rank structure and compression schemes all become first-class design choices rather than afterthoughts. The leaderboard summaries bear this out: GPTQ calibration, AWQ-lite mixed quantization, per-group lrzip compression and quant-gate scaling all appear as scoring-relevant changes. A model that trains beautifully and serializes to 17MB scores nothing.

The ten-minute cap on 8xH100s is described in the README as a deliberate compromise. The README states plainly that it would ideally allow arbitrary compute, and that the cap exists so the challenge does not become inaccessibly expensive. Submissions that exceed the cap are not rejected outright; the README says they belong in a Non-record Submissions section, framed as pushing the infinite frontier of parameter-limited performance. That is a sensible split, but it also means the leaderboard number is a joint property of the model and the hardware budget, not a pure capability measure. A 1.0565 at ten minutes and a better loss at sixty minutes are not comparable, and the README does not pretend otherwise.

## Where Parameter Golf is the wrong tool

Do not reach for this if you want a training framework. There is no documented API, no versioned release, and no install guide beyond a requirements file with unpinned core dependencies. The last push to the repository was on 2026-05-04, and the README states the challenge ran from March 18th to April 30th, so the competitive window described in the README has closed. Nothing here suggests the code is being maintained as a general-purpose library.

The evaluation setup is also a poor fit for most production questions. Bits per byte on FineWeb validation is a compression metric, and it rewards tokenizer design in ways that a downstream task will not. A submission that wins by adopting a lossless CaseOps transform with byte-sidecar BPB accounting, as one leaderboard row describes, is optimizing for the scoring rule as much as for language modeling. That is the point of a competition, and it is a liability if you copy a winning configuration into a product. The three-seed and five-seed averaging in the leaderboard rows is a signal worth heeding: differences at the fourth decimal place are small enough that the README's own table reports p-values to argue they are real.

## How it differs from NanoGPT Speedrunning

NanoGPT Speedrunning, which the README names as the inspiration, fixes the target loss at 3.28 FineWeb validation loss and asks how fast you can get there. The constraint is time. Parameter Golf inverts that: the time budget is fixed at ten minutes on 8xH100s, and the loss is what you minimize, subject to an artifact-size ceiling. The practical consequence is that the two challenges reward different work. Speedrunning rewards step-time engineering, fused kernels and learning-rate schedules that converge fast. Parameter Golf rewards anything that buys loss per byte of serialized model, which is why its leaderboard is full of quantization calibration, parameter tying, test-time training and tokenizer changes rather than pure throughput work.

The README also places NanoGPT Slowrun alongside these, optimizing loss under a fixed dataset size. Read together, the three are the same optimization viewed along different axes: L(T), L(N) and L(D). If your interest is the parameter axis specifically, Parameter Golf is the one with a public leaderboard and a compression-based score. If your interest is wall-clock training efficiency, Speedrunning is the better-studied target.

## Licence and the cost of tracking the frontier

The repository is MIT licensed, with a THIRD_PARTY_NOTICES.md at the top level. MIT is permissive, so reusing train_gpt.py or pieces of a submission in your own work carries few obligations beyond preserving the notice. The notices file exists for a reason: the leaderboard stacks pull in techniques from many PRs, and the dependency list includes kernels, tiktoken and sentencepiece, each with its own terms. Check THIRD_PARTY_NOTICES.md before redistributing a merged configuration, and note that this is a description of what the repository contains, not legal advice.

The upgrade cost is the more interesting question. There are no releases, so there is no upgrade path in the usual sense: you track main or you pin a commit. Because scores depend on seed and on library versions, and because requirements.txt pins only typing-extensions, reproducing a leaderboard row means reconstructing its environment from the PR, not from the repository. The README's leaderboard rows link to specific PRs and, in at least one case, a specific commit hash, which is the closest thing to a reproducibility contract on offer.

## Conclusion

Adopt this if you already know how to train a small transformer and want a hard, externally scored constraint to optimize against; the leaderboard entries are public PRs with reproducible seeds, which is a genuinely useful learning artifact. Do not adopt it if you need a maintained training library, a stable API, or something with documented installation and versioning: the README does not document a supported install path, and the last push was on 2026-05-04, so treat the code as a snapshot of a finished competition rather than an evolving tool. Before spending compute, verify the artifact-size accounting rules in the challenge page and the evaluation script, because the 16MB limit is enforced by the scoring pipeline, not by an assertion in train_gpt.py.

## FAQ

### What is OpenAI's Parameter Golf challenge?

It is the OpenAI Model Craft Challenge, a competition to train the best language model that fits in a 16MB artifact and trains in under 10 minutes on 8xH100s, scored by compression on the FineWeb validation set in bits per byte. The README says it runs from March 18th to April 30th and is inspired by the NanoGPT Speedrunning challenge.

### What is parameter golf?

Parameter Golf is a challenge framed by the README as L(N) optimization: minimize loss with the parameter count fixed, leaving data, compute, steps and architecture unconstrained in principle. Leaderboard submissions are limited to ten minutes on 8xH100s to keep the challenge affordable.

### What did Parameter Golf teach us?

The README presents the challenge as a test of tackling unfamiliar problems with creativity and rigor, and expects the parameter constraint to push people toward unusual architectures, compression schemes and other creative submissions. It does not publish a retrospective of lessons learned, so what the leaderboard demonstrates is a chain of incremental, seed-averaged improvements.

## Sources

- [Issues](https://github.com/openai/parameter-golf/issues)
- [License: MIT](https://github.com/openai/parameter-golf/blob/main/LICENSE)
- [openai/parameter-golf on GitHub](https://github.com/openai/parameter-golf)
- [README](https://github.com/openai/parameter-golf/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/openai-parameter-golf
