# sapientinc/HRM: a 27M-parameter reasoning model you train from scratch

> Hierarchical Reasoning Model is a recurrent architecture with two coupled modules and no chain-of-thought data. The repository is a training and evaluation harness for Sudoku, mazes and ARC, not a chat model.

**sapientinc/HRM** — Hierarchical Reasoning Model Official Release

- Repository: https://github.com/sapientinc/HRM
- Website: https://sapient.inc
- Stars: 12,641 · Forks: 1,822
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/sapientinc-hrm

## The problem HRM targets: reasoning without chain-of-thought traces

Most current reasoning systems are language models that write intermediate steps. The README frames that approach as having three costs: brittle task decomposition, extensive data requirements, and high latency. HRM is the authors' answer to all three at once. It is a recurrent architecture that performs sequential reasoning in a single forward pass, with no explicit supervision of the intermediate process. The README states the model has 27 million parameters and reaches its results with 1000 training samples, without pretraining and without CoT data. The headline tasks are complex Sudoku, optimal path finding in large mazes, and the Abstraction and Reasoning Corpus. That scope matters. HRM is not positioned as a general assistant. It is a research artifact for structured puzzle domains where a correct answer is checkable, and the repository reflects that: the top level holds pretrain.py, evaluate.py, a dataset/ directory of puzzle builders, and a puzzle_visualizer.html. If your problem is open-ended text generation, this is the wrong project, and nothing in the README suggests otherwise.

## Two recurrent modules, one forward pass: how the HRM architecture is wired

The mechanism the README describes is two interdependent recurrent modules running at different timescales. A high-level module does slow, abstract planning. A low-level module does rapid, detailed computation. Their interaction is what produces computational depth without a long chain of emitted tokens. Training configuration exposes this directly. In the full Sudoku-Hard command the README gives, arch.L_cycles=8 sets the number of cycles and arch.halt_max_steps=8 sets a halting bound, while arch.pos_encodings=learned selects learned positional encodings and arch.loss.loss_type=softmax_cross_entropy selects the loss. Those keys tell you the architecture is configured through Hydra and OmegaConf rather than hardcoded: pretrain.py takes overrides on the command line, and requirements.txt lists hydra-core, omegaconf, pydantic and argdantic, which is the config stack doing that work. The repository layout matches the story. models/ holds the architecture, config/ holds the Hydra configs whose keys appear in the commands, and puzzle_dataset.py feeds batches. The practical consequence is that you tune reasoning depth by changing config values rather than by prompting for more steps.

## Installing sapientinc/HRM: CUDA 12.6, FlashAttention, then requirements

There is no pip package and no wheel. Installation is a source checkout plus a CUDA toolchain, and the README's prerequisites section is explicit that the repo needs CUDA extensions to be built. The first block installs the CUDA 12.6 toolkit and the matching PyTorch build, then the build helpers the extensions need. Expect a large download and a silent installer; the export of CUDA_HOME is what later build steps read.

```bash
CUDA_URL=https://developer.download.nvidia.com/compute/cuda/12.6.3/local_installers/cuda_12.6.3_560.35.05_linux.run
wget -q --show-progress --progress=bar:force:noscroll -O cuda_installer.run $CUDA_URL
sudo sh cuda_installer.run --silent --toolkit --override
export CUDA_HOME=/usr/local/cuda-12.6
pip3 install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu126
pip3 install packaging ninja wheel setuptools setuptools-scm
```

FlashAttention comes next, and the README splits the instructions by GPU generation. On Hopper you clone the repository and build the hopper directory from source. On Ampere or earlier you install the published package. This is the step most likely to fail, because it compiles against your CUDA and PyTorch versions.

```bash
git clone git@github.com:Dao-AILab/flash-attention.git
cd flash-attention/hopper
python setup.py install
```

With the toolchain in place, the Python dependencies come from requirements.txt, and experiment tracking needs a Weights & Biases login because the training scripts log there.

```bash
pip install -r requirements.txt
wandb login
```

After that, the quickest real use is the Sudoku demo. The first command builds a 1000-example dataset with 1000 augmentations into data/sudoku-extreme-1k-aug-1000; the second starts training. The README gives roughly 10 hours on an RTX 4070 laptop GPU for this configuration.

```bash
python dataset/build_sudoku_dataset.py --output-dir data/sudoku-extreme-1k-aug-1000 --subsample-size 1000 --num-aug 1000
OMP_NUM_THREADS=8 python pretrain.py data_path=data/sudoku-extreme-1k-aug-1000 epochs=20000 eval_interval=2000 global_batch_size=384 lr=7e-5 puzzle_emb_lr=7e-5 weight_decay=1.0 puzzle_emb_weight_decay=1.0
```

You should see training progress in W&B. For Sudoku and maze runs, the README points at eval/exact_accuracy as the metric to watch.

## Training from scratch versus loading the published checkpoints

You do not have to train anything to inspect HRM's behaviour. The README links three checkpoints on Hugging Face: HRM-checkpoint-ARC-2, HRM-checkpoint-sudoku-extreme for 9x9 extreme Sudoku from 1000 examples, and HRM-checkpoint-maze-30x30-hard for 30x30 hard mazes from 1000 examples. Loading them means using evaluate.py, and requirements.txt includes huggingface_hub for that path. The README's evaluation section is short: check eval/exact_accuracy in W&B, and for ARC run the evaluation script and then finish with the arc_eval.ipynb notebook.

```bash
OMP_NUM_THREADS=8 torchrun --nproc-per-node 8 evaluate.py checkpoint=<CHECKPOINT_PATH>
```

Training yourself is where the cost sits. The README's full-scale commands assume an 8-GPU setup. ARC-1 and ARC-2 small-sample runs are quoted at roughly 24 hours each, with a note that a checkpoint after 8 hours is often sufficient for ARC-2. Maze 30x30 Hard takes about an hour at 8 GPUs, and Sudoku Extreme about 10 minutes. The Sudoku-Hard full run is about 2 hours and changes more than the dataset path: it uses global_batch_size=2304, lr=3e-4, lr_min_ratio=0.1, and a different weight decay, so the small-sample recipe does not transfer unchanged. The README also notes that small-sample learning shows accuracy variance of around plus or minus 2 points, which means a single run is a weak basis for comparison.

## Where HRM breaks: overfitting, variance and the missing inference story

The README's Notes section is the most useful part for anyone deciding whether to depend on this. It states that for the Sudoku-Extreme 1000-example dataset, late-stage overfitting may cause numerical instability during training and Q-learning, and advises early stopping once training accuracy approaches 100 percent. That is a real failure mode, not a tuning detail: the same small-sample regime that makes HRM cheap to train is the regime where it destabilizes. The plus or minus 2 point variance on small-sample learning compounds this, because early stopping decisions are being made on noisy numbers.

The second limitation is scope. Every dataset builder in the repository produces puzzles: Sudoku, mazes, ARC. There is no text corpus loader, no tokenizer, and no serving path in the README. The RELATED SEARCHES list contains phrases like HRM-Text-1B and sapient inc/hrm-text, but nothing in this repository's README, file listing or requirements.txt describes a text model, and no such checkpoint is linked. Treat those queries as pointing at something this repository does not document. The third gap is operational: the README documents training and evaluation, and says nothing about exporting to ONNX, quantizing, or serving the model behind an API. If you need inference at low latency in a service, you would be building that layer yourself from models/.

## How HRM differs from a chain-of-thought language model

The natural comparison is a chain-of-thought language model, and the difference is architectural rather than a matter of scale. A CoT system emits intermediate tokens, so its reasoning depth is bounded by context length and its latency grows with the number of tokens generated. HRM keeps the intermediate computation inside the recurrent modules and produces an answer in one forward pass, which is why the README can claim computational depth without intermediate supervision. The trade is that HRM has no language interface and no pretrained knowledge: it learns a task family from scratch on 1000 examples. A CoT model arrives with broad priors and handles tasks you never trained on, at the cost of the data, latency and decomposition problems the README lists. Neither replaces the other. If your inputs are natural language and your outputs are natural language, a CoT language model is the right tool and HRM is not. If your problem is a structured puzzle with a verifiable answer and you want a 27M-parameter model you can train on a single laptop GPU, HRM is addressing exactly that case.

## Licence, maintenance and what upgrading actually costs

The repository is Apache-2.0, which permits commercial use and modification and includes an explicit patent grant. The checkpoints are hosted separately on Hugging Face, so confirm the terms attached to each checkpoint rather than assuming the repository licence covers them. Nothing here is legal advice.

On maintenance: the repository is not archived, and the last push was on 2026-03-31. There are no retrieved releases, so there is no versioned upgrade path to follow. Upgrading means pulling main and re-checking your environment, and the environment is the expensive part. The README pins CUDA 12.6 and splits FlashAttention instructions by GPU generation, so a CUDA or PyTorch bump can invalidate a working build. Hydra configs add a second axis: the commands in the README override keys such as arch.L_cycles and arch.halt_max_steps, and a config rename upstream breaks those overrides silently at parse time. Budget for a rebuild of the CUDA extensions, not just a git pull.

## Conclusion

Adopt sapientinc/HRM if you want to study or reproduce a small recurrent reasoning architecture on puzzle domains and you have a CUDA GPU, PyTorch, and the patience to build FlashAttention before anything runs. Skip it if you need a hosted endpoint, a general-purpose assistant, or a CPU-only install: the README requires CUDA extensions, and the shipped checkpoints are for ARC, Sudoku and maze puzzles rather than open-ended text. Before committing, verify three things against your own hardware: that the CUDA 12.6 toolkit and the FlashAttention variant for your GPU generation build cleanly, that the Sudoku demo fits your GPU memory at global_batch_size=384, and that the evaluation path you need (W&B eval/exact_accuracy for Sudoku and maze, arc_eval.ipynb for ARC) produces numbers you can compare with the checkpoints on Hugging Face.

## FAQ

### What is the AI model HRM?

HRM is the Hierarchical Reasoning Model from sapientinc, a recurrent architecture with a high-level module for slow planning and a low-level module for fast computation. The README states it has 27 million parameters and is trained without pretraining or chain-of-thought data, reaching strong results on Sudoku, mazes and ARC from 1000 samples.

### What is HRM in AI?

In this repository, HRM refers to a brain-inspired recurrent architecture that executes sequential reasoning in a single forward pass instead of emitting intermediate reasoning tokens. The README contrasts it with chain-of-thought LLMs, which it describes as suffering from brittle task decomposition, extensive data requirements and high latency.

### What is an HRM model?

It is a model built from two interdependent recurrent modules operating at different timescales, configured through Hydra keys such as arch.L_cycles and arch.halt_max_steps. The repository ships training and evaluation scripts plus three Hugging Face checkpoints for ARC-AGI-2, Sudoku 9x9 Extreme and Maze 30x30 Hard.

## Sources

- [Issues](https://github.com/sapientinc/HRM/issues)
- [License: Apache-2.0](https://github.com/sapientinc/HRM/blob/main/LICENSE)
- [Project website](https://sapient.inc)
- [README](https://github.com/sapientinc/HRM/blob/main/README.md)
- [sapientinc/HRM on GitHub](https://github.com/sapientinc/HRM)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/sapientinc-hrm
