# antidoom rejects the one token that starts a repetition loop

> Antidoom is a narrow tool for a narrow failure: a reasoning model that falls into a repetition loop. It samples completions, finds the token where the loop starts, marks that token as rejected, and trains a LoRA adapter with Final Token Preference Optimization on the plausible alternatives at the same position.

**Liquid4All/antidoom** — Antidoom generates and trains targeted preference data for reducing model repetition loops (doom loops).

- Repository: https://github.com/Liquid4All/antidoom
- Stars: 389 · Forks: 38
- Language: Python
- License: Apache-2.0
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/liquid4all-antidoom

## The training signal is one token, not a rewritten answer

Most preference training for reasoning models starts from gold answers. Antidoom does not. For each prompt it generates a completion, scans for inner repetition, and refines the boundary in token space so the rejected token is the first readable token of the repeated segment. That single token becomes the negative example, and the positives are alternatives sampled from what the model already considered plausible at that same position.

Each FTPO row therefore carries four things: a context prefix ending immediately before the rejected token, that one rejected token, one or more chosen tokens, and metadata about the source prompt, the generated completion, and the detected loop.

The design is adapted from the single-token preference training idea in Antislop, and the project credits both that work and its auto-antislop implementation. What changes is the target: Antislop addresses slop, this addresses runaway repetition during reasoning. The README is explicit that this is a narrow tool for a narrow failure mode, which is worth taking seriously as a scope statement. It is not a general alignment toolkit.

One consequence of the single-token framing is worth holding on to. The adapter is not learning a better answer to the prompt. It is learning that one particular next token, in one particular position, was a mistake.

## Three forces have to line up before a model loops

The project gives a three-part account of why doom loops appear, and each part is separately actionable.

The first is overtrained tokens. Common reasoning tokens such as `Wait`, `So`, `But`, and `Alternatively` can become unusually attractive after heavy synthetic reasoning training. When the model is uncertain, those tokens dominate the next-token distribution without moving the reasoning anywhere.

The second is self-reinforcing context. Once a short sequence appears, the prior context makes it more likely to appear again, so across repeated turns the probability of each token climbs toward certainty. This is the part that makes the failure self-sustaining rather than a one-off slip.

The third is low-temperature sampling. At temperature 0 or near 0, the model keeps taking the highest-probability continuation, so a locally reinforced loop has no natural escape route. That is a direct hint about configuration, and it sits awkwardly with the fact that the documented run uses `--temp 0.01`. Low temperature makes loops easier to detect during generation, which is what the generation stage wants, but it is also the condition described as helping them form.

## The install is uv sync, and the ROCm path needs a different lock file

The NVIDIA path is the default and the short one:

```bash
git clone https://github.com/Liquid4All/antidoom
cd antidoom
uv sync
```

Antidoom needs a GPU runtime for both PyTorch and vLLM, and that is not a soft requirement. The dependency list in pyproject.toml is short but unforgiving: torch, vllm, trl, peft, transformers, datasets, tensorboard, and bitsandbytes, with Python 3.12 or newer required.

AMD is a separate setup, and the instruction is to not use the default `uv.lock` because it is CUDA-only. Instinct MI-series cards need their own environment built against `configs/default_amd.yaml`, which adds two overrides: `attention_backend: TRITON_ATTN` and `optim: adamw_torch`. Anyone porting this to ROCm is doing real configuration work, not flipping a flag.

Two command names point at the same entry point, `antidoom.cli:main`: `antidoom` and `doomloop`. That is a small convenience worth knowing if you are scripting around it.

## One run directory holds the completions, the pairs, the adapter, and the merge

The full flow is a single command:

```bash
uv run antidoom -c configs/default.yaml -r runs/antidoom1 \
  --temp 0.01 \
  --model-name LiquidAI/LFM2.5-1.2B-Base
```

Set `model_name` in the config to the checkpoint you want to treat, which is how the same pipeline gets pointed at a different base model. The defaults read prompts from a prompt-only ShareGPT mixture on Hugging Face, generate the FTPO pairs, train the adapter, and merge it.

The output layout is the part that makes iteration practical. Everything for one experiment lands under the run directory: generated completions, the FTPO pair file, the trained adapter, and the merged model. Train on the pair file rather than regenerating means you can change `max_train_examples` or the learning rate and re-run training without paying for generation again.

The config keys worth knowing are grouped by stage. `generation.hf_dataset` or `generation.input_jsonl` picks the prompt source, `generation.prompt_field` picks the column, `generation.target_pairs` sets how many pairs to collect, and `generation.temperature` or `generation.temperatures` sets sampling. On the training side, `train.max_train_examples`, `train.output_dir`, and `train.merged_output_dir` control what gets used and where it lands.

## prompt_field is authoritative, and the parser will not guess your format

Generation reads whichever column you name in `prompt_field` and then decides what to do with the value by its shape. A string is a plain user prompt. An OpenAI-style list of `{"role": ..., "content": ...}` objects is used as chat messages. A ShareGPT-style list of `{"from": ..., "value": ...}` objects is converted to chat messages.

What it will not do is infer the format from the column name. That is a deliberate choice and it is the one that bites people. A dataset whose prompts live in a column called `messages` but hold ShareGPT rows will be parsed as something it is not, and the failure will show up as poor generations rather than as an error, because the shapes are close enough to pass through.

The default dataset is built for this pipeline and named for it. It is prompt-only by design, and the omissions are stated rather than accidental: no gold answers, no rationales, no hidden tests, no verifier targets, no answer labels. That is consistent with the method, since the training signal comes from the model's own sampled alternatives rather than from a correct answer, but it also means you cannot use this pipeline as a general supervised fine-tuning path. The build scripts and licence notes live in `datasets/`.

## Cap max_train_examples at 70% of what you generated, or regularisation has no room

The sizing rule is stated as a hint and the reason is mechanical. Start with at least 15k prompts and aim to produce roughly 15k to 20k preference rows. Then set `max_train_examples` below that, for example 12k rows drawn from a 15k to 20k generated set. The stated ceiling is 70% of the actual number of preference rows.

The leftover rows are not waste. They are what `rejected_regularisation_strength` works on, and that setting flattens the distribution of rejected-token frequency by culling samples. Regularisation is what stops the adapter from simply learning to suppress one word. Without it, the generated set can be badly unbalanced, which the project says can create poor training outcomes.

This also means your sample count is not fully under your control. How many preference rows you get depends on how many prompts the generation step has, how many temperature passes you run, and how doom-loopy the checkpoint actually is. A model that barely loops produces fewer rows, so a 15k target can quietly become a much smaller set. Check the generated pair count before you decide what fraction to train on.

The remaining `generation.*` defaults are described as a reasonable starting point, which is where the README stops giving guidance on them.

## chosen_win is the dial, and 0.5 is where the trainer starts hurting

Early stopping keys off `chosen_win`, the share of samples where the chosen token is beating the rejected token. The documented band is `0.15` to `0.3`, where a strong reduction in doom looping should usually show up. Past `0.5` you may be overtraining, and the project is candid that more ablations are needed on that boundary. The usage hint is to try 0.4 and adjust from there.

The failure mode of overtraining here is specific and worse than useless. The README states that an overtrained model can degrade and produce more doom loops. So the metric you are optimising is not monotonically good, and a run left to converge on its own can end up worse than the checkpoint you started from.

The learning rate needs the same caution. The stated starting range is `0.00001` to `0.00002` for training on about 12k samples, and the trainer can undertrain or overtrain, so the right value takes some trial and error. Early stopping is the recommended guard against overtraining.

Taken together, these three settings describe a training run that is meant to be stopped while it still looks like it is working. If you are used to pipelines where more steps help, that assumption is the wrong one to carry in here.

## Conclusion

Antidoom fits a team that has a specific checkpoint falling into repetition loops on reasoning prompts and can afford a GPU run to generate 15k or more prompts worth of preference rows. It does not fit anyone hoping for a general-purpose RLHF pipeline, because the dataset deliberately carries no gold answers, no rationales, and no verifier targets. Before you start, set `early_stopping_chosen_win` to 0.4 and leave it there for the first run, because the trainer will overtrain past `chosen_win` of 0.5 and an overtrained adapter produces more doom loops than the base model did.

## FAQ

### What is an example of a doom loop in a language model?

Antidoom describes three forces that line up: overtrained reasoning tokens such as `Wait` or `Alternatively` becoming unusually attractive, self-reinforcing context that raises the probability of each repeated token, and low-temperature sampling that leaves the locally reinforced loop no escape route.

### What does antidoom do to a model?

It samples completions, scans for inner repetition, marks the first readable token of the repeated segment as rejected, and trains a LoRA adapter with Final Token Preference Optimization on plausible alternatives at that position.

### How do I install and run antidoom?

Clone the repository and run `uv sync`, then execute `uv run antidoom -c configs/default.yaml -r runs/antidoom1 --temp 0.01 --model-name LiquidAI/LFM2.5-1.2B-Base`. You need a GPU runtime for PyTorch and vLLM, and Python 3.12 or newer.

### Can antidoom run on AMD GPUs?

Yes, on Instinct MI-series cards, but not with the default `uv.lock`, which is CUDA-only. You need a separate ROCm environment and `configs/default_amd.yaml`, which adds the `attention_backend: TRITON_ATTN` and `optim: adamw_torch` overrides.

### What is chosen_win in antidoom and when should I stop training?

chosen_win is the share of samples where the chosen token beats the rejected token. A value around 0.15 to 0.3 usually shows a strong reduction in doom looping, past 0.5 may be overtraining, and the suggested starting point is 0.4.

### Does the antidoom dataset contain correct answers?

No. The default prompt-only ShareGPT mixture deliberately excludes gold answers, rationales, hidden tests, verifier targets, and answer labels, because the training signal comes from sampled alternatives rather than from a correct answer.

## Sources

- [Issues](https://github.com/Liquid4All/antidoom/issues)
- [License: Apache-2.0](https://github.com/Liquid4All/antidoom/blob/main/LICENSE)
- [Liquid4All/antidoom on GitHub](https://github.com/Liquid4All/antidoom)
- [README](https://github.com/Liquid4All/antidoom/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/liquid4all-antidoom
