Model or dataset
Liquid4All/antidoom avatar
Liquid4All/antidoom

Antidoom: Generating FTPO Training Data to Fix LLM Repetition Doom Loops

Antidoom generates and trains targeted preference data for reducing model repetition loops (doom loops).

389 stars38 forksPythonApache-2.0

At a glance

What is it?
Antidoom is a Python tool from LiquidAI that targets a specific LLM failure mode: runaway repetition during reasoning, also called a doom loop. Rather than training on full gold responses, it detects where a repetition begins, marks the first loop-starting token as rejected, samples plausible alternatives, and trains a LoRA adapter using Final Token Preference Optimization. It is aimed at LLM engineers working on reasoning models who see doom-loop behavior under low-temperature sampling.
Who is it for?
Antidoom is the right tool for engineers who have a reasoning model that falls into repetition loops at low temperature and want to fix the behavior with targeted LoRA training rather than changing the sampling strategy. It is a poor fit for engineers whose models do not exhibit doom-loop behavior, or who want to improve output quality more broadly: the README is explicit that this is a narrow tool for a narrow failure mode.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 86 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The Problem Antidoom Addresses

A doom loop in an LLM is a runaway repetition during generation: the model keeps producing the same token sequence in a cycle, or oscillates between a small set of tokens without advancing the reasoning. This is distinct from occasional repetition. A doom loop effectively freezes the output until it hits a token or length limit.

The README describes three conditions that tend to cause them. First, overtrained tokens: common reasoning tokens like Wait, So, But, and Alternatively can become overrepresented after heavy synthetic reasoning training. Second, self-reinforcing context: once a short sequence appears, the prior context makes it more likely to appear again, and across repeated turns the probability of each token can approach certainty. Third, low-temperature sampling: at temperature 0 or near 0, the model keeps selecting the highest-probability token and cannot escape a locally reinforced loop.

Antidoom attacks the problem at its origin: the specific token position where the loop begins. Its approach is to generate model completions, detect where a repeated span starts, and train a local preference at that exact token position.

How FTPO Works: Targeting the Loop's First Token

Final Token Preference Optimization adapts the single-token preference training idea from the Antislop paper (arxiv 2510.15061) to repetition loops. Antislop applied this technique to overused phrases; Antidoom applies it to the specific token that starts a runaway repetition.

For each prompt, Antidoom generates a completion and scans it for inner repetition. When it finds a loop, it refines the boundary in token space so the rejected token is the first readable token of the repeated segment. Each FTPO training row contains a prompt and context prefix ending immediately before the rejected token, one rejected token (the token that begins the loop), one or more chosen tokens sampled from filtered alternatives at that same position, and metadata about the source prompt and generated completion.

The training step includes regularization that flattens the distribution of rejected-token frequency. This prevents the adapter from learning to suppress one specific common token while ignoring others, ensuring the trained preference generalizes across the space of loop-starting tokens rather than patching one word.

Installing and Running the Pipeline

Antidoom uses uv for dependency management. Clone the repository and sync the environment:

bash
git clone https://github.com/Liquid4All/antidoom
cd antidoom
uv sync

Edit configs/default.yaml and set model_name to the checkpoint you want to apply the anti-doom training to. The default config reads prompts from LiquidAI/antidoom-mix-v1.0 on HuggingFace.

Run the full generate-and-train flow:

bash
uv run antidoom -c configs/default.yaml -r runs/antidoom1 \
  --temp 0.01 \
  --model-name LiquidAI/LFM2.5-1.2B-Base

This writes generated completions and FTPO pairs to runs/antidoom1/, trains on the pair file, and writes the LoRA adapter and the merged model under the same run directory.

For AMD/ROCm hardware, do not use the default uv.lock file, which is CUDA-only. The README provides a separate setup path for ROCm using configs/default_amd.yaml with overrides for attention_backend: TRITON_ATTN and optim: adamw_torch.

Key Configuration Parameters and Their Practical Meaning

Three parameters in configs/default.yaml have the largest effect on training quality.

max_train_examples controls how many FTPO rows are actually used for training. The README advises setting it to at most 70% of the total generated preference rows. Starting with at least 15,000 prompts and aiming for 15,000 to 20,000 generated preference rows before training gives the regularization step room to balance the rejected-token distribution. A 12,000-row training set from a 15,000-to-20,000-row generated set is cited as a reasonable starting configuration.

learning_rate has a narrow useful range. The README suggests starting between 0.00001 and 0.00002 when training on about 12,000 samples. Overtraining is a real risk: an overfit adapter can produce more doom loops than the original model.

early_stopping_chosen_win stops training when the share of samples where the chosen token beats the rejected token exceeds a threshold. The README states that around 0.15 to 0.3 chosen_win is typically enough for a strong reduction in doom looping, and suggests trying 0.4 as a starting value while noting that going past 0.5 may be overtraining.

Limitations and What Antidoom Cannot Fix

Antidoom is a narrow tool for a narrow failure mode. The README states this directly: if a model does not exhibit doom loops, or if its main quality problems are factual accuracy, instruction following, or output diversity, Antidoom is not the right tool.

The method depends on being able to elicit doom loops from the model during the generation phase. A model that only loops on rare or unusual inputs may require very large prompt sets to generate enough preference pairs. The README's suggestion of at least 15,000 prompts is a heuristic starting point; highly stable models may need more.

The current supported GPU paths are NVIDIA/CUDA (the default) and AMD/ROCm. The README states that CUDA performance is the focus. CPU-only inference is possible with PyTorch but is not designed as a practical training path given the scale of generation needed.

The learning rate and early stopping thresholds require empirical tuning. The README provides starting points but notes that more ablations are needed before confident recommendations can be made. A poorly tuned training run can leave the model worse than before.

The project has no GitHub releases as of the last push on 2026-07-07.

Compared with Full-Response Preference Training

Standard RLHF and DPO training compare full chosen and rejected response pairs. A chosen response is the better complete answer; a rejected response is the worse one. This approach improves overall response quality but requires human-rated or model-rated full responses.

Antidoom takes the opposite approach: it trains on a single token position, not a full response pair. The rejected token is the one that starts the loop. The chosen tokens are plausible alternatives that do not continue the loop. No human rating is needed; the preference labels come from the generated completion's own repetition structure.

This makes Antidoom faster to apply and more targeted than full-response DPO, but also more limited. It cannot improve response quality, factual accuracy, or style. It can only reduce the specific failure mode of doom looping.

Dataset, Maintenance, and Licence

The default prompt dataset is LiquidAI/antidoom-mix-v1.0, available on HuggingFace. The README describes it as a prompt-only ShareGPT mixture built for this pipeline. It intentionally excludes gold answers, rationales, hidden tests, verifier targets, and answer labels, so the only signal the pipeline uses is the repetition detected in the model's own completions on those prompts. Dataset build scripts and licence notes are in the datasets/ directory.

The last push to the repository was on 2026-07-07. The project is not archived. There are no GitHub releases.

The project is licensed under Apache-2.0. Apache-2.0 permits use in commercial products, modification, and redistribution, provided attribution and the licence notice are preserved. It includes an explicit patent grant, which is absent from MIT and BSD licences.

Editorial conclusion

Antidoom is the right tool for engineers who have a reasoning model that falls into repetition loops at low temperature and want to fix the behavior with targeted LoRA training rather than changing the sampling strategy. It is a poor fit for engineers whose models do not exhibit doom-loop behavior, or who want to improve output quality more broadly: the README is explicit that this is a narrow tool for a narrow failure mode. Before running, generate at least 15,000 prompts and verify that the number of preference rows produced reaches the training target set in max_train_examples.

Frequently asked questions

What is FTPO (Final Token Preference Optimization)?

According to the README, FTPO is a training method that adapts the single-token preference training idea from the Antislop paper to the problem of doom loops. Instead of training on full chosen and rejected responses, it trains a preference at a single token position: the first token of a detected repetition is the rejected token, and plausible alternatives at that same position are the chosen tokens.

How much training data does Antidoom need?

The README recommends starting with at least 15,000 prompts to generate 15,000 to 20,000 preference rows. The max_train_examples setting should be at most 70% of the generated row count, so a 12,000-example training run from a 15,000 to 20,000 row generated set is a reasonable starting configuration. Models with fewer doom loops may require more prompts.

What hardware does Antidoom require?

Antidoom needs a supported GPU runtime for PyTorch and vLLM. NVIDIA/CUDA is the default path and uses the pinned uv.lock file. AMD/ROCm on Instinct MI-series hardware is supported but requires a separate environment and the configs/default_amd.yaml configuration file, which adds required attention_backend and optimizer overrides. CPU-only is not a practical training path.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. Liquid4All/antidoom on GitHub
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/liquid4all-antidoom.svg)](https://hysenlabs.com/projects/liquid4all-antidoom)