OpenMythos: A Theoretical Recurrent-Depth Transformer Implementation
A theoretical reconstruction of the Claude Mythos architecture, built from first principles using the available research literature.
At a glance
- What is it?
- OpenMythos is an independent, community-built PyTorch implementation of what the README calls a Recurrent-Depth Transformer (RDT): a three-stage architecture where a subset of transformer layers runs in a loop for a configurable number of iterations before producing output. It is explicitly not affiliated with or endorsed by Anthropic.
- Who is it for?
- OpenMythos is a research tool for ML engineers and researchers who want to experiment with Recurrent-Depth Transformer architectures in PyTorch. The pre-configured scale variants from 1B to 1T parameters and the training script for FineWeb-Edu provide a starting point for pre-training experiments.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 130 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What OpenMythos Is and What It Claims
OpenMythos is an open-source PyTorch implementation of a Recurrent-Depth Transformer architecture. The README includes a prominent disclaimer: OpenMythos is an independent, community-driven theoretical reconstruction based solely on publicly available research and speculation. It is not affiliated with, endorsed by, or connected to Anthropic or any of their proprietary systems.
The central hypothesis stated in the README is that a model called Claude Mythos uses a looped transformer design, also called a Recurrent-Depth Transformer or RDT. The README explains the concept: rather than stacking many unique layers, a subset of layers is reused and run through multiple times per forward pass. The README states that this reasoning happens silently inside a single forward pass, in continuous latent space, and is distinct from chain-of-thought, which produces intermediate token outputs.
OpenMythos is a theoretical reconstruction, not a reverse-engineered model. It provides a trainable architecture that implements the RDT design hypothesis, along with a training script and pre-configured scale variants. Whether the hypothesis is correct is outside the scope of the code.
The Three-Stage Architecture: Prelude, Recurrent Block, and Coda
The architecture divides layers into three functional blocks. The README describes the structure as:
Input
↓
[Prelude P] — standard transformer layers, run once
↓
[Recurrent Block R] — looped T times
↑_______↓ (hidden state h updated each loop with input injection e)
↓
[Coda C] — standard transformer layers, run once
↓
OutputThe Prelude runs once, encoding the input. The Recurrent Block runs up to `max_loop_iters` times, updating a hidden state `h` at each step. The Coda runs once on the final hidden state to produce the output. The recurrent block update rule shown in the README is `h_{t+1} = A·h_t + B·e + Transformer(h_t, e)`, where `e` is the encoded input from the Prelude injected at every loop step. The parameters `A` and `B` are learned. The README notes that this input injection prevents the model from drifting away from the original prompt signal during recurrence.
The feed-forward component is a sparse Mixture of Experts (MoE) with both routed and shared experts. The attention layer is switchable between two implementations.
Installing OpenMythos and Running a Forward Pass
OpenMythos is distributed as the `open-mythos` package on PyPI:
pip install open-mythosFlash Attention 2 support in the GQA attention variant requires CUDA and build tools:
pip install open-mythos[flash]The package requires Python 3.10 or newer. The pyproject.toml pins `torch = 2.11.0` and `transformers = ">=4.40.0"`. Running a forward pass uses the `OpenMythos` class and a `MythosConfig`:
import torch
from open_mythos.main import OpenMythos, MythosConfig
attn_type = "mla"
base = {
"vocab_size": 1000,
"dim": 256,
"n_heads": 8,
"max_seq_len": 128,
"max_loop_iters": 4,
"prelude_layers": 1,
"coda_layers": 1,
"n_experts": 8,
"n_shared_experts": 1,
"n_experts_per_tok": 2,
"expert_dim": 64,
"lora_rank": 8,
"attn_type": attn_type,
}The full implementation is in `open_mythos/main.py`. The `docs/open_mythos.md` file in the repository provides the API reference.
Pre-Configured Scale Variants from 1B to 1T Parameters
The README provides a table of seven pre-configured model scales. The `mythos_1b` variant uses a dimension of 2048, 64 experts, and 16 loop iterations with a 4k context window. At the other end, `mythos_1t` uses a dimension of 16384, 512 experts, and 64 loop iterations with a 1M context window.
The scales are imported from the package:
from open_mythos import (
mythos_1b,
mythos_3b,
mythos_10b,
mythos_50b,
mythos_100b,
mythos_500b,
mythos_1t,
OpenMythos,
)Each function returns a `MythosConfig` object that can be passed to `OpenMythos(cfg)` to instantiate the model. The parameter counts for these variants are not stated in the README for all sizes, but the table maps each variant to its dimension, expert count, loop iteration count, and context length. The `mythos_100b` and larger variants use a 1M context window.
The Two Attention Implementations: GQA and MLA
The attention layer is configurable via the `attn_type` parameter in `MythosConfig`. The README documents two options.
The `gqa` option uses Grouped Query Attention. The README cites Ainslie et al. 2023 and describes the mechanism: fewer KV heads than Q heads reduces KV-cache memory by a factor of `n_heads / n_kv_heads`. When the `flash-attn` package version 2.8.3 or newer is installed, GQA uses Flash Attention 2 for I/O-efficient computation. When Flash Attention is absent, it falls back to manual scaled dot-product attention.
The `mla` option uses Multi-Latent Attention, citing DeepSeek-V2. This variant caches a compressed KV latent rather than full K/V tensors. The README states the compression uses a `kv_lora_rank` parameter and splits RoPE and non-RoPE head dimensions for positional encoding. This approach reduces KV-cache size during inference at the cost of slightly more complex configuration.
The choice between the two affects memory usage, inference speed with Flash Attention available, and the number of configuration parameters required in `MythosConfig`.
Training on FineWeb-Edu and the Multi-GPU Setup
A training script for the 3B scale model on the FineWeb-Edu dataset is included at `training/3b_fine_web_edu.py`. The README provides single-GPU and multi-GPU launch commands.
Single GPU:
python training/3b_fine_web_edu.pyMulti-GPU with automatic GPU count detection:
torchrun --nproc_per_node=$(python -c "import torch; print(torch.cuda.device_count())") training/3b_fine_web_edu.pyThe training configuration documented in the README uses AdamW as the optimizer, the `HuggingFaceFW/fineweb-edu` dataset (the `sample-10BT` split by default), the `openai/gpt-oss-20b` tokenizer through a `MythosTokenizer` wrapper, PyTorch DDP via torchrun for parallelism, bfloat16 precision on H100 and A100 GPUs, and a linear warmup of 2000 steps followed by cosine decay. The training target is 30 billion tokens.
Limitations, Project Status, and License
OpenMythos is a theoretical architecture implementation without released model weights. The README makes no claim that any trained checkpoint is publicly available. Training the 3B model to 30 billion tokens requires significant GPU resources that the project does not provide.
The project has no GitHub releases as of the last push on 2026-05-23. The version in pyproject.toml is 0.5.0, and the development status classifier is `3 - Alpha`. The `torch = 2.11.0` pin in pyproject.toml is exact, which means the package may conflict with environments that have a different PyTorch version installed. Users should create a dedicated virtual environment.
As a theoretical reconstruction, OpenMythos cannot be validated against the design it hypothesizes. The README's disclaimer is clear on this point. The project is best understood as an ML research sandbox for experimenting with the looped transformer concept rather than as a production-grade model framework.
The license is MIT. Use, modification, and redistribution are unrestricted, including for commercial purposes. The pyproject.toml lists the homepage as `github.com/The-Swarm-Corporation/OpenMythos`.
Editorial conclusion
OpenMythos is a research tool for ML engineers and researchers who want to experiment with Recurrent-Depth Transformer architectures in PyTorch. The pre-configured scale variants from 1B to 1T parameters and the training script for FineWeb-Edu provide a starting point for pre-training experiments. It is not a working chatbot, not a deployed model, and not affiliated with any commercial AI lab. The last push was on 2026-05-23 and there are no GitHub releases, which means the project is active but has not reached a versioned release milestone. Before adopting it, verify that your PyTorch version matches the `torch = 2.11.0` pin in pyproject.toml and that you have the GPU resources for any meaningful training run. The license is MIT.
Frequently asked questions
What is the code behind Claude Mythos?
OpenMythos is a community-built theoretical reconstruction, not official code from Anthropic. The README includes a disclaimer stating it is not affiliated with, endorsed by, or connected to Anthropic or any of their proprietary systems. It implements a Recurrent-Depth Transformer hypothesis based on publicly available research.
How do I use OpenMythos?
Install with pip install open-mythos. Import OpenMythos and MythosConfig from open_mythos.main, configure the model dimensions, attention type, loop iterations, and MoE parameters in MythosConfig, then pass it to OpenMythos(cfg). A forward pass returns logits. The docs/open_mythos.md file in the repository provides the full API reference.
How do I install OpenMythos?
Run pip install open-mythos. Python 3.10 or newer is required, and torch 2.11.0 must be installed. For Flash Attention 2 support in GQA, run pip install open-mythos[flash], which additionally requires CUDA and build tools.
What is OpenMythos?
OpenMythos is an independent PyTorch implementation of a Recurrent-Depth Transformer architecture, described in the README as a theoretical reconstruction of a looped transformer design. It provides seven pre-configured model scales from 1B to 1T parameters, two attention implementations (GQA and MLA), a sparse MoE feed-forward layer, and a training script for the FineWeb-Edu dataset.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/kyegomez-openmythos)