BDH (Dragon Hatchling): A Biologically Inspired Alternative to the Transformer
BDH (Dragon Hatchling) – Architecture and Code
At a glance
- What is it?
- BDH, or Dragon Hatchling, is a Python research implementation of a biologically inspired neural architecture developed at Pathway. It replaces the Transformer's matrix-based attention with a scale-free network of locally interacting neurons and matches GPT-2-scale Transformers on language benchmarks at equivalent parameter counts.
- Who is it for?
- Researchers exploring interpretable neural architectures or biologically grounded alternatives to the Transformer will find BDH's codebase a concrete starting point. The baseline variant in this repository does not reproduce the 97.4% Sudoku benchmark result reported in Pathway's internal implementation.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 137 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What BDH Is and the Problem It Addresses
Large language models based on the Transformer architecture process tokens through matrix multiplication and dot-product attention. This design is efficient on GPUs but operates in a latent space that is difficult to interpret, and it processes information token-by-token in a way that the README describes as limiting for search-heavy, non-linguistic reasoning tasks.
BDH, short for Dragon Hatchling, is the reference implementation of an architecture described in the paper 'The Dragon Hatchling: The Missing Link between the Transformer and Models of the Brain' by researchers at Pathway. The core claim is that attention in neural networks can emerge naturally from a graph-based network of locally interacting neurons, rather than being imposed through explicit matrix operations. The README frames this as a bridge between machine-based language understanding and the principles of biological neural computation.
The practical target audience for this repository is researchers studying interpretable architectures, neuroscience-inspired machine learning, and alternatives to the Transformer scaling paradigm. It is not a production-ready language model toolkit; it is a research baseline with a single training script and a minimal dependency list.
Architecture: Scale-Free Networks and Hebbian Memory
BDH represents a network topology described as scale-free, meaning the connectivity between neurons follows a power-law distribution similar to what is observed in biological neural systems. Within this topology, neurons interact locally rather than through global attention over all positions simultaneously.
The key architectural properties listed in the README are: excitatory and inhibitory neuron dynamics, Hebbian working memory based on synaptic plasticity (described as showing monosemanticity, meaning individual neurons or small groups encode single concepts), a GPU-friendly state-space formulation, and sparse, positive activations that the README calls interpretable.
BDH and the Transformer share attention-inspired computation at a conceptual level. The difference the README emphasises is that BDH's attention emerges from neuron-level graph interactions, whereas the Transformer's attention is a direct matrix operation over a sequence. The README also notes that BDH follows Transformer-like scaling laws, maintaining parameter efficiency while adding interpretability at any scale. Empirically, the README states that BDH matches GPT-2-scale Transformers on language and translation tasks at parameter scales from 10M to 1B.
Installing BDH and Running the Training Script
The repository structure is minimal: `bdh.py` contains the architecture implementation, `train.py` is the training entry point, `requirements.txt` lists the dependencies, and `figs/` holds figures from the paper. Installation requires Python and pip:
pip install -r requirements.txtThe requirements file specifies only three packages: `torch`, `numpy`, and `requests`. After installation, the toy training run uses:
python train.pyThe README credits Andrej Karpathy's nanoGPT for the training code and the tiny Shakespeare dataset used in this demonstration. That means the toy training setup is a modified version of nanoGPT's character-level language model training loop, adapted to use the BDH architecture instead of a standard Transformer. The repository does not include documentation on training configurations, hyperparameter choices, or how to swap datasets beyond the toy example.
The Sudoku Benchmark and What This Repository Does Not Include
The README describes a benchmark on Extreme Sudoku puzzles where a Pathway BDH implementation reached 97.4% accuracy across roughly 250,000 difficult puzzles, without chain-of-thought reasoning, solution backtracking, or external tool use. The README includes a comparison table showing leading LLMs achieving approximately 0% accuracy on the same benchmark.
This is an important limitation to understand clearly: the README explicitly states that this result refers to Pathway's internal BDH implementation, not the current open-source repository. The public repository contains the baseline variant as described in the paper, and the README says it does not reproduce the 97.4% benchmark result out of the box. The Sudoku Extreme implementation is separate and not included here.
Researchers who want to reproduce the Sudoku benchmark result will need to contact Pathway directly or wait for a separate release. The publicly available code is sufficient for studying the baseline architecture described in the arXiv paper (arXiv:2509.26507) and for running small-scale language experiments, but it is not the full production system.
Community Ports and Research Context
Since the initial release, several community members have created ports and extensions. The README lists: `adamskrodzki/bdh` adding dynamic vocabulary and stateful attention, `mosure/burn_dragon_hatchling` porting BDH to the Burn deep learning framework written in Rust, `severian42/bdh` porting to MLX (Apple's machine learning framework for Apple Silicon), and two other forks with custom modifications.
The paper has been discussed on Hugging Face Papers, Alphaxiv, and EmergentMind, and covered in media outlets including Forbes, Semafor, The Turing Post, Quantum Zeitgeist, and Golem. The research team includes Adrian Kosowski, Przemek Uznanski, Jan Chorowski, Zofia Stamirowska, and Maciej Bartoszkiewicz. A 72-minute podcast episode on SuperDataScience features Adrian Kosowski explaining the architecture in detail.
The last push to this repository was on 2026-05-16. The repository has no published GitHub releases; the arXiv paper is the primary reference for the architecture specification.
How BDH Differs from the Standard Transformer
The Transformer, as implemented in the PyTorch ecosystem and the Hugging Face Transformers library, uses multi-head self-attention through matrix multiplications that let every token attend to every other token in the sequence. This global attention is computed in parallel and scales quadratically with sequence length. Interpretability requires post-hoc attribution techniques because the weights do not correspond directly to neuron-level computations.
BDH replaces this with a graph where each neuron particle interacts only with its local neighbors, and attention arises from those local dynamics rather than from a global matrix operation. The activations are sparse and positive by design, which the README argues makes them directly interpretable without additional tools. The state-space formulation makes BDH trainable on GPUs despite its graph structure.
The practical limitation compared to a standard Transformer is that the community tooling, pre-trained checkpoints, fine-tuning utilities, and deployment infrastructure built around the Hugging Face ecosystem do not apply to BDH. Anyone adopting BDH for a project beyond research exploration is responsible for building that infrastructure from scratch, starting from the toy training script in `train.py`.
Editorial conclusion
Researchers exploring interpretable neural architectures or biologically grounded alternatives to the Transformer will find BDH's codebase a concrete starting point. The baseline variant in this repository does not reproduce the 97.4% Sudoku benchmark result reported in Pathway's internal implementation. Before using BDH for serious experiments, verify that the `train.py` entry point and the single `bdh.py` module cover the aspects of the architecture you want to study; the repository does not include training infrastructure beyond the toy Shakespeare dataset inherited from nanoGPT.
Frequently asked questions
What is BDH in AI?
BDH, or Dragon Hatchling, is a biologically inspired neural architecture developed at Pathway that replaces the Transformer's matrix attention with a scale-free graph of locally interacting neurons. The architecture matches GPT-2-scale Transformers on language benchmarks at equivalent parameter scales and produces sparse, interpretable activations.
What does Pathway Company do?
Pathway is a research and engineering company that developed BDH. The README identifies Pathway as the organisation behind this biologically inspired language model architecture research. The company publishes related research on its website at pathway.com.
Does the BDH repository include the Sudoku benchmark implementation?
No. The repository contains the baseline variant described in the public arXiv paper. The README explicitly states that the 97.4% Extreme Sudoku accuracy result refers to Pathway's internal implementation, which is not included in this open-source repository and cannot be reproduced from the code here.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/pathwaycom-bdh)