Library / SDK
keon/seq2seq avatar
keon/seq2seq

keon/seq2seq: A Minimal PyTorch Seq2Seq Implementation with Bahdanau Attention for Neural Machine Translation

Minimal Seq2Seq model with Attention for Neural Machine Translation in PyTorch

706 stars166 forksPythonMIT

At a glance

What is it?
keon/seq2seq is a concise PyTorch implementation of a sequence-to-sequence model with attention, trained on German-to-English translation using the Multi30k dataset. It is built for readability and modularity rather than production performance, making it a practical starting point for understanding attention-based NMT.
Who is it for?
keon/seq2seq is the right reference when you need a short, self-contained PyTorch seq2seq implementation to study, modify, or integrate into a research project. The code is minimal by design, which means it lacks beam search decoding, BLEU evaluation, model checkpointing, and the architectural depth of a production NMT system.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 136 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What keon/seq2seq Is and Who It Is For

keon/seq2seq is a minimal implementation of a sequence-to-sequence model with attention for neural machine translation in PyTorch. The README states three design goals: modular structure for reuse in other projects, minimal code for readability, and full utilization of batches and GPU. These goals make it useful as a teaching reference, a starting point for experimentation, or a component to adapt into a larger system.

The target user is a machine learning engineer or researcher who wants to understand how attention-based seq2seq works at the code level, without the complexity of a production implementation. The dataset is Multi30k German-to-English, loaded via the HuggingFace datasets library. Tokenization uses spaCy with the de_core_news_sm and en_core_web_sm models.

The last push was on 2026-05-18. The repository has no GitHub releases.

Architecture: Bidirectional GRU Encoder and Attention Decoder

The encoder is a bidirectional GRU (Gated Recurrent Unit). Bidirectional processing means each encoder hidden state carries information from both the preceding and following tokens in the source sequence, giving the decoder richer context than a unidirectional encoder would provide.

The decoder is a GRU with an attention mechanism. The attention implementation follows the paper by Bahdanau, Cho, and Bengio: Neural Machine Translation by Jointly Learning to Align and Translate (arXiv:1409.0473). In this mechanism, the decoder computes an alignment score between each encoder hidden state and the current decoder state, producing a context vector as a weighted sum of encoder states. This context vector is concatenated with the decoder input at each step, allowing the decoder to selectively attend to different parts of the source sequence as it generates each output token.

The model.py file contains the complete architecture. The modular structure documented in the README means the encoder, decoder, and attention modules can be imported and reused in other projects without modification.

Installing and Running the Training Script

The requirements file lists three dependencies:

bash
pip install -r requirements.txt

The requirements.txt content:

code
torch>=2.0
datasets>=2.14
spacy>=3.7

After installing Python dependencies, download the spaCy language models:

bash
python -m spacy download de_core_news_sm
python -m spacy download en_core_web_sm

Start training with the default hyperparameters:

bash
python train.py -epochs 30 -batch_size 32 -lr 3e-4

Device selection is automatic: CUDA is used if available, then Apple MPS, then CPU. For CPU smoke runs on a machine without a GPU, the README recommends passing smaller -hidden_size and -embed_size flag values to reduce model capacity and make runs feasible on CPU. The sanity check table in the README was produced with hidden_size=128 and embed_size=64.

Training Dynamics: Loss Curve from the README

The README includes a sanity check table from a CPU run with hidden_size=128 and embed_size=64 over 500 batches:

| step | train loss | perplexity | |------|-----------|------------| | init | 9.19 | 9803 | | 50 | 6.98 | 1071 | | 100 | 5.48 | 239 | | 250 | 5.15 | 173 | | 500 | 4.84 | 127 |

The README notes that the random-initialization prior is approximately log(|V|) = 9.19, where |V| is the vocabulary size. Starting at 9.19 means the model begins at chance performance and improves from there. The final validation loss of 4.93 represents meaningful learning from the initial prior, though perplexity of 127 on a small model trained for 500 CPU batches is not competitive with a full training run on GPU.

The table gives a concrete benchmark for debugging: if a training run on similar hardware does not reproduce approximately these loss values, something is likely wrong with the data loading, tokenization, or model configuration.

Dataset Loading and the HuggingFace Dependency

The implementation uses the HuggingFace datasets library to load Multi30k, which replaced the previously common torchtext approach. Multi30k is a paired image-caption dataset with approximately 30,000 English-German sentence pairs, commonly used as a benchmark for neural machine translation research because it is small enough to train quickly on a GPU while being large enough to show meaningful learning.

The spaCy tokenizer requires downloading two language models: de_core_news_sm for German and en_core_web_sm for English. These are rule-based tokenizers that split text into tokens before the vocabulary lookup step. The choice of spaCy over a simpler split-on-whitespace tokenizer means the model handles punctuation and compound words more accurately.

The utils.py file in the repository handles vocabulary building and data preprocessing. The train.py script handles the training loop, loss computation, and device selection.

Where keon/seq2seq Ends: Missing Features and Comparison to Transformers

The README explicitly describes the implementation as minimal. Features common in production NMT systems that are not present here include beam search decoding (only greedy decoding is documented), BLEU score evaluation, model checkpointing (saving and resuming training), teacher forcing schedule annealing, and attention visualization.

The seq2seq with attention architecture predates the Transformer. The Transformer architecture (introduced by Vaswani et al. in 2017) replaced recurrent encoders and decoders with self-attention layers, allowing parallelization across the entire sequence rather than processing tokens sequentially. Modern high-quality translation models are all Transformer-based. The GRU encoder-decoder with attention in keon/seq2seq processes sequences recurrently, which is slower to train and less effective at long sequences than a Transformer.

The distinction matters for choosing how to use this repository: it is an educational reference for the attention mechanism and seq2seq architecture in their original GRU form, not a competitive baseline for translation quality. For learning how Transformers work, a separate implementation would be needed. For a competitive seq2seq baseline, a pretrained model from the HuggingFace model hub would be more appropriate.

Editorial conclusion

keon/seq2seq is the right reference when you need a short, self-contained PyTorch seq2seq implementation to study, modify, or integrate into a research project. The code is minimal by design, which means it lacks beam search decoding, BLEU evaluation, model checkpointing, and the architectural depth of a production NMT system. If you need a seq2seq baseline for a research comparison, this implementation gives you a clean starting point; if you need a competitive translation model, a Transformer-based system trained on a larger corpus is the appropriate direction. The last push was on 2026-05-18.

Frequently asked questions

What is a seq2seq model?

A seq2seq (sequence-to-sequence) model maps an input sequence to an output sequence of a different length. The encoder processes the input and produces a fixed-length context representation; the decoder generates the output sequence one token at a time, conditioned on that context. keon/seq2seq implements this with a bidirectional GRU encoder and a GRU decoder with Bahdanau attention.

Is seq2seq an encoder-decoder architecture?

Yes. A seq2seq model is a specific application of the encoder-decoder architecture where both the input and output are sequences (such as sentences in two different languages). keon/seq2seq uses a bidirectional GRU as the encoder and a GRU with attention as the decoder.

How does seq2seq differ from a Transformer?

A seq2seq model with GRU processes tokens recurrently, one at a time, which limits parallelization and performance on long sequences. A Transformer replaces recurrent layers with self-attention, allowing all positions to be processed in parallel. keon/seq2seq is a GRU-based seq2seq with Bahdanau attention and does not implement the Transformer architecture.

Is seq2seq an RNN-based model?

The seq2seq architecture was originally proposed with RNNs (recurrent neural networks). keon/seq2seq uses GRUs, which are a gated variant of RNNs that address the vanishing gradient problem. GRUs are RNNs, so yes: this implementation is RNN-based.

Official sources

  1. Issues
  2. keon/seq2seq on GitHub
  3. License: MIT
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/keon-seq2seq.svg)](https://hysenlabs.com/projects/keon-seq2seq)