Sentence-VAE reimplements Bowman 2015 without LSTM, and its interpolation list repeats two lines
PyTorch Re-Implementation of "Generating Sentences from a Continuous Space" by Bowman et al 2015 https://arxiv.org/abs/1511.06349
At a glance
- What is it?
- timbmg/Sentence-VAE is a PyTorch re-implementation of Generating Sentences from a Continuous Space, trained on Penn Treebank with an RNN or a GRU. Its own results page and sample list show where the write-up stops short.
- Who is it for?
- Read it as a starting point for your own experiment, not as a reproduction. The cell type you would try first is not implemented, the results table never names the cell it used, and no checkpoint exists to sample from.
- Can I use it commercially?
- Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
- Is it still maintained?
- Yes. The repository last received commits 122 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 10, 2026, and from our analysis. They are not legal advice.
Editorial analysis
A note under the title rules out LSTM before anything else is explained
One sentence into the opening block of a repository built to re-implement Bowman et al. 2015, a note settles the compatibility question before any other detail: the implementation does not support LSTMs at the moment, but RNNs and GRUs. That sentence is the whole of it. The command line reference gives `-rnn, --rnn_type` exactly two accepted values, 'rnn' or 'gru', with no third choice and no switch that turns one on. A reader arriving with the paper in hand therefore has one decision to make before anything runs, and it is a choice between two of the three recurrent cells such a model is normally assembled from. The note also carries a time phrase, at the moment, and nothing in the repository dates it or says whether it still holds. The top level holds nine entries, README.md, dowloaddata.sh, figs/, inference.py, model.py, ptb.py, requirements.txt, train.py and utils.py, and none of them is a run configuration, an experiment record or a saved model.
Four epochs of training, about one epoch of ELBO, and two columns that move apart
The performance paragraph gives the run behind the table in a single breath: training was stopped after 4 epochs, and the true ELBO was optimized for approximately 1 epoch. Elapsed epochs and useful epochs are different measures, and the paragraph points at a graph to make the second one visible. Results are averaged over entire split, a statement that applies to the whole table. Train shows 99.821 NLL with 7.944 KL, Validation 103.220 with 7.346, Test 103.967 with 7.269. Read down the columns and they drift at different rates. NLL climbs 3.399 from train to validation and another 0.747 on to test, 4.146 in total, while KL drops 0.675 across the same two steps and 0.598 of that happens before validation. A Test row exists at all, which means the run passed `--test`, and that flag measures performance on the test set, so a test split has to be present in the data directory. The data paragraph only requires `ptb.train.txt` and `ptb.valid.txt` to be there.
Eight distinct sentences are printed across ten interpolation lines
Ten lines sit under the interpolation heading and eight of them differ. The fourth and the fifth are identical strings, and the seventh and the eighth are identical to each other, which leaves two places where neighbouring points in the traversal decode to exactly the same output. The method gets one sentence: samples come from drawing twice from z ~ N(0, I) and interpolating the two samples. What that leaves out is the parameter a reader needs. There is no step count, no range for the interpolation coefficient, no indication of the order in which the ten lines should be read, and no statement that any particular line is an endpoint, even though the source sets the first and the last apart from the rest by typographic weight alone. Interpolating a continuous latent is the reason the 2015 paper exists, and this list is the only place in the repository where that idea gets exercised.
The printed samples spell an unknown word as a bare letter
Five sample sentences are drawn from z ~ N(0, I) and printed with tokenization left in. `u . s .` arrives as four spaced tokens and `company ' s` as three, with uniform spacing around every period, so what a reader sees is output that has not been detokenized. Unknown words surface as a bare letter rather than a word. The first sample reads `mr . n who was n't n with his own staff and the n n n n n`, which is seven bare n tokens in eighteen, and it opens on the two-token fragment `mr . n` that shares nothing with the sentence behind it. The argument reference spells the replacement token as `<unk>`, and that form appears in none of the five samples or the ten interpolation lines, so the printed output and the option documentation name unknown words two different ways. Nothing in the repository says whether the bare letter is the token itself, a fragment of one, or an artifact of the printing step.
Twenty-three flags, every one explained, not one default stated
The argument reference is careful about what each option does and silent about what each one starts at. It walks through twenty-three entries, from `--data_dir` and `--create_data` at the top down to `-v, --print_every` and `-bin, --save_model_path` at the bottom, and not a single line carries a default value. Two entries enumerate what they accept, `--rnn_type` with 'rnn' or 'gru' and `-af, --anneal_function` with 'logistic' or 'linear', which hints at the shape of the rest without giving any starting points. The annealing pair is where that silence costs the most. `-x0, --x0` is the mid-point where the weight is 0.5 under the logistic function and the denominator under the linear one, so one flag carries a different quantity depending on the choice made one entry earlier. `-k, --k` is the steepness of the logistic curve, which leaves it with no work to do once 'linear' is selected. The KL column is the quantity that schedule exists to shape, and no line here says which schedule produced 7.944.
Four knobs reshape the decoder input, and the auxiliary files decide whether they count
Four options change what reaches the decoder, and they act at different stages. `--min_occ` works on the corpus, replacing a word that occurs fewer than min_occ times with the `<unk>` token before training starts. `-wd, --word_dropout` works per batch, replacing decoder input words with `<unk>` at a probability, and `-ed, --embedding_dropout` drops embeddings at that same decoder input. `--max_sequence_length` cuts off long sentences. Unknown tokens therefore arrive from a static frequency count and from a random draw, on top of whatever the cut-off removes. The derived files holding those counts live in `--data_dir`, and `--create_data` is the documented way to build new ones from the source data, so the vocabulary options are only as fresh as the auxiliary files sitting in that directory. The corpus itself is the Penn Treebank archive linked from Tomas Mikolov's page, an unversioned tarball fetched over plain http, and the script that retrieves it is spelled `dowloaddata.sh`.
The sampling command needs a checkpoint the repository has never published
Sampling runs through one line:
python3 inference.py -c $CHECKPOINT -n $NUM_SAMPLES`-c` has to name a checkpoint, and checkpoints are written wherever `-bin, --save_model_path` points while training runs. There are no GitHub releases, and the nine top level entries contain no weights file, so there is nothing to download and nothing to hand to `-c`. Training is one command:
python3 train.pySo the first user has to train before sampling, which means running the ELBO and KL schedule from the previous section blind. The two shell placeholders above are the only guidance on those variables, and the README does not repeat the argument reference for inference.py, so the options that script accepts beyond `-c` and `-n` are written down nowhere in the repository.
Two sections named Training, a KL heading one level too high, and a script that misspells download
Two defects sit in how the file is arranged rather than in what it claims. The first is heading structure. Results opens with a Training subsection holding ELBO and Negative Log Likelihood as its own children, KL Divergence then appears one level higher than those two graphs, and a separate Training section much further down covers the command line. Two different sections carry the same name at two different depths, and the KL heading sits apart from the pair it belongs with. The second is spelling, carried into prose and into the filesystem: Sentenes appears twice, the download sentence reads donwloaded, the remark about the graph reads as can bee see, the inference section opens on senteces, one line reads and the interpolating the two samples, and the script meant to fetch the corpus is `dowloaddata.sh`. The dependency file mixes styles in four lines, leaving numpy>=1.22 open while nltk==3.9, torch==2.2.0 and tensorboardX==2.0 are pinned exactly, and no Python version appears anywhere even though every command says python3.
Editorial conclusion
Read it as a starting point for your own experiment, not as a reproduction. The cell type you would try first is not implemented, the results table never names the cell it used, and no checkpoint exists to sample from. Before trusting a run, confirm which value went to `--rnn_type`, keep `-k` and `-x0` tied to the schedule you chose with `-af`, and check whether the auxiliary files in `--data_dir` were rebuilt after you touched `--min_occ` or `--max_sequence_length`. Anyone who needs published figures or a pretrained model will find neither in this repository.
Frequently asked questions
Does timbmg/Sentence-VAE support LSTM encoders and decoders?
No. A note in the opening block says the implementation does not support LSTMs at the moment, and the cell flag accepts only 'rnn' or 'gru' as its value for `--rnn_type`.
Which recurrent cell produced the NLL and KL numbers in timbmg/Sentence-VAE?
The repository does not say. It prints 99.821, 103.220 and 103.967 NLL for train, validation and test after four epochs, but never names the value that was passed to `--rnn_type`.
Where does timbmg/Sentence-VAE get the Penn Treebank data it trains on?
From the archive linked on Tomas Mikolov's page. The code expects at least `ptb.train.txt` and `ptb.valid.txt` in the directory passed as `--data_dir`, and `dowloaddata.sh` can fetch the data.
Can timbmg/Sentence-VAE generate samples without training a model first?
No. No checkpoint is published and the repository has no GitHub releases, so the line `python3 inference.py -c $CHECKPOINT -n $NUM_SAMPLES` needs a checkpoint written by your own `python3 train.py` run.
What are the word_dropout and min_occ options in timbmg/Sentence-VAE for?
`--min_occ` replaces every corpus word occurring fewer than min_occ times with the `<unk>` token, while `-wd, --word_dropout` replaces decoder input words with that token at a probability, and `-ed, --embedding_dropout` drops embeddings instead.
What license does timbmg/Sentence-VAE carry?
No license value appears in the repository metadata, and the nine top level entries include no LICENSE file, so the terms for reuse are not stated anywhere in the project.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/timbmg-sentence-vae)