Model or dataset
haykgrigo3/TimeCapsuleLLM avatar
haykgrigo3/TimeCapsuleLLM

TimeCapsuleLLM: a from-scratch language model trained only on 1800-1875 London text

A LLM trained only on data from certain time periods to reduce modern bias

1,982 stars76 forksPythonMIT

At a glance

What is it?
TimeCapsuleLLM trains GPT-style models exclusively on period documents from 1800 to 1875 London to cut modern bias out of the output. The repository is honest about the cost: small corpora, OCR noise and a high hallucination rate are documented in the README itself.
Who is it for?
Adopt TimeCapsuleLLM if you are studying historical language, want a controllable corpus, or need a base model whose vocabulary is bounded by a date range rather than by a 2024 web crawl. Do not adopt it as a general assistant or for factual question answering: the README documents the v0.5 model as still having a high factual hallucination rate, and the v2mini-eval1 release notes describe a tokenization defect that splits words apart.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 82 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem TimeCapsuleLLM was built to attack

Ask a modern chat model what London smelled like in 1834 and you get a plausible paragraph assembled from twenty-first-century summaries. The vocabulary is period-flavored, the assumptions are not. TimeCapsuleLLM starts from the opposite end: the model is trained from scratch on documents from a fixed window, so the only words it can produce are words that appeared in that window. The README frames the goal as emulating the voice, vocabulary and worldview of the era, and its one-line pitch is that the model should not pretend to be historical but actually be constrained by historical text.

The audience is narrow and that is the point. Digital humanities researchers, corpus linguists, and people building period-specific generative text want a model whose priors come from a known document set rather than from an undisclosed web scrape. If your question is what a Londoner in 1850 would have said, a model that has never read a 2020 blog post is a different instrument from one that has.

How the training pipeline is put together

The repository is a training and data-collection project, not an inference server. At the top level it holds download_texts_improved.py, london_corpus_dataset.py, internet_archive_ids.txt, and one directory per model generation: london_1800_1850_v0, london_1800_1875_v0.5, london_1800_1875_v2, and london_1800_1875_v2mini_eval1. A file named Copy of London Documents for Time Capsule LLM.txt sits alongside them.

The lineage is documented rather than hidden. v0 and v0.5 use the training scripts and architecture from nanoGPT by Andrej Karpathy. v1 is built on Microsoft's Phi 1.5. v2 is built on LlamaForCausalLM. That means the project's contribution is the corpus and the training configuration, not a novel architecture, and the README says so directly. The dataset lives on Hugging Face under the name postgrammar/london-llm-1800, titled Historic London English (1800-1875), with a BibTeX entry crediting Hayk Grigorian and Hamed Yaghoobian.

Data flow is therefore: Internet Archive identifiers feed a download script, the corpus script assembles period text, and a generation-specific directory holds the training run and its outputs. The v2mini-eval1 release is described as trained on a 15GB sample drawn from v2's 90GB dataset, which gives a rough sense of scale between the full run and the evaluation checkpoint.

Running the corpus scripts and finding the weights

The README does not publish a pip package or a single install command, so treat this as a source checkout. Clone the repository, then run the corpus scripts in the order the filenames imply: fetch document identifiers first, then build the dataset.

bash
git clone https://github.com/haykgrigo3/TimeCapsuleLLM.git
cd TimeCapsuleLLM
python download_texts_improved.py
python london_corpus_dataset.py

The download script reads internet_archive_ids.txt and pulls the source documents; the corpus script assembles them into training text. Expect OCR artifacts in the result, because the README notes that noise such as "Digitized by Google" survives into model output for v0.5. That is a property of the source scans, not a bug you can configure away.

For the trained weights, the project points to Hugging Face rather than shipping checkpoints in the repository. The README links a collection under haykgrigorian named timecapsulellm-1800-1875-london, and the citation entry gives the dataset as postgrammar/london-llm-1800.

Once you have a checkpoint, the README's own examples show what to expect. Prompting v0 with "Who art Henry?" returned "I know that man, I have did not a black, the storm." Prompting v1 with "It was the year of our Lord 1834" produced a long passage about protest and petition that names Lord Palmerston. The first is era-flavored noise; the second is the result the project considers its first real success.

Where the model breaks, according to its own README

The limitations section is unusually blunt, and it is the most useful part of the repository. v0 is described as producing mostly incoherent sentences, which the README attributes to roughly 187MB of training data. v0.5 fixes the grammar and punctuation but is still described as having a high factual hallucination rate, and it emits OCR boilerplate from the scanned sources. Fluent period prose and factual accuracy are separate properties here, and the project only claims the first.

The v2mini-eval1 release has a more specific failure. The README states there was an issue with tokenization, and the sample output shows it plainly: "W ho is Charles D ic ens ?" with spaces inserted inside words. Tokenization defects are not cosmetic, because the model is learning from a corrupted view of the text. The README also notes that this checkpoint was trained to 10K steps only, so it is an evaluation artifact rather than a finished model.

The wrong-tool case follows from that. If you need answers about history, this is the wrong instrument: a model trained on a bounded corpus with no retrieval layer will confabulate names and events with period grammar, which is harder to spot than modern-sounding nonsense. It is also the wrong choice if you need broad English coverage, since the corpus is London-specific and the releases are labeled 1800-1875.

How it differs from fine-tuning a modern base model

The obvious alternative is taking a capable open base model and fine-tuning it on the same 1800-1875 corpus. That approach is cheaper, converges faster, and inherits working tokenization and instruction following. The difference in approach is what the model knows before it sees your data. A fine-tuned modern model still carries a web-scale prior; period style becomes a layer painted over it, and the underlying model can drift back toward modern phrasing, modern named entities, and modern framing when a prompt pulls it there.

TimeCapsuleLLM removes that prior by construction. There is no modern text in the training set, so the model cannot produce a concept it never saw, which is exactly the property the README claims for v0: no mention of modern concepts. The price is everything a large pretraining run buys you. Coherence, factual grounding and robustness all have to be earned from a corpus measured in hundreds of megabytes for v0, and the README says that is why v0's sentences fall apart.

The choice is therefore about which failure you can tolerate. Fine-tuning gives you a capable model that occasionally sounds modern. Training from scratch gives you a model that cannot sound modern and frequently sounds wrong. For stylometric work the second failure is measurable and the first is fatal; for anything user-facing the reverse holds.

Maintenance, licensing and what a fork costs

The repository is not archived, and the last push was on 2026-07-11, which is recent enough that the project is still moving. The release history shows a cadence of roughly one substantial checkpoint every few months across 2025 and 2026, from v2mini-eval2 in December 2025 through v2 in January 2026 to v3mini-eval1 in July 2026. The v3mini-eval1 release is labeled English (1800-1875), a shift from the London-specific labeling of v2, so the corpus scope appears to be widening.

The licence is MIT, which is permissive and places few conditions on reuse or redistribution of the repository code. Two caveats belong next to that. First, the trained weights and the dataset are distributed through Hugging Face, not through this repository, so their terms are a separate question from the MIT licence on the code. Second, the underlying source documents come from the Internet Archive and other scans, and the project's own BibTeX entry asks that academic use cite the dataset. This is a description of what the repository states, not legal advice; check the terms on the Hugging Face dataset page before redistributing the corpus.

Upgrade cost is dominated by data, not by code. Each generation changed its base architecture, from nanoGPT to Phi 1.5 to LlamaForCausalLM, so a fork pinned to one generation cannot simply pull the next one's training scripts. Reproducing a run means re-downloading the period texts and re-running the corpus pipeline, which is the expensive part.

Editorial conclusion

Adopt TimeCapsuleLLM if you are studying historical language, want a controllable corpus, or need a base model whose vocabulary is bounded by a date range rather than by a 2024 web crawl. Do not adopt it as a general assistant or for factual question answering: the README documents the v0.5 model as still having a high factual hallucination rate, and the v2mini-eval1 release notes describe a tokenization defect that splits words apart. Before building on it, verify which release you are pulling from Hugging Face, check whether the tokenizer issue applies to that checkpoint, and read the corpus scripts to confirm which documents your target period actually covers.

Frequently asked questions

What is TimeCapsuleLLM?

It is a language model trained from scratch only on documents from a fixed place and time period, in the released versions London between 1800 and 1875, with the goal of reducing modern bias and reproducing the vocabulary and worldview of the era. The repository contains the corpus scripts and one directory per model generation.

Where are the TimeCapsuleLLM models hosted?

The README points to Hugging Face rather than shipping weights in the repository, linking a collection under the name haykgrigorian for the 1800-1875 London models. The dataset is published separately as postgrammar/london-llm-1800.

What are the known limitations of TimeCapsuleLLM?

The README states that v0 produces mostly incoherent sentences, that v0.5 still has a high factual hallucination rate and emits OCR noise such as "Digitized by Google", and that v2mini-eval1 has a tokenization issue that splits words apart in its output.

How is TimeCapsuleLLM trained?

The training scripts and architecture come from other projects rather than being original: v0 and v0.5 build on nanoGPT, v1 builds on Microsoft's Phi 1.5, and v2 builds on LlamaForCausalLM. The corpus is assembled by download_texts_improved.py and london_corpus_dataset.py from Internet Archive document identifiers.

Official sources

  1. haykgrigo3/TimeCapsuleLLM on GitHub
  2. Issues
  3. License: MIT
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/haykgrigo3-timecapsulellm.svg)](https://hysenlabs.com/projects/haykgrigo3-timecapsulellm)