# texts-to-transformer trains a 1.38M-parameter model on your iMessage history, and refuses to train below a million tokens

> A from-scratch decoder-only Transformer pipeline that snapshots chat.db read-only, pseudonymises identities with keyed HMACs, splits sessions chronologically with guard bands, and trains on MLX without PyTorch. Honest about being a style model rather than an assistant, and explicit that pseudonymisation is not anonymisation.

**Doriandarko/texts-to-transformer** — Train a tiny Transformer from scratch on your iMessage history, entirely on your Mac.

- Repository: https://github.com/Doriandarko/texts-to-transformer
- Stars: 470 · Forks: 42
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/doriandarko-texts-to-transformer

## The default preset is four layers, 1.38M parameters, and a 256-token context

The scale is stated before anything else, which is unusual and welcome. The default small preset is a four-layer, 1.38M-parameter decoder-only Transformer paired with a custom 4,096-token byte-level BPE tokenizer and a 256-token context window. A larger 6.16M-parameter preset is included for unusually large message histories.

Nothing is pretrained. Both the tokenizer and the model start from zero, which means the byte-level BPE vocabulary is learned from your messages rather than inherited from a public corpus. That has an upside beyond novelty: because the vocabulary is byte-level and trained on your data, it preserves emoji, casing, punctuation, multilingual text, slang, and unusual spelling without a vocabulary that silently maps your idioms to unknown tokens.

The project is direct about what that buys you. Its own warning says this builds a small personal style model and not a generally capable assistant, that it can learn phrasing, rhythm, slang, and common responses, and that it will not reliably reason or answer factual questions. It adds that the resulting model may memorise private text and must remain private.

For calibration, an example development run used roughly eight million training tokens and produced a 1.38M-parameter model that generated short replies in the owner's writing style. That is the claim being made: recognisable stylistic echo, not capability.

The pipeline itself runs as a chain, from the Messages database through a private snapshot, extraction and pseudonymisation, conversation sessions, chronological splits, tokenizer training, random initialisation of the Transformer, MLX training and evaluation, and finally a local reply generator.

## The live chat.db is never opened for writing, and the copy is hashed

The privacy design is specific enough to audit, which is rarer than a privacy page. The Messages database at `~/Library/Messages/chat.db` is opened in SQLite read-only mode and is never modified. Processing happens from a consistent private backup under `work/`, never from the live database.

The snapshot step uses SQLite's online backup API rather than a file copy, writes to `work/snapshot/chat.db`, hashes the snapshot, and runs `PRAGMA quick_check`. That combination addresses the two things that actually go wrong when copying a live database: catching a torn copy mid-write, and not knowing whether what you got is what you hashed. The integrity check and the hash are both recorded so the copy can be verified later.

Attachments are never opened or copied, which is both a privacy measure and a performance one, since the bulk of an iMessage history is media.

Identity handling happens before anything is written. Handles and chat identifiers are replaced with keyed HMAC pseudonyms before the JSONL is produced, and URLs, email addresses, and phone-number-shaped strings are redacted by default. A keyed HMAC rather than a plain hash is the correct primitive here, because a plain hash of a handle is trivially reversible by guessing the identifier space, while a keyed one requires the key.

Two further guarantees are stated plainly. Raw messages are never printed in normal logs; the commands print aggregate counts and metrics instead. And no command uploads data or sends an iMessage, with the `chat` command only printing a suggestion in your terminal.

## doctor checks six things, and the Full Disk Access loop is deliberate

Before any data is touched, one command verifies the environment:

```bash
git clone https://github.com/Doriandarko/texts-to-transformer.git
cd texts-to-transformer

# Skip this if uv is already installed.
brew install uv

uv sync
uv run imessage-mlx doctor
```

The doctor checks six specific things: that the machine is Apple Silicon, that MLX Metal support is present, that there is enough disk space, that Git ignore coverage is in place, that the private directories have the right permissions, and that the Messages database is readable in read-only mode. Running that list before touching a personal database rather than after is the right order.

Full Disk Access is where people get stuck, and the documentation handles the failure mode explicitly. If a flag called `safe_to_snapshot_real_data` reads as false, you are told to open System Settings, then Privacy and Security, then Full Disk Access, enable the application running the command, completely restart that application, and rerun the doctor.

Three details in that instruction are the ones people usually miss. You have to restart the application completely, because the permission is read at process start. You rerun the doctor rather than the snapshot, so you find out before the real step. And there is an explicit prohibition: do not copy the live database manually or change its permissions as a workaround. That last line is doing real work, since loosening permissions on your Messages database to route around macOS is exactly the kind of change that persists after the project is uninstalled.

## The extractor reads your local Messages schema instead of trusting an example

There is a dedicated step for inspecting the schema, and the reason given is the one that matters: Apple changes the Messages schema between macOS versions, so the extractor inspects the local schema rather than assuming that an example found on the internet is correct. This is the single most common way a project like this breaks, silently, on the day you run it.

Recovering the text is itself non-trivial, and the dependency list explains why. Ordinary text is one path; a large share of messages instead carry their body in an Apple typedstream structure stored in a column called `attributedBody`, which is a compact binary format. A package called `pytypedstream` is used to decode it, and the prepare step explicitly recovers both ordinary text and typedstream text.

The preparation stage then does seven things in order. It recovers both text forms. It filters out reactions, system events, deleted messages, and attachment-only rows. It redacts obvious identifiers and pseudonymises database identities. It groups messages into six-hour conversation sessions. It removes duplicate sessions. It creates chronological 90/5/5 splits with seven-day guard bands. And it verifies that no duplicate session hash appears across splits.

That last group is the part worth dwelling on. Splitting by time rather than at random, leaving a seven-day gap between train, validation, and test, and then checking for a session hash that appears in two splits is a deliberate defence against leakage, because near-duplicate conversations either side of a boundary inflate every metric. And the command stops rather than silently continuing when a recovery or privacy check fails, which means a failed audit is a hard stop rather than a warning you scroll past.

## Model size is chosen for you, and a bad token-to-parameter ratio gets labelled

You do not pick the architecture by hand. A corpus statistics step runs against the splits and writes a report, and you are told to read `work/reports/model-selection.json`. The project then does two things with that measurement: it refuses real training below one million tokens, and it selects the largest preset the corpus can support.

The second guard is subtler and more useful. A low token-to-parameter ratio is explicitly marked as a memorisation-prone experiment rather than left for you to infer from the loss curve.

That ratio is worth doing the arithmetic on, because it explains the project's insistence on evaluating overlap separately. The example run had roughly eight million training tokens against 1.38 million parameters, which is on the order of six tokens per parameter. Published language models are trained on orders of magnitude more data per parameter than that. A model in that regime is in the range where next-token prediction is efficiently solved by memorisation, which is exactly what the label warns about.

Training itself is conventional and competently specified: next-token cross-entropy, AdamW, learning-rate warmup with cosine decay, gradient clipping, compiled MLX updates, validation-based checkpoint selection, and resumable checkpoints. Interrupted runs continue from a `last` checkpoint via a `--resume-from` flag rather than starting over, which matters when you are training on a laptop and the process will get suspended.

The tokenizer step is separate and comes first, trained only on the training split, so the validation and test text never influences the vocabulary.

## Evaluation reports overlap and PII counts without saving your messages

The evaluation step reads the untouched test split and produces a report whose contents are the point. It includes overall perplexity and a `me`-turn perplexity, which separates your own writing from everyone else's. It includes a unigram baseline, so you can tell whether the model beats just predicting the most frequent words in your history. It includes exact train n-gram overlap aggregates, which is the memorisation check: how often the model reproduces training text verbatim. And it includes counts of obvious personal-data patterns.

The privacy property that makes this safe to share a screenshot of is stated directly. Matching private text is never persisted in the report. You get the number, not the string that produced it.

That design is the honest answer to the tension the project names in its own warning. A model trained on your messages will, at some point, emit something from your messages, and there is no configuration that prevents it. What the tooling can do is make the risk measurable, and an aggregate overlap figure you can watch across runs is more actionable than a promise.

A separate `privacy-audit` command runs as its own step after preparation, which keeps the audit independent of the command it audits. Export follows evaluation, and then there is a local terminal chat interface that prints suggestions rather than sending anything, so the model stays a suggestion engine and not a messaging client.

One caveat on scope: the final export step and the chat section are cut off in this copy of the documentation, so the export flags and the exact invocation are not readable here.

## Python is pinned to a single minor version and MLX to an exact one

The dependency constraints are unusually tight in one specific place, and it matters for anyone setting this up. The project declares `requires-python` as greater than or equal to 3.11 and less than 3.12, which is not a floor but a single-minor-version cage. A machine on 3.12 or 3.13 is refused rather than warned about. A `.python-version` file at the repository root reinforces the intent.

MLX is pinned to an exact version with no range at all, while the other five runtime dependencies use bounded ranges: numpy from 2.0 to below 3, the typedstream decoder from 0.1 to below 1, PyYAML from 6.0 to below 7, the tokenizers library from 0.22 to below 0.23, and the CLI framework from 0.16 to below 1. The asymmetry is defensible, since MLX is the one dependency whose internals the training loop compiles against, and a lock file is committed alongside.

PyTorch is not a dependency. That is the point of using MLX on Apple Silicon, and it is why the requirement list names an M1 or newer and at least 16 GB of unified memory as recommended rather than asking for a discrete GPU.

The repository is small and conventionally organised: source under `src/`, a `configs/` directory holding the YAML the commands reference, `docs/`, `tests/`, and top-level `LICENSE`, `SECURITY.md`, and `Makefile`. The build backend is Hatchling and the CLI entry point is a single `imessage-mlx` command.

The Makefile is thin, with `sync`, `lint`, `format`, and `test`, and one target worth noting. `check` depends on lint and test, and then additionally verifies formatting and runs `git diff --check`, which catches whitespace errors including the trailing ones that matter when a file has to stay diff-clean. The version is 0.1.0, there are no GitHub releases, and the last push to the default branch was 8 July 2026.

## Conclusion

texts-to-transformer fits someone who wants to understand what training a language model actually involves, on data they already own, with a hard ceiling on scale. The privacy engineering is the reason to read it even if you never run it: read-only snapshot, hashed copy, integrity check, keyed pseudonyms, Git-excluded artefacts, and an explicit statement that none of that is anonymisation. Two things to check first. Whether your Python is 3.11, since the project pins the interpreter to that single minor version with an upper bound and will refuse anything else. And whether your history is large enough, because the pipeline will not train below one million tokens and marks a poor token-to-parameter ratio as memorisation-prone. Treat the output as a private artefact in the literal sense: a 256-token window trained on your own messages will reproduce your phrasing, and the documentation says so plainly rather than implying the privacy controls change that.

## FAQ

### What does texts-to-transformer actually build?

A decoder-only Transformer trained from random initialisation on your iMessage history, plus a tokenizer trained from zero. The default preset is four layers and 1.38M parameters with a 4,096-token byte-level BPE tokenizer and a 256-token context window; a 6.16M-parameter preset is included for unusually large histories.

### Does texts-to-transformer modify my Messages database?

No. The database at ~/Library/Messages/chat.db is opened in SQLite read-only mode and is never modified. The snapshot uses SQLite's online backup API to write work/snapshot/chat.db, then hashes the copy and runs PRAGMA quick_check, and all processing happens from that private copy.

### What does `uv run imessage-mlx doctor` verify?

Six things: that the machine is Apple Silicon, that MLX Metal support is available, that there is enough disk space, that Git ignore coverage is in place, that the private directories have correct permissions, and that the Messages database is accessible read-only.

### What hardware and software does texts-to-transformer need?

An Apple Silicon Mac (M1 or newer), macOS 14 or newer, at least 16 GB of unified memory recommended, and Full Disk Access for whichever application runs the snapshot command. It requires Python 3.11, pins MLX to 0.32.0, and does not require PyTorch.

### How does texts-to-transformer avoid train and test leakage?

Messages are grouped into six-hour conversation sessions and split chronologically 90/5/5 with seven-day guard bands between the splits. The pipeline then verifies that no duplicate session hash appears across splits, and stops rather than continuing if a recovery or privacy check fails.

### Is the trained model private?

The documentation says pseudonymization is not anonymization, and that work/ and outputs/ must stay on a FileVault-protected Mac and never be committed, uploaded, or shared. Datasets, tokenizers, checkpoints, and weights are Git-excluded, and no command uploads data or sends an iMessage.

## Sources

- [Doriandarko/texts-to-transformer on GitHub](https://github.com/Doriandarko/texts-to-transformer)
- [Issues](https://github.com/Doriandarko/texts-to-transformer/issues)
- [License: MIT](https://github.com/Doriandarko/texts-to-transformer/blob/main/LICENSE)
- [README](https://github.com/Doriandarko/texts-to-transformer/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/doriandarko-texts-to-transformer
