Library / SDK
Doriandarko/texts-to-transformer avatar
Doriandarko/texts-to-transformer

texts-to-transformer: a tiny Transformer trained on your own iMessage history

Train a tiny Transformer from scratch on your iMessage history, entirely on your Mac.

470 stars42 forksPythonMIT

At a glance

What is it?
Doriandarko's texts-to-transformer is a local-first MLX pipeline that snapshots your Mac Messages database, pseudonymizes it, and trains a 1.38M-parameter decoder-only model from random initialization. It is a personal style model, not an assistant.
Who is it for?
Adopt it if you have an Apple Silicon Mac with 16 GB of unified memory, roughly 8M tokens of message history, and you want a small model that imitates your phrasing rather than answers questions. Skip it if you need factual question answering, if your history is short, or if you cannot keep work/ and outputs/ on a FileVault-protected disk.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 85 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What texts-to-transformer actually produces

The README is unusually direct about scope: this builds a small personal style model, not a generally capable assistant. It can learn phrasing, rhythm, slang, and common responses, but the README says it will not reliably reason or answer factual questions. That framing matters, because the name invites the wrong expectation. You are not fine-tuning a foundation model. Nothing is pretrained. The tokenizer and the model both start from zero.

The default small preset is a 4-layer, 1.38M-parameter decoder-only Transformer with a custom 4,096-token byte-level BPE tokenizer and a 256-token context window. A 6.16M-parameter preset is included for unusually large message histories. The README cites an example development run that used roughly 8M training tokens and produced a 1.38M-parameter model generating short replies in the owner's writing style. That is the realistic target: short replies in your register, not paragraphs of reasoning.

The audience is narrow on purpose. You need an Apple Silicon Mac, macOS 14 or newer, and 16 GB of unified memory recommended. The project pins MLX 0.32.0 and does not require PyTorch. If you are on an Intel Mac or Linux, the MLX dependency is the wall.

The pipeline from chat.db to a local reply generator

The README ships a Mermaid flowchart that is worth reading as a contract, because each stage writes to disk and the next stage reads only that artifact. The live Messages database is opened in SQLite read-only mode and never modified. A consistent private backup is written under work/ using SQLite's online backup API, then hashed and checked with PRAGMA quick_check. Processing happens from that backup, never from the live database. Attachments are never opened or copied.

Extraction recovers ordinary text plus Apple typedstream attributedBody text, filters reactions, system events, deleted messages, and attachment-only rows, redacts obvious identifiers, and replaces handles and chat identifiers with keyed HMAC pseudonyms before JSONL is written. Messages are grouped into six-hour conversation sessions, duplicates are removed, and the corpus is split chronologically 90/5/5 with seven-day guard bands. The README states that the pipeline verifies no duplicate session hash appears across splits, and stops instead of silently continuing when recovery or privacy checks fail. That guard-band design is the part most homegrown scripts skip, and it is what keeps near-duplicate conversations out of both train and test.

Tokenization happens only on the training split. The byte-level BPE tokenizer preserves emoji, casing, punctuation, multilingual text, slang, and unusual spelling. Then corpus-stats encodes the splits and writes work/reports/model-selection.json. The README says the project refuses real training below one million tokens and selects the largest preset supported by the corpus, marking a low token-to-parameter ratio as a memorization-prone experiment. Training itself uses next-token cross-entropy, AdamW, learning-rate warmup and cosine decay, gradient clipping, compiled MLX updates, validation-based checkpoint selection, and resumable checkpoints.

Install and first run: doctor, snapshot, prepare

The project uses Python 3.11 and pins MLX 0.32.0. It installs through uv, which the README suggests getting via Homebrew if you do not have it. Clone the repository, sync the environment, and run doctor, which verifies Apple Silicon, MLX Metal support, disk space, Git ignore coverage, private directory permissions, and read-only access to the Messages database.

bash
git clone https://github.com/Doriandarko/texts-to-transformer.git
cd texts-to-transformer

# Skip this if uv is already installed.
brew install uv

uv sync
uv run imessage-mlx doctor

If doctor reports safe_to_snapshot_real_data as false, the README points to System Settings, Privacy & Security, Full Disk Access. Enable the application running the command, completely restart that application, and rerun doctor. The README explicitly warns against copying the live database manually or changing its permissions as a workaround.

Once doctor passes, take the snapshot and inspect the local schema. The README notes that Apple changes the Messages schema between macOS versions, so the extractor inspects your local schema instead of trusting an internet example.

bash
uv run imessage-mlx snapshot --config configs/data.yaml

uv run imessage-mlx inspect-schema \
  --database work/snapshot/chat.db \
  --output work/schema/schema.json

The snapshot command writes work/snapshot/chat.db, hashes it, and runs PRAGMA quick_check. Next, build the dataset and audit it. The commands print aggregate counts and metrics, not message text.

bash
uv run imessage-mlx prepare --config configs/data.yaml
uv run imessage-mlx privacy-audit

After that the sequence is train-tokenizer against work/splits/train.jsonl with --vocab-size 4096, then corpus-stats against work/splits and outputs/tokenizer writing to work/tokens, then train with configs/model-1m.yaml unless model-selection.json explicitly selects model-7m. Evaluate against the untouched test split with --checkpoint outputs/runs/my-model/best. An interrupted run resumes with --resume-from outputs/runs/my-model/last.

Where the design gives up ground

Pseudonymization is not anonymization, and the README says so in those words. Handles and chat identifiers become keyed HMAC pseudonyms, but the message bodies remain, and a 1.38M-parameter model trained on your own text can memorize it. The README's warning is unambiguous: the resulting model may memorize private text and must remain private. Keep work/ and outputs/ on a FileVault-protected Mac and never commit, upload, or share them. Datasets, tokenizers, checkpoints, and final weights are excluded from Git, but that is a default, not a guarantee against your own mistakes.

The one-million-token floor is the other hard edge. Below it, the project refuses real training. That is a sensible guard, and it also means short message histories simply cannot use this pipeline. The README's own example used roughly 8M training tokens for a 1.38M-parameter model. If your history is a few thousand messages, you are below the floor and the tool will not pretend otherwise.

There is also a capability ceiling that no configuration removes. A 256-token context window and a 1.38M-parameter model will not hold a long conversation, follow multi-step instructions, or retrieve facts. The README's own warning that it will not reliably reason or answer factual questions is the honest boundary. If you want a local assistant that answers questions about your notes, this is the wrong tool regardless of how much data you feed it.

How this differs from fine-tuning an existing model

The obvious alternative is fine-tuning or prompting a pretrained model, whether a local one or a hosted one. The difference is not just size. Fine-tuning starts from weights that already encode grammar, world knowledge, and instruction following, and adapts them to your data. texts-to-transformer starts from random initialization and learns everything from your messages alone. That is why the README can promise your phrasing and rhythm but not reasoning: there is no pretrained substrate to reason with.

The trade is legibility. A from-scratch 1.38M-parameter model is small enough to inspect, and the pipeline is built around that: memorization checks, exact train n-gram overlap aggregates in the evaluation report, a unigram baseline, and obvious-PII pattern counts. Fine-tuning a much larger model gives you more capability and far less visibility into what it retained. If your goal is a private model that sounds like you and you accept it cannot answer questions, the from-scratch route is defensible. If your goal is a useful assistant, the pretrained route wins and this project is not competing there.

Maintenance, licence, and what a fork inherits

The licence is MIT, which permits commercial and private use, modification, and redistribution with the copyright notice and permission notice preserved. That is a permissive default and it does not change the privacy obligation, which comes from the data rather than the licence. Nothing here is legal advice; if you plan to redistribute anything derived from your message history, the pseudonymization caveat in the README is the thing to read first.

The repository is not archived, and the last push was on 2026-07-08. The dependency surface is small and pinned: mlx 0.32.0, numpy 2.x, pytypedstream, pyyaml, tokenizers 0.22.x, and typer 0.16.x, with pytest, pytest-cov, and ruff as dev dependencies. Python is constrained to >=3.11,<3.12. That pinning is good for reproducibility and it is also the upgrade cost: MLX moves, and the MLX pin means a fork that wants a newer MLX has to retest training and compiled updates. The Makefile exposes sync, lint, format, test, and check targets, with check running lint, pytest, ruff format --check, and git diff --check, so the repository ships its own verification loop. The schema-inspection step exists because Apple changes Messages between macOS releases, which is the maintenance burden you inherit with the data source rather than the model code.

Editorial conclusion

Adopt it if you have an Apple Silicon Mac with 16 GB of unified memory, roughly 8M tokens of message history, and you want a small model that imitates your phrasing rather than answers questions. Skip it if you need factual question answering, if your history is short, or if you cannot keep work/ and outputs/ on a FileVault-protected disk. Before training, run uv run imessage-mlx doctor and confirm safe_to_snapshot_real_data is true, then read docs/privacy.md, because the README states plainly that the model may memorize private text and must remain private.

Frequently asked questions

Can I create my own transformer with texts-to-transformer?

Yes, and that is the whole point of the project. It trains a decoder-only Transformer from random initialization on your iMessage history, with a custom 4,096-token byte-level BPE tokenizer, using MLX on Apple Silicon.

Does texts-to-transformer upload my iMessage data anywhere?

The README states that no command uploads data or sends an iMessage, and that chat only prints a suggestion in the terminal. The pipeline works from a private snapshot under work/ rather than the live database.

How much iMessage history do I need before texts-to-transformer will train?

The README says the project refuses real training below one million tokens and selects the largest preset supported by the corpus. Its example development run used roughly 8M training tokens for a 1.38M-parameter model.

Can texts-to-transformer answer factual questions about my messages?

No. The README warns that it will not reliably reason or answer factual questions, and describes the output as a small personal style model that learns phrasing, rhythm, slang, and common responses.

What Mac does texts-to-transformer require?

An Apple Silicon Mac with an M1 or newer chip running macOS 14 or newer, with at least 16 GB of unified memory recommended. The project pins MLX 0.32.0 and does not require PyTorch.

Official sources

  1. Doriandarko/texts-to-transformer on GitHub
  2. Issues
  3. License: MIT
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/doriandarko-texts-to-transformer.svg)](https://hysenlabs.com/projects/doriandarko-texts-to-transformer)