Model or dataset
AutoArk/TinyEngram avatar
AutoArk/TinyEngram

TinyEngram: two Python files, a pinned CUDA index, and a composability claim about hash collisions

Research of DeepSeek Engram Architecture based on Qwen-3 and Stable Diffusion series.

1,380 stars89 forksPythonLicense varies

At a glance

What is it?
TinyEngram is a research codebase that reimplements DeepSeek's Engram architecture on Qwen and extends it to Stable Diffusion, so that visual concepts can be injected through the text encoder without touching the image backbone. The repository is small and unusually honest about its own shape: the architecture lives in two files at the root, every dependency is pinned with an equals sign, and the strongest claim in the file is argued from mechanism rather than from a table.
Who is it for?
TinyEngram is worth reading if you are interested in memory injection as an alternative to low-rank adaptation and want a codebase small enough to hold in your head, because the implementation is two Python files at the root and the surrounding directories are training scaffolding rather than library code. Three things to check before you try to run it.
Can I use it commercially?
Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
Is it still maintained?
Yes. The repository last received commits 136 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The composability claim is argued from hash collisions, not from a table

The strongest statement in the file is that memories strictly do not interfere with each other, and it rests on a single mechanism: Engram relies on exact N-gram matching, described in parentheses as hard hash collisions. Because a memory only fires when its exact phrase appears in the prompt, stacking thousands of them is presented as safe, with zero degradation to the base model's general capabilities claimed as the result.

That is an argument from construction rather than from evaluation. If two memories cannot both match the same n-gram, they cannot fire together, and the interference question answers itself. The claim is also scoped to exact matching, which means the safety property is a property of the lookup, not a learned separation: anything that changes the n-gram, shortens the name, or introduces a near miss moves the problem outside the argument.

The rest of the same section is more modest and more useful. Minimal and surgical describes building a vocabulary for one target phrase, which is a constraint rather than a feature, since the vocabulary is specific to what you are injecting. The example given is a caption about information injection in text to image work.

No table, figure or ablation sits beside the non-interference claim in the file. The ablation studies announced in the timeline are described as parameter ablation studies with convergence observations, which is a different measurement.

The dependency file pins one CUDA build and hands you the fallback

Every line of the training requirements is pinned with an equals sign, and the first significant line is not a package at all:

bash
conda create -n tinyengram python=3.10 -y
conda activate tinyengram
pip install --upgrade pip
pip install -r requirements.txt

The requirements file opens with an extra index URL pointing at the PyTorch wheel index for a specific CUDA build, cu126. The PyTorch stack below it is pinned to three matching wheels at 2.8.0, 0.23.0 and 2.8.0 for torch, torchvision and torchaudio.

The comment above those pins says the file was tested with Python 3.10 on CUDA 12.x GPUs, and the comment below the index URL says that if your CUDA runtime is not compatible with cu126 you should install the matching torch, torchvision and torchaudio wheels from pytorch.org instead. So the compatibility story is: one tested configuration, and a manual repair instruction for everything else.

The extra index URL is the mechanism that makes the pins resolve to CUDA builds at all, and it is also the part most likely to surprise you. A pip install on a machine with a different CUDA runtime will still succeed and then fail at import time or at device initialisation, which is a worse failure than a resolver error. The setup path does not detect the mismatch for you.

The architecture is two Python files and everything else is scaffolding

The top level of the repository reads like a small experiment rather than a library. Two Python modules sit directly in the root, one named for the architecture and one named for the model it runs on. Around them sit directories for training, evaluation scripts, vision work, data and documentation, plus two separate requirement files and a DeepSpeed configuration at the root.

The naming carries the whole design. The architecture is not a package with a public interface, it is a module you import, and the model-specific adaptation is a second module beside it rather than a configuration of the first. That is what makes the codebase readable end to end, and it is also why there is nothing to install as a dependency.

The two requirement files split the problem in two. The one shown covers direct dependencies for training and the vision demos, and a second file at the root is reserved for evaluation. The setup instructions keep both out of the quick path, pointing instead to a document under the reproduction directory for CUDA notes and the optional evaluation dependencies.

The DeepSpeed configuration at the root, named for the ZeRO stage two, matches a pinned DeepSpeed in the requirements. ZeRO two shards the optimizer and gradient state across devices, which is the standard choice for fine-tuning where the model weights are still replicated, and its presence tells you the training script expects more than one GPU to be useful.

The vision path injects through the text encoder and leaves the image model frozen

The extension to Stable Diffusion treats visual concepts as memories that can be injected into the text encoder. The mechanism is described precisely: the system recognises specific N-grams in the prompt and injects learned embeddings that guide generation, all without fine-tuning the massive U-Net or DiT backbone.

Two design consequences follow from that choice. The backbone stays frozen, so the original weights are never overwritten and there is no risk of degrading general image quality through a full fine-tune. And the injection point is the text encoder, which is a much smaller surface than the denoiser, which is why the section calls the approach lightweight and composable.

The stated use is teaching the model new subjects, with specific characters given as the example, while keeping the original weights frozen. The writeup covers two model families, SD1.5 and SD3.5, and the documentation points to a technical report available both as a PDF in the repository and as an arXiv entry.

The vision stack is pinned separately from the language stack in the same requirements file: a diffusers version, a modelscope version, Pillow, sentencepiece and protobuf. TensorBoard and TensorBoardX are pinned as a pair, so a training run can log to either.

LoRA is the baseline, and peft is the dependency that defines it

Both headline findings are comparisons rather than absolute claims. Engram is presented as a parameter efficient fine-tuning method, and separately as outperforming LoRA in catastrophic forgetting. The summary line states the comparison in both directions: better parameter efficiency and better resistance to forgetting.

The baseline is not hand-rolled. The low-rank adaptation path comes from a pinned parameter-efficient fine-tuning library in the requirements, so the comparison is against that library's implementation at a fixed version rather than against a strawman written for the experiment.

Catastrophic forgetting is the more interesting of the two claims to think about, because it is the failure mode people actually hit when fine-tuning: a model that learns the new task and loses the old one. The timeline records the comparison as added on 2026-01-30, the same day as the parameter ablation studies, so both analyses landed together, one week after the initial commit.

What the file does not do is show the numbers. The summary asserts the direction of both comparisons and points to the report, and the reproduction scripts for the Engram versus LoRA experiment were released separately on 2026-02-02. So the claim is checkable, and the checking is deferred to code and a paper rather than presented in the overview.

Four months of work, then the announcements box stops

The README carries its own changelog inside a styled panel, and that panel is set to a fixed height with a scrollbar rather than being allowed to grow the page. Six entries fit inside it. The most recent is dated 2026.05.20 and announces that the vision technical report is ready, with links to a PDF in the repository and to an arXiv version.

Reading downward, the entries chart the scope of the project rather than a release history. On 2026.02.12 the vision work landed, described as injecting visual concepts into Stable Diffusion through Engram. On 2026.02.02 the reproduction scripts for the Engram versus LoRA experiment were released. On 2026.01.30 came both the catastrophic forgetting comparison and the parameter ablation studies. And on 2026.01.23 there was an initial commit.

So the whole body of work runs from late January to late May 2026. The last push to the default branch was 2026-05-21, and the repository has no GitHub releases, so there are no version tags to anchor a dependency on. The panel is dated rather than versioned, which suits a research log and gives a consumer nothing to pin.

No licence and no homepage sit alongside a link to an upstream Engram repository

Two repository fields are empty. No licence is recorded, and no homepage is set. The header does link to a license file in the repository, so a licence document is present even though none is declared in the metadata, and the two cannot both be taken as the whole story.

For a research codebase the licensing question has a specific shape. The architecture being reimplemented is Engram, credited in the introduction and linked to an upstream repository under the DeepSeek organisation, and the base model is Qwen, linked to its own repository. Neither of those is the same as the licence on this repository, and the file does not say which terms apply to code derived from an upstream implementation.

That makes it worth reading the license file directly before reusing any of it. It also frames the rest of the project accurately. The introduction describes a lightweight, ready-to-train codebase for anyone to reproduce, experiment with, or extend Engram-style models, and says the repository is both a toolkit and a living research notebook. A notebook invites copying; a toolkit invites depending. Those want different licensing clarity, and this repository currently offers neither field.

Editorial conclusion

TinyEngram is worth reading if you are interested in memory injection as an alternative to low-rank adaptation and want a codebase small enough to hold in your head, because the implementation is two Python files at the root and the surrounding directories are training scaffolding rather than library code. Three things to check before you try to run it. The environment is pinned hard to one CUDA build, with an extra index URL selecting cu126 wheels and a comment telling you to install matching wheels by hand if your runtime differs, so a machine on a different CUDA version needs manual work before the first import. The claims about composability and non-interference are argued from exact N-gram matching and hard hash collisions, which is a mechanism argument and not a measurement, so read the report rather than the summary if the claim matters to you. And note the cadence: the first commit is dated 2026-01-23, the vision report landed 2026-05-20, and the last push was 2026-05-21, with no release tags at all, so this is a notebook that stopped rather than a library on a support schedule. No licence is recorded for the repository and no homepage is set, so resolve the licensing question before you build anything on it.

Frequently asked questions

What is DeepSeek Engram, in TinyEngram's terms?

It is described as an LLM enhancement that boosts phrase-level understanding by integrating a compact N-gram memory module and a gated retrieval mechanism into key transformer layers. TinyEngram reimplements that architecture on Qwen as a lightweight, ready-to-train codebase, and points to the upstream Engram repository for the original.

How do I set up the TinyEngram training environment?

Create a conda environment on Python 3.10, activate it, upgrade pip, and install from the pinned requirements file with pip install -r requirements.txt. That file pins the PyTorch stack to specific versions and selects CUDA wheels through an extra index URL, so a machine on a different CUDA runtime needs the matching wheels installed by hand.

Does TinyEngram modify the Stable Diffusion backbone?

No. The stated approach injects learned embeddings into the text encoder after recognising specific N-grams in the prompt, and the file says this happens without fine-tuning the U-Net or DiT backbone. The purpose is to teach subjects such as specific characters while the original weights stay frozen.

What does TinyEngram compare itself against?

Against LoRA, on two axes: parameter efficiency and resistance to catastrophic forgetting. The low-rank baseline comes from a pinned parameter-efficient fine-tuning library rather than a hand-written one, and reproduction scripts for the comparison were released separately from the initial commit.

Can I install TinyEngram as a Python package?

There is no installable package here. The implementation sits in two Python modules at the root of the repository, one for the architecture and one for the model it runs on, and the surrounding directories are training, evaluation, vision and data folders. The setup path is a conda environment plus the pinned requirements file.

Official sources

  1. AutoArk/TinyEngram on GitHub
  2. Issues
  3. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/autoark-tinyengram.svg)](https://hysenlabs.com/projects/autoark-tinyengram)