Model or dataset
ai-in-pm/Titans---Learning-to-Memorize-at-Test-Time avatar
ai-in-pm/Titans---Learning-to-Memorize-at-Test-Time

This repository asks seven models about Titans, it does not implement Titans

Multi-agent demo platform for Titans (arXiv:2501.00663) — neural networks that learn to memorize at test time. 7 AI agents, native desktop UI.

289 stars53 forksPythonMIT

At a glance

What is it?
The ai-in-pm Titans platform is a multi-agent desktop demo that puts seven commercial language models behind a role each, derived from the three ways the paper integrates a memory module into a transformer. It is a discussion harness with telemetry and charts, not a memory architecture: there is no model, no memory module and no training code in the repository, and one of the seven agents is named for benchmarking it does not run.
Who is it for?
This platform fits someone teaching or writing about test-time memory who wants seven models answering the same question side by side, since the value is the comparison and the streamed transcript, not any capability the models do not already have. It does not fit anyone looking for a Titans implementation, and nothing in the repository is one.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 110 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The seven roles are the paper's three integration modes plus four job titles

The agent table is the design, and it is worth reading as two overlapping axes. The first three agents are named for the paper's actual taxonomy. Memory as Context prepends memory tokens to the attention context so the model reads from a persistent external store. Memory as Gate multiplies the attention output by the memory output, so memory controls information flow rather than feeding it. Memory as Layer inserts the memory module as a standalone layer in the network stack. The other four are not architectural roles at all: an experimental validation agent, an innovations agent, and two analysis agents, one of which is explicitly a cross-agent synthesis. So three of seven agents have a defined technical referent and four are editorial. Pairing each with a different vendor, from OpenAI through Anthropic, Mistral, Groq, Google, Cohere and Emergence, means the comparison you get is a comparison of model personality on an architecture prompt, not a comparison of memory mechanisms.

The dependencies list Streamlit, the interface is Tkinter, and three entry points sit at the root

The requirements file and the interface described in the documentation do not match, and the repository layout explains why rather than excuses it. The dependency list includes a web application framework and a plotting library alongside the data stack and seven vendor clients, which suggests a browser-based interface. The documented interface is a native desktop one built with a toolkit that ships in the standard library, with a selector panel, a streamed console, live telemetry, a chart and a split pane that remembers its position across sessions. Meanwhile the top level of the repository contains three candidate entry points: the desktop application entry, a second module with a graphics-suffixed name, and an application module alongside a static assets directory. That is a codebase partway through a migration between two UI approaches, with both left in place. For a reader, the practical consequence is that the documented launch command may not be the only way in, and the web dependency may be vestigial or may be the newer path.

The paper PDF and a prebuilt Windows binary both sit in the repository root

Two files at the top level tell you what kind of repository this is. The paper itself is committed as a PDF next to the licence and the readme, so the primary source is available without a trip to the preprint server, which is convenient for teaching and means the citation in the documentation is verifiable offline. Alongside it is a prebuilt Windows executable, and the documentation says the executable gives a bundled experience that needs no Python at all, while a batch launcher handles path setup for people who do have Python. Shipping a compiled binary in a source repository has obvious costs: it is an opaque artefact that no reviewer can read, it is platform-locked, and it has to be rebuilt whenever the code changes or it silently becomes a different program from the source beside it. The documentation does at least give a reason for the batch file, which is that running through it surfaces the terminal error output when the application closes immediately on launch. That is the honest version of this pattern: keep the launcher, and treat the binary as a convenience rather than the product.

There is one release, named after the repository rather than a version

The release history has a single entry, and its name is the repository name with a suffix. There is no semantic version anywhere in it. That tells you the release was cut automatically rather than chosen, which is consistent with a project that distributes by cloning the branch and running an entry point rather than by installing a pinned artefact. It also means a reader cannot say which version they have without comparing commits, and cannot upgrade to a specific version because there is nothing to upgrade to. The last push was on 2026-06-15, which is over three months before this article was written, so the project is not moving quickly either. The absence of open issues is consistent with a small demonstration rather than a tool people depend on. None of that is criticism of a teaching demo; it is the practical boundary of what you can build on top of it, and anyone planning to extend this codebase should fork it rather than wait for a release train.

Graceful degradation is the design: any subset of seven keys is a valid configuration

The key configuration is the best-engineered part of the documentation, and the reason is stated twice. The environment template says the application works with any subset of keys and that unavailable agents are shown with a status message, and the configuration section repeats that you do not need all of them. The whole path from a fresh clone to a running window is four steps:

bash
git clone https://github.com/ai-in-pm/Titans---Learning-to-Memorize-at-Test-Time.git
cd Titans---Learning-to-Memorize-at-Test-Time

pip install -r requirements.txt

cp .env.sample .env
# Edit .env and add your API keys (only the providers you want to use)

python main.py

So the worst case is that you configure one vendor and watch one agent answer. That is a real design choice rather than a convenience, because seven vendor clients means seven different authentication models, seven different response formats and seven different rate-limit behaviours, and any of them can fail on a given afternoon. The troubleshooting table names the symptom for the common case directly: an agent shows as unavailable when that provider's key is missing or invalid in the environment file. The other three entries are equally unglamorous, and one of them is a deprecation warning from a Google client library that the documentation explicitly calls non-fatal.

The surprise metric is described but never implemented here

This is the point to be clear about before anyone builds on the repository. The documentation explains the paper's central mechanism accurately: the memory module updates its own parameters during inference, driven by a surprise metric, which is what lets the model hold on to information that contradicts what it already believes, with no additional training. That is a description of someone else's architecture. What the repository contains is seven vendor client modules, one per provider, and a desktop shell that streams their output. There is no memory module, no test-time update loop, no surprise computation, no transformer, and no trained weights. The agent named experimental validation is described as covering benchmarking and ablation analysis, which in practice means it is asked to write about those things. So this is a prompt-level exploration of an idea, and a genuinely useful one for a group that wants to compare how seven models reason about memory architectures, but it is not a reproduction and it is not a benchmark.

The troubleshooting table teaches you to read the error, not to avoid it

Four rows, and three of them are about getting a message out of a program that is not giving you one. If the application closes immediately on launch, run it through the batch launcher, which is described as handling path setup automatically and which is implicitly where the terminal output goes. If the direct launch fails with a path error, change into the project directory first, which is a reminder that this is a relative-path application rather than an installed one. The third row is the interesting one, because it is an admission rather than a fix: the Google generative client raises deprecation warnings, and the documentation says they are non-fatal and the application still works correctly. That is the correct way to handle a dependency that has moved on and cannot yet be replaced, and it is also a signal about the project's age. A dependency pinned only by a lower bound, with a deprecation already visible, is the normal state of a demonstration codebase that gets attention in bursts rather than continuously.

Editorial conclusion

This platform fits someone teaching or writing about test-time memory who wants seven models answering the same question side by side, since the value is the comparison and the streamed transcript, not any capability the models do not already have. It does not fit anyone looking for a Titans implementation, and nothing in the repository is one. Before you rely on it for a serious comparison, note that the agent named for experimental validation produces prose about benchmarking rather than benchmark results, so treat the interface as a way to read seven answers and not as evidence that any of them is right.

Frequently asked questions

Does the ai-in-pm Titans repository implement test-time memory?

No. It contains seven vendor client modules and a desktop shell that streams their output, and the documentation describes the paper's architecture, including test-time parameter updates driven by a surprise metric, without implementing any of it. There is no memory module, no test-time update loop and no trained model in the repository.

How many API keys do I need to run the Titans platform?

Any subset. The environment template and the configuration section both say the application works with any subset of keys and shows a status for agents whose provider is unavailable. The launch is a clone, a dependency install, a copied environment file and one command, so a single configured vendor gives you one working agent.

Which language model providers does the Titans platform talk to?

Seven, one agent each: OpenAI, Anthropic, Mistral, Groq, Google, Cohere and Emergence. Each has its own module in the agents directory, and each agent is assigned a role drawn from the paper's taxonomy where one exists, such as memory as context, memory as gate and memory as layer, plus four editorial roles for validation, innovation and analysis.

What interface does the Titans desktop application use?

A native desktop one built with a toolkit that ships in the standard library, providing an agent selector, a live streamed console, per-agent timing and token telemetry, a chart extracted from agent output with a scrub control, a cross-agent insights panel, and a split pane that remembers its position between sessions. The dependency list also includes a browser application framework and a plotting library, and the repository root holds three candidate entry-point modules, so the interface appears to be partway through a migration.

How do I run the Titans platform on Windows?

Use the batch launcher, which handles path setup automatically and is also where terminal error output appears if the application closes on launch, or run the bundled executable for a dependency-free experience that needs no Python. If the direct launch fails with a path error, change into the project directory first.

Official sources

  1. ai-in-pm/Titans---Learning-to-Memorize-at-Test-Time on GitHub
  2. Issues
  3. License: MIT
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/ai-in-pm-titans-learning-to-memorize-at-test-time.svg)](https://hysenlabs.com/projects/ai-in-pm-titans-learning-to-memorize-at-test-time)