# AgentOS: A TypeScript Agent Framework With Cognitive Memory and Runtime Tool Forging

> AgentOS is an Apache-2.0 TypeScript framework for agents that keep a session transcript, decay memories on a schedule, and write new tools mid-run. Here is what the README and package layout actually commit to, and where the design gets awkward.

**framerslab/agentos** — TypeScript AI agent framework: cognitive memory, runtime tool forging, multi-agent orchestration, 11 LLM providers.

- Repository: https://github.com/framerslab/agentos
- Website: https://docs.agentos.sh
- Stars: 673 · Forks: 97
- Language: TypeScript
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/framerslab-agentos

## What AgentOS Solves, and Who It Is Actually For

Most agent libraries give you a loop: prompt the model, run a tool, feed the result back. The state that makes the agent feel like it knows you lives outside that loop, in whatever database you bolted on. AgentOS takes the position that memory, tool availability and personality are framework concerns, not application concerns.

The README frames the project as agents that "remember, adapt, and write their own tools." Concretely, that means three subsystems ship together: a persistent cognitive memory layer, a runtime tool forging path, and a dispatch interface that the README says spans 11 LLM providers. The package description in package.json adds multi-tier guardrails, a voice pipeline, and graph orchestration.

Who is this for? The install example is a tutoring agent with a personality object and a memory config, which tells you the intended entry point is application code, not infrastructure. If you are writing TypeScript and you want a single dependency that owns the agent's long-term state, this is the pitch.

It is a worse fit if you already have a memory service you trust. AgentOS will want to own that layer, and the README does not describe an adapter interface for swapping in an external memory store.

## Cognitive Memory: Eight Mechanisms and What They Cost You

The memory subsystem is the most distinctive part of the project. The README lists eight neuroscience-backed mechanisms and names four: Ebbinghaus decay, retrieval-induced forgetting, reconsolidation, and source-confidence decay.

These are not cosmetic labels. Ebbinghaus decay implies memories lose strength over time unless reinforced. Retrieval-induced forgetting implies recalling one memory can suppress a competing one. Reconsolidation implies a retrieved memory can be rewritten before it is stored again. Source-confidence decay implies the agent tracks how much it trusts where a fact came from, and that trust erodes.

That combination produces behaviour you cannot get from a plain vector store. It also produces behaviour that is hard to debug. If the agent forgets something, the cause could be decay, suppression from a competing retrieval, or a low source-confidence score. The README does not document a per-memory audit trail that would let you tell those apart.

The memory types are configurable. The install example passes `memory: { types: ['episodic', 'semantic'], working: { enabled: true } }`, so episodic and semantic stores are selectable and a working-memory tier can be turned on separately.

Benchmark claims appear as badges: the README cites 85.6% on LongMemEval-S at $0.0090 per correct answer with gpt-4o, and 70.2% on LongMemEval-M, described as the only open-source library above 65% on M with reproducible methodology. Those numbers come from the project's own bench repository, so treat them as the project's claim rather than an independent result.

## Runtime Tool Forging and the node:vm Sandbox

The second distinctive mechanism is that an agent can write a new tool during a session. The README describes the flow precisely: the agent writes a TypeScript function with a Zod schema, an LLM judge approves it, and it runs in a hardened `node:vm` sandbox before joining the catalog for the rest of the session.

Four steps, four places things can go wrong. The generated function can be syntactically valid and semantically wrong. The judge can approve something it should not. The sandbox can be less hardened than the word suggests. And the forged tool persists for the session, which means a bad tool keeps getting called until the session ends.

That last point is the real design trade-off. Session-scoped persistence is a reasonable middle ground between one-shot code execution and writing to a permanent registry, but the README does not document a way to revoke a forged tool mid-session or to list what has been forged so you can inspect it.

The README points to `examples/emergent-hierarchical-spawning.mjs` as a reproduction of the demo GIF, where three agents with distinct HEXACO personalities collaborate on a code review and forge a tool once they hit a gap their static toolkit cannot cover. If you want to evaluate this feature before trusting it, that file is the concrete starting point rather than the prose.

## Installing AgentOS and Running a First Session

The package is published to npm as `@framers/agentos`. The README gives a single install command:

```bash
npm install @framers/agentos
```

The package is ESM only. package.json sets `"type": "module"` and the exports map points both `import` and `default` at `./dist/index.js`, so a CommonJS `require` is not the path the package advertises.

The quickstart builds an agent with a provider, instructions, a personality, and a memory config, then opens a session keyed by a user id:

```typescript
import { agent } from '@framers/agentos';

const tutor = agent({
  provider: 'anthropic',
  instructions: 'You are a patient CS tutor.',
  personality: { openness: 0.9, conscientiousness: 0.95 },
  memory: { types: ['episodic', 'semantic'], working: { enabled: true } },
});

const session = tutor.session('student-1');
await session.send('Explain recursion with an analogy.');
await session.send('Can you expand on that?');
```

The README states the provider is auto-detected from the environment when `provider` is omitted, so you need the matching API key set before the first `send` resolves. The second `send` is the interesting one: it relies on session history, not on the memory subsystem, which is the change described for version 0.10.

If you want a stateless session, the README shows the opt-out explicitly:

```typescript
const stateless = agent({ model, memory: false, history: false }).session('job-1');
const bounded = agent({ model, history: { maxTokens: 60_000 } }).session('job-2');
```

Note that `memory: false` alone no longer makes a session stateless. The README states that before 0.10, disabling memory kept no history, and that from 0.10 sessions remember by default. Anyone upgrading across that boundary should read the session section before assuming old behaviour.

## Session Bounds, Reseeding, and the Cache Caveat

Sessions in 0.10 carry what the README calls a lossless conversation transcript: assistant tool calls, tool results, thinking blocks. It is independent of the memory subsystem and bounded by default, with whole-block eviction once a rough 120K-token estimate is passed.

Whole-block eviction is the right call for cache stability. Dropping half a tool call would corrupt the transcript. But it means a single large tool result can push you past the bound and evict more than you expected, and the README does not describe a way to pin a block so it survives eviction.

For long tool-driving loops, the README offers three tools: `session.reseed(snapshot)` for atomic history replacement with in-flight epoch guarding, `session.messages()` as checkpoint material, and per-send generation overrides covering `toolChoice`, `requestTimeout`, `cache`, `cacheDiagnostics` and `blockLabel`.

The cache story comes with an honest caveat that is worth quoting rather than paraphrasing: history byte-stability holds for the stored transcript between eviction events, but the wire request can still differ when dynamic memory context or message-mutating hooks inject per-call content. In other words, a stable transcript does not guarantee a stable prompt, and provider-side prompt caching may miss more often than the transcript suggests.

## Where AgentOS Is the Wrong Choice

The project is at 0.10.x, with three releases on 2026-08-07 alone and a package.json version of 0.10.18. Pre-1.0 versioning means the 0.10 session change is the kind of break you should expect again. If you need an API that will not move under you, this is not that.

The framework is TypeScript and ESM only. There is no Python package described in the README, and the exports map does not advertise a CommonJS entry. Teams with Python agent code cannot reuse this.

The cognitive memory model is opinionated in a way that can hurt. Decay, suppression and reconsolidation are good for an assistant that should forget gracefully. They are bad for a compliance agent that must recall every prior statement exactly. The README does not document a mode that disables the neuroscience mechanisms while keeping persistence.

Runtime tool forging is the feature most likely to be misapplied. In a setting where a wrong tool call has real consequences, an LLM judge approving generated code is a thin control. The README describes the judge and the sandbox but does not describe a human approval step, and it does not document revoking a forged tool before the session ends. If your threat model requires that, the forging path is not ready for you.

Finally, the repository is a library, not a platform. There is no hosted control plane in the README. You supply the provider keys, the process, and the persistence substrate.

## How AgentOS Differs From a Plain LLM SDK

The obvious alternative is a provider SDK or a thin wrapper like the Vercel AI SDK: you get streaming, tool definitions and message arrays, and you own everything else. The difference in approach is where state lives. A plain SDK treats a conversation as an array you pass in each call. AgentOS treats the session as an object that owns its own transcript and eviction policy, and layers memory on top of it.

If your agent is a single request with a fixed set of tools, that extra machinery is overhead. The AgentOS session object, the reseed path and the memory config all exist to serve runs that last long enough to need them.

A second comparison is a vector database plus your own retrieval code. That gives you semantic recall, which is roughly the `semantic` memory type in the AgentOS config. What it does not give you is decay, retrieval-induced forgetting or source-confidence tracking, because a vector store has no notion of time or of competing memories suppressing each other. If those mechanisms are the reason you are looking at AgentOS, no vector database substitute will match them. If they are not, a vector store plus a prompt template is less code to own.

The README also lists six orchestration strategies and points at `examples/agency-graph.mjs`, `examples/agency-roundtable.mjs` and `examples/agency-shared-memory.mjs`. Those are the files that show how the multi-agent story differs from running several independent SDK clients, since shared memory is the part a plain SDK cannot express.

## Licence, Maintenance and Upgrade Cost

AgentOS is Apache-2.0. The repository contains a LICENSE file at the top level and the README and package metadata both state the licence. Apache-2.0 includes an express patent grant and requires that you preserve notices and state significant changes when redistributing. For most teams embedding the library in a service, that is a low-friction arrangement, but it is not legal advice and the licence text is the authority.

The last push to the repository was on 2026-08-07, and the most recent release, v0.10.14, carries a timestamp of 2026-08-07T11:00:23Z. The repository is not archived. Three releases landed within hours of each other that day, which suggests a burst of work rather than a steady cadence, and the repository shows no push after that date.

Upgrade cost is the real budget line. The README documents one behaviour change between pre-0.10 and 0.10 that silently alters semantics: `memory: false` used to mean stateless, and now it does not. Code that relied on the old meaning will keep history it did not expect, which affects token spend and possibly what the model can see. The escape hatch is `history: false`, and the README shows it in a code comment. Any upgrade plan should include a check for `memory: false` call sites.

There is a CHANGELOG.md at the repository root, so the release-by-release detail exists in the tree even though it is not reproduced in the README.

## Conclusion

Adopt AgentOS if you are building a TypeScript service where an agent must hold context across a long session and occasionally extend its own toolkit, and you are willing to read the source when the docs run out. Skip it if you need a stable 1.0 API, prefer Python, or want a hosted agent platform rather than a library you assemble. Before committing, verify two things yourself: that your provider key resolves through the dispatch layer as the README describes, and that the session eviction behaviour around the ~120K-token estimate matches your workload, since the README states the wire request can differ from the stored transcript when dynamic memory context is injected.

## FAQ

### How do I install AgentOS?

Install it from npm with the command the README gives, `npm install @framers/agentos`. The package is ESM only, since package.json sets `"type": "module"` and the exports map points at `./dist/index.js`.

### What is AgentOS?

AgentOS is an open-source Apache-2.0 TypeScript framework for AI agents, published as `@framers/agentos`. It combines persistent cognitive memory, runtime tool forging, multi-agent orchestration and a single dispatch interface across 11 LLM providers.

### What is an agent OS in the sense AgentOS uses the term?

In this project the term describes a framework layer rather than an operating system: it owns an agent's session transcript, its long-term memory, its tool catalog and its provider dispatch. The README's framing is agents that remember, adapt and write their own tools.

### What is the alternative to AgentOS if I do not want its memory model?

A plain LLM SDK or a vector database plus your own retrieval code covers the same ground with less machinery, but neither provides decay, retrieval-induced forgetting or source-confidence tracking. The README does not describe an adapter for plugging in an external memory store.

## Sources

- [framerslab/agentos on GitHub](https://github.com/framerslab/agentos)
- [License: Apache-2.0](https://github.com/framerslab/agentos/blob/master/LICENSE)
- [Project website](https://docs.agentos.sh)
- [README](https://github.com/framerslab/agentos/blob/master/README.md)
- [Releases](https://github.com/framerslab/agentos/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/framerslab-agentos
