# Deuz SDK: a zero-dependency TypeScript agent runtime with durable execution and long-term memory

> Deuz SDK bundles memory, compaction, checkpoints, MCP tool calling and human approval into one MIT-licensed package. It is aimed at teams shipping long-running agents on Node, Bun, Deno or the edge, and its value depends on whether you want those seams in your own database.

**Deuz-AI/Deuz-SDK** — Zero-dependency TypeScript framework for production AI agents: durable execution, long-term memory, hybrid RAG, MCP tool calling, human-in-the-loop approval, planning and CodeAct sandboxes. One streaming API for Claude, GPT, Gemini, Grok, Mistral and DeepSeek — Node, Bun, Deno, serverless and edge.

- Repository: https://github.com/Deuz-AI/Deuz-SDK
- Website: http://deuz-sdk.tech/
- Stars: 697 · Forks: 3
- Language: TypeScript
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/deuz-ai-deuz-sdk

## The gap Deuz SDK targets: everything around the model call

The README opens with a claim worth taking seriously: calling a model is a solved problem, and what is not solved is everything around it. Deuz SDK is a TypeScript framework that tries to answer the surrounding questions in one package. Remembering a user across sessions. Staying inside a context window on turn forty. Asking a human before an irreversible action. Resuming after the process dies mid-run. Connecting a tool server without hand-rolling OAuth.

The intended audience is engineers building production agents in TypeScript, on Node 22 or newer, Bun, Deno, serverless or the edge. The package is published as @deuz-sdk/core, with an optional @deuz-sdk/react for useChat, useObject and headless UI. The repository is a monorepo with packages/ and examples/ workspaces, licensed MIT, and the last push was on 2026-08-13.

The design constraint that shapes everything else is stated plainly: zero runtime dependencies, and nothing ambient. Clock, randomness, fetch, keys and logging are injected rather than imported. That is what lets the same code run across runtimes and keeps tests deterministic. It is also a tax on you as the integrator, because you supply those seams.

## Canonical delta stream: the one design rule behind retry, resume and budgets

The README names a single rule that explains most of the codebase: normalize provider bytes to a canonical delta stream first. Every provider is parsed into one internal event language before anything else touches it. Retry, failover, resume, budgets, sub-agents and typed UI events then operate on that shared stream instead of each re-implementing provider quirks.

That choice shows up in the public API. streamChat returns synchronously and never throws, and failures arrive as typed stream parts rather than exceptions. You consume res.textStream as an async iterable and await res.usage separately, so token accounting is a first-class value rather than a log line.

The README claims 28 chat providers across four wires, plus embeddings, images, speech, transcription and video. Four wires for 28 providers is the interesting number: it implies providers are grouped by protocol shape, and adding one means mapping it onto an existing wire. The README does not enumerate which wire each provider uses, so if you depend on an unusual provider, verify its subpath exists in the export table before you commit to the framework.

## Installing @deuz-sdk/core and streaming a first response

The README gives two install commands. The runtime is the only required package; the React bindings are optional. Node 22 or newer is required, or any edge runtime with fetch. Optional peer dependencies are pulled in only when you use the corresponding feature: zod (or any Standard Schema library), @modelcontextprotocol/sdk, react, pg or redis, unpdf, mammoth, xlsx, playwright, and @opentelemetry/api.

```bash
npm install @deuz-sdk/core     # the runtime
npm install @deuz-sdk/react    # optional: useChat, useObject, headless UI
```

After installing, create a provider with createAnthropic and pass an API key. The README's first example streams a terse greeting and then reads usage. Note that streamChat returns synchronously: the call itself does not throw, and failures surface as typed parts on the stream. The loop below writes each text chunk to stdout as it arrives.

```ts
import { streamChat } from '@deuz-sdk/core';
import { createAnthropic } from '@deuz-sdk/core/anthropic';

const anthropic = createAnthropic({ apiKey: process.env.ANTHROPIC_API_KEY });

const res = streamChat({
  model: anthropic('claude-opus-4-8'),
  instructions: 'You are terse.',
  prompt: 'Hello!',
});

for await (const chunk of res.textStream) process.stdout.write(chunk);
const usage = await res.usage;
```

The repository also ships a skills directory and a command to teach a coding agent the API surface. The README states the skills are gated: every @deuz-sdk symbol is resolved against the real export table on every commit, every code example is compiled against the built package, and a freshness check fails when the version or the locked API contract moves.

```bash
npx skills add Deuz-AI/Deuz-SDK
```

If you prefer to read runnable code, examples/ contains numbered projects: 01-basic-stream, 02-tool-loop, 03-next-chat, 04-structured-output, 05-durable-resume and 06-autonomous-agent. The root package.json exposes npm run example to run a workspace example after npm run example:setup builds the packages.

## Memory is a reconciliation pipeline, not a message array

The README is explicit that memory is not a message array. It is a pipeline that extracts durable facts from a conversation, reconciles them against what it already knows using add, update or delete rather than blind appends, scores them for importance, expires them, and pulls the relevant ones back on the next call. It runs on a vector store, a Postgres table, or an Obsidian vault.

Configuration is passed inline to generateText. The seams object supplies the store, embedder and llm. The scope object carries the userId. The recall object caps retrieval at topK 6, maxChars 2000 and expandLinks 1. The writePolicy is set to each-turn, meaning extraction runs every turn rather than on session end.

```ts
await generateText({
  model, messages,
  memory: {
    seams: { store, embedder, llm: model },
    scope: { userId },
    recall: { topK: 6, maxChars: 2000, expandLinks: 1 },
    writePolicy: 'each-turn',
  },
});
```

The each-turn write policy is the detail to think about. It keeps memory current, and it means an extra extraction and reconciliation pass on every turn, plus embedding calls. The README does not publish latency or cost figures for that path, so the trade-off is yours to measure. If your agent runs hundreds of cheap turns, this is the setting to revisit first.

## Compaction: what happens when the context window fills

Compaction is the second feature the README singles out. When the window fills, the runtime prunes stale tool output, drops old reasoning, and folds the earliest turns into a single running summary, described as one block that gets updated rather than a stack that grows. Enabling it is a single string.

```ts
await generateText({ model, messages, maxSteps: 30, compaction: 'auto' });
```

The recovery path is the part worth noting. When a provider still rejects a request as too long, the loop force-compacts and retries that step instead of failing the run. That is a real distinction from frameworks where an over-length error ends the turn and leaves you to restart.

What the README does not document is what the running summary preserves. It says old reasoning is dropped and stale tool output is pruned, but it does not specify whether tool results referenced later in the run survive pruning. For an agent whose later steps depend on an early tool result, that is the question to test before trusting compaction on a long run.

## Durable execution without a workflow vendor, and where it stops

The durability story is step checkpoints written to your own database, with resumeFromCheckpoint called later. The README frames this as avoiding a workflow vendor, and the practical consequence is that the checkpoint table lives in infrastructure you already operate. Store packs exist for SQLite, Redis and Postgres behind the memory, chat, session and run seams. The README example wires chat and session stores from createPostgresStores in the same call, with a runId carrying transcript and checkpoints over one connection.

```ts
import { generateText, handoff } from '@deuz-sdk/core';
import { promptInjectionGuardrail, maxOutputLength } from '@deuz-sdk/core/guardrails';
import { createPostgresStores } from '@deuz-sdk/core/stores/postgres';

const stores = createPostgresStores({ connectionString: process.env.DATABASE_URL });

await generateText({
  model: triage,
  messages,
  maxSteps: 8,
  tools: { ...handoff({ billing, support }), search },
  guardrails: { onInput: promptInjectionGuardrail(), onOutput: maxOutputLength(4000) },
  mcp: [{ url: 'https://mcp.example.com/mcp' }],
  chat: { store: stores.chats, chatId, scope: { userId } },
  session: { store: stores.sessions, runId },
  runtimeContext: { tenantId, db },
});
```

The boundary is the word step. Checkpoints are taken between steps, so work inside a step is not covered. If a tool call is long-running and the process dies mid-call, the step restarts from its checkpoint, and any side effect the tool already performed is your problem to make idempotent. The README does not document rollback of external side effects, and no framework can roll back an email that was already sent. That is the honest limit of this design, and it is worth stating before you put an irreversible action behind a tool.

The same example shows the human-in-the-loop surface. needsApproval is available at any depth, approvals are carried by HMAC-signed expiring tokens, and a missing verdict denies. Deny-by-default is the right default for irreversible actions, and it means an approval that is never delivered blocks the run rather than letting it proceed.

## MCP connections and guardrails in the same call

MCP support is configured with mcp: [{ url }] and, per the README, the connection is namespaced and closed for you. OAuth 2.0, reconnect, sampling and roots are listed as supported. That is a meaningful amount of protocol surface to not hand-roll, and it is the feature most likely to save a week of work if you were about to write your own MCP client.

Guardrails run at three points: on input, on each tool call, and on the final answer, with three outcomes: pass, block or rewrite. The README ships promptInjectionGuardrail() and maxOutputLength(4000) as built-ins. The rewrite outcome is the one to design around, because a rewritten answer is still an answer, and your downstream code needs to know which of the three happened.

The remaining surface is broad: handoff() moves the conversation, including history, tools and model, to another agent; planTasks, CodeAct sandboxes and verifyStep cover plan, act and verify; createAgent produces a frozen value rather than a class; tracing is available without an account through versioned events, a JSONL observer, a standalone HTML run report and an OpenTelemetry bridge. Breadth at this scale is a trade-off. It gives you one dependency instead of eight, and it means the API surface is large enough that the skills directory, which routes tasks to thirteen reference files across 53 subpaths, is not optional reading.

## Deuz SDK versus the Vercel AI SDK

The README links a migration guide at docs/content/docs/migration/from-vercel-ai-sdk.mdx, which tells you who the comparison is against. The Vercel AI SDK is the default choice for TypeScript teams, and it is excellent at the part both projects agree is solved: calling models and streaming results to a UI. Its provider ecosystem is broad and its React integration is mature.

The difference is where the framework stops. The Vercel AI SDK gives you generateText and streamText and leaves memory, compaction, checkpoints and approval to you or to adjacent libraries. Deuz SDK ships those as part of the runtime, with the seams pointed at your database. If you have already built a memory layer and a checkpoint table on top of the Vercel AI SDK, Deuz SDK replaces working code rather than filling a gap, and the migration guide exists precisely because that port is real work.

A second difference is determinism. Deuz SDK injects clock, randomness, fetch, keys and logging, which is what makes the same code runnable on Node, Bun, Deno and the edge and makes tests deterministic. That is a stronger position than most SDKs take, and it costs you the convenience of ambient globals.

## Conclusion

Adopt Deuz SDK if you are building a long-running TypeScript agent and you want memory, compaction, checkpoints and MCP connections in one package rather than assembled from several. Skip it if your agent is a single model call behind a form, or if your team is standardized on the Vercel AI SDK and unwilling to port. Before committing, read docs/content/docs/migration/from-vercel-ai-sdk.mdx, check the @deuz-sdk/core export table against the subpaths you need, and confirm that the store pack for your database exists in packages/.

## FAQ

### What is Deuz SDK?

It is a zero-dependency TypeScript framework for production AI agents, published as @deuz-sdk/core with an optional @deuz-sdk/react package. It bundles durable execution, long-term memory, hybrid RAG, MCP tool calling, human-in-the-loop approval, planning and CodeAct sandboxes behind one streaming API across 28 chat providers.

### How do I install Deuz SDK?

Run npm install @deuz-sdk/core for the runtime, and optionally npm install @deuz-sdk/react for useChat, useObject and headless UI. Node 22 or newer is required, or any edge runtime with fetch, and optional peer dependencies are only needed when you use the features that require them.

### Does Deuz SDK support durable execution and resuming a crashed run?

Yes. Step checkpoints are written to your own database and resumeFromCheckpoint resumes the run later, with store packs for SQLite, Redis and Postgres behind the memory, chat, session and run seams. Checkpoints are taken between steps, so work inside a step is not covered and side effects already performed by a tool are not rolled back.

### What happens when the context window fills up in Deuz SDK?

With compaction set to auto, the runtime prunes stale tool output, drops old reasoning, and folds the earliest turns into a single running summary that is updated rather than stacked. If a provider still rejects a request as too long, the loop force-compacts and retries that step instead of failing the run.

## Sources

- [Deuz-AI/Deuz-SDK on GitHub](https://github.com/Deuz-AI/Deuz-SDK)
- [License: MIT](https://github.com/Deuz-AI/Deuz-SDK/blob/main/LICENSE)
- [Project website](http://deuz-sdk.tech/)
- [README](https://github.com/Deuz-AI/Deuz-SDK/blob/main/README.md)
- [Releases](https://github.com/Deuz-AI/Deuz-SDK/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/deuz-ai-deuz-sdk
