deepagent-code writes every prompt to a log before running it
DeepAgent Code: AI coding agent with persistent memory and control plane
At a glance
- What is it?
- DeepAgent Code is a durable coding agent whose runtime follows one rule: record first, then act. Prompts, tool calls and events go to an append-only log in the same transaction as the state they describe, so a crash resumes where it stopped instead of starting a second copy of the work.
- Who is it for?
- Fit for work that outlives one prompt, where a migration has an objective completion criterion and a long run has to survive a restart, since the durable log and the per-commit worktree review are the features other agents approximate. A poor fit for a single small edit, since the same machinery is overhead there, and a poor fit if you cannot accept a strict plan gate that rejects writes outside the approved plan with an explicit error.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 2 days ago.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Record first, then act
Version 2.0 rebuilds the runtime around a single principle: record first, then act. Every prompt, tool call and event is written to the durable log before it happens.
That one decision produces most of the behavioural differences the project claims, and it is worth tracing how rather than taking the claims at face value.
Durable sessions come from it directly. A prompt is saved as a durable record before execution starts, so if the process dies mid-run the session resumes exactly where it stopped, and a retry picks up the same work instead of silently starting a second copy. The failure mode being designed against is duplication rather than loss.
Event-sourced history is the same idea applied to the log itself. Session activity is an append-only event log written in the same transaction as the state it describes, which is what makes replay, audit and recovery agree on what actually happened. Three different consumers of the history, one source of truth.
The comparison table puts the other side of it plainly: a typical coding agent forgets everything when the process exits. Durability here is a property of the storage layer, not of the model.
Writes need an approved plan, and the default gate is strict
Write isolation is the feature with the most operational consequence, and it is on by default.
A strict plan gate rejects writes outside the approved plan with an explicit error rather than letting them land silently in your tree. The alternative behaviour, an agent writing a file nobody approved, is the one that produces the worst afternoon: the change is real, it is in the tree, and nothing in the transcript said it should not be.
Durable multi-agent collaboration extends the same idea across agents. Delegated work is itself a durable record, every write-capable agent works in an isolated git worktree, and each commits only its own scoped changes. Review happens against the exact commit before anything merges, which is stated as the property that matters: an out-of-date worker can never overwrite newer work.
That is per-commit review rather than per-branch review, and it is the difference between a stale worker failing loudly at review time and a stale worker quietly reverting a colleague's file.
Autonomy and permission are separate settings
The collaboration model and the access model are independent, which is the design choice worth internalising before you configure anything.
There are three collaboration modes. Auto takes a request and defines the objective, designs and plans as needed, then executes end to end. Loop takes a goal and writes an editable `goal+plan.md`, advancing it through plan, execute, verify and iterate cycles. Design takes your own `goal+plan.md` and executes it faithfully without redefining the objective or completion criteria.
Design mode is the interesting one. It is the mode for a handed-over migration, because the person who wrote the plan keeps ownership of the completion criteria and the agent is explicitly forbidden from rewriting them.
Then there are three permission levels, Read-only, Request approval, and Full access, and the documentation is explicit that you can change any of them without changing the collaboration mode. So a read-only Auto run and a Full access Design run are both valid combinations, which is what lets you start with the most restrictive setting and widen it only for the task that needs it.
Steering lands at the next pause, never mid-tool
The control surface is built around not interrupting work, and there are five distinct mechanisms rather than one.
Live steering sends new guidance while a model turn or a tool is running. Your message is saved to the log first, then applied at the next safe pause between model calls, and work in progress is never aborted. The ordering matters: the instruction is durable before it is applied, so a crash between the two loses nothing.
Goal steering is the variant for long runs. Guidance sent to an active goal is folded into the next cycle, which preserves the current tool and plan state rather than resetting the cycle.
Hot plan editing lets you edit a running or paused goal, with stable step IDs, evidence, completed work and the new plan version carrying into the next cycle. The stable step IDs are what make a live plan change reviewable: you can tell which steps a revision actually changed.
Explicit queueing covers the case where an instruction should start after the current one rather than modify it. And every long-running workflow has a human control path: pause, resume, take over, or roll back, each with a durable audit trail.
Memory is five scopes with a governed lifecycle
The stated principle is that memory is not hidden in an opaque prompt. Project state lives in typed, versioned documents with provenance, confidence, scope, status and links, and there are five scopes.
Session-private working context stays with the current conversation. Project-shared facts and decisions follow the repository. User-global preferences can travel across projects. Built-in skills and domain packs remain versioned system knowledge. And sealed evaluator material stays audit-only and never enters model context.
That last one is the item with a security reading: some material is deliberately excluded from what the model sees, which is only possible if memory is a system with scopes rather than one undifferentiated blob.
Learning follows a lifecycle rather than a write. Evidence creates a candidate, isolated review or a human decision changes its status, and regression and ablation gates publish a reproducible knowledge snapshot. Rejection reasons remain durable so discarded patterns are not silently relearned, which is the part that stops a bad pattern from returning as if it were new.
The Repo and Wiki view is how you audit any of it: browse knowledge and execution archives, search across the repository, follow docs-to-code links, inspect lineage, and promote run evidence into governed knowledge.
Four graphs feed one budgeted context window
Context is assembled from four views of the project rather than from a longer prompt.
The code graph covers files, symbols, imports, calls, diagnostics and references. The knowledge graph holds strategies, methodologies, facts, skills and failure dossiers. Project memory holds decisions, constraints, environment facts and learned conventions. The document graph holds plans, designs, worklogs, evaluations and run context.
The mechanism that keeps it inside a budget is described precisely. Stable instructions stay unchanged from turn to turn while changing state is appended at the end. Sent history is therefore byte-stable and shrinks only in one batched pass when the budget is crossed.
That last property is doing double duty. It keeps the prompt small, and because the prefix does not change between turns, provider-side prompt caching stays effective. The measured prefix-cache hit rates are stated as above 96 percent, and the same byte-stability is the reason the token comparison against the baseline comes out ahead.
The numbers are benchmark numbers, and here is the baseline
The performance claims are specific, and it is worth reading them with their comparison attached.
They come from DeepSWE tasks, measured against the mini-swe-agent baseline with the same model. Output tokens drop to roughly a third of the baseline, and the stated reason is that structured tool calls replace long stretches of shell-output reasoning, with generated tokens being the most expensive part of the bill.
On the harder tasks the average fix-to-pass rate rises from 71.8 percent to 98.4 percent. Steps fall by about 10 percent on average while end-to-end wall time is on par with the baseline, which is the honest framing: the saving is in tokens and steps, not in elapsed time.
Total input runs at roughly 40 to 80 percent of the baseline on DeepSWE. Old tool outputs are trimmed by fixed rules and past reasoning is never replayed, which keeps the prompt prefix byte-stable, and the measured prefix-cache hit rates are above 96 percent.
Two limits are worth stating. These are numbers for one benchmark against one baseline on the same model, and a task suite that is built around software engineering fixes will not tell you what happens on a migration with no pass or fail criterion.
A Bun monorepo on the dev branch, with an SST deployment
The repository layout is worth reading before anything else, because it explains the scale.
The package is private, marked as an ES module, and pinned to Bun 1.4.2 as its package manager. The workspaces list is where the product surface shows: `packages/*`, then three sub-workspaces for a console, a stats application and a Slack integration, plus a JavaScript SDK. The root scripts name what exists inside them, with separate dev entries for the core agent, the desktop app, the web app, the console and the stats app, a storybook, an OpenTUI upgrade script and a live LLM test runner.
Two repository-level files are unusual in a good way. One script is named for asserting the documentation claims, which is a machine check that what the README says matches the code. Another generates the live LLM runs across providers.
The infrastructure is a mix of three stacks. There is an SST config and a generated environment type declaration, a Nix flake for the development shell, and a Husky setup with a prepare script for commit hooks. Linting is oxlint, formatting is Prettier with an ignore file, and the task runner is turbo.
One structural fact to check before you build: the default branch is dev, not main.
Editorial conclusion
Fit for work that outlives one prompt, where a migration has an objective completion criterion and a long run has to survive a restart, since the durable log and the per-commit worktree review are the features other agents approximate. A poor fit for a single small edit, since the same machinery is overhead there, and a poor fit if you cannot accept a strict plan gate that rejects writes outside the approved plan with an explicit error. Before trusting the token figures, note that they are measured on DeepSWE tasks against a mini-swe-agent baseline with the same model, and read the branch situation first, because the default branch is dev rather than main.
Frequently asked questions
What is DeepAgent Code?
It is an AI coding workspace for work that lasts longer than a single prompt. Version 2.0 is built around a durable core: every prompt, tool call and event is recorded before it runs, session activity is an append-only log written in the same transaction as the state it describes, and delegated work is reviewed per commit in isolated git worktrees before anything merges.
What are the key differences between DeepAgent Code and a typical coding agent?
The documentation contrasts five things. A typical agent forgets everything when the process exits, while this one records before it acts and resumes where it stopped. It keeps project memory you can inspect rather than only understanding the current prompt. It uses structured tool calls instead of narrating shell output, measured at about a third of the baseline output tokens on DeepSWE. It plans and collaborates through goal loops and isolated worktrees with per-SHA review. And it runs as desktop or terminal against 75 or more model providers with your own keys.
When should I use DeepAgent Code?
When the work has an objective completion criterion and may outlive one prompt: a migration handed over with its own goal and plan file, or a long run that has to survive a restart. The project also documents the case against it, since a strict plan gate rejects writes outside the approved plan by default, so it suits a task where you want that constraint rather than a single small edit.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/deepagent-ltd-deepagent-code)