Model or dataset
codejunkie99/agentic-stack avatar
codejunkie99/agentic-stack

agentic-stack: one portable memory folder that plugs into thirteen coding agents

One brain, many harnesses. Portable .agent/ folder (memory + skills + protocols) that plugs into Claude Code, Cursor, Windsurf, OpenCode, OpenClaw, Hermes, or DIY Python — and keeps its knowledge when you switch.

2,285 stars281 forksPythonApache-2.0

At a glance

What is it?
This project keeps an agent's memory, skills and protocols in a single folder that thirteen different coding harnesses can read, and the two sentences worth reading are a security disclaimer that says outright the supervisor is not an operating-system sandbox, and an upgrade specification that lists exactly which files it refuses to touch.
Who is it for?
agentic-stack is worth trying if you genuinely move between coding assistants and lose your accumulated context every time you do, because the portable unit is a folder in your repository rather than a setting in somebody else's application, and the blast radius of its upgrade command is documented precisely enough to run on a live project.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 4 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The most important sentence in the readme is a refusal

Somewhere in the middle of the section on bounded loops, after describing worktrees and budgets and deny-path gates and checkpoints, there is a sentence that deserves to be read on its own.

The supervisor bounds and audits child processes. It is not an operating-system sandbox. The readme then tells you what to do instead: use the isolation that the harness itself provides, along with its approval prompts, if you need stronger separation.

That is a small paragraph and it is worth more than most of the feature list, for a specific reason. This tool exists to run an automated loop that makes a change, runs a verifier, and repeats. The code in that loop is generated. It is not code the author wrote, reviewed or tested. It is code produced by a model in response to a task description, and the project's own design acknowledges that by including a verifier and a checker as separate stages.

Given that, the boundary between the loop and the machine is the security boundary, and a supervisor that bounds a process is not a boundary in the sense a reader might assume. It can limit attempts, limit runtime, limit output, limit tokens, and refuse to write to certain paths. It cannot stop generated code from reading your files, opening a socket, or writing somewhere you did not think to deny. The readme says this, in the plainest possible terms, in the middle of the section that would otherwise read as a list of safety features.

The second sentence in that paragraph is an operational instruction rather than a warning, and it is easy to miss. Schedulers are told to invoke one bounded run at a time and to check its exit status before starting another. That is a statement about concurrency: the loops are not designed to be run in parallel, and a scheduler that fans them out will produce overlapping writes to the same working tree with nobody arbitrating.

Taken together, the two sentences tell you how to use this responsibly. Keep one run at a time, treat each run as untrusted code execution, and get your actual isolation from the harness rather than from this tool.

The upgrade command is specified by what it refuses to touch

Most installers tell you what they will do. This one spends a paragraph telling you what it will not do, and that paragraph is more useful than most upgrade documentation.

There is a dry run you can perform first, and a real run that applies it:

bash
./install.sh upgrade --dry-run
./install.sh upgrade --yes

The real run copies the latest harness code, memory code, tool code, the generated skill index and any new skill directories. And then the readme enumerates everything it will leave alone: your main agent instruction file, your agent settings file, all four kinds of memory, a set of candidate records, and your existing skill directories.

That is an unusually precise blast radius, and every item on it is something a user would be upset to lose. An instruction file is months of accumulated guidance. The settings file contains local permissions. The memory files, described as personal, semantic, episodic and working, are the entire point of the project. The skill directories are where your own work lives alongside the ones it shipped.

Stating the negatives matters more than stating the positives here, because the failure mode of an upgrade tool is silent data loss. If a refresh overwrites an instruction file, you find out when your agent starts behaving differently, and by then you may not remember what the file said. If it overwrites a memory directory, you have lost the thing you installed the tool for. The readme is protecting against exactly that, and it does so by writing down a promise a future maintainer can be held to.

There is also a repair command for a specific drift problem, which rebuilds the generated skill manifest from the frontmatter of the installed skill files. That implies the manifest is a derived artefact that can fall out of step with what is on disk, which is the classic failure of a generated index in a directory that users can also write to. Having a documented repair command for it is better than having an integrity check that silently repairs it, because at least the user knows the index was wrong.

The migration story for old installs is handled with the same care. There is a read-only audit command, and the readme says to run it before upgrading from an early version, because it synthesises the state file from signals found on disk so the newer backend can see previous installations. And it says plainly that installing on top without that step would orphan them, which is a sentence that saves somebody an afternoon of wondering where their configuration went.

A maker, a verifier and a checker is three roles rather than two

The loop contract is described as a lifecycle with three stages, and the third one is the reason to read it closely.

The first is a maker, which produces a change. The second is a deterministic verifier, which means the check has a fixed answer: tests, a type checker, a build, anything that returns the same result for the same input. The third is an independent checker.

Two stages would be the obvious design, and most agent loops stop there. Generate, then test. The problem with two stages is that a verifier can be satisfied by the wrong thing. A generated change can make the tests pass by weakening the test, by skipping a case, by special-casing the input, or by replacing the thing being tested with a stub. The tests go green, the verifier reports success, and the loop declares victory. Nothing in the two-stage design can distinguish that from a correct fix.

A third, independent stage is the answer. Its job is not to run the verifier again. Its job is to ask whether the verifier's result means what it appears to mean. That is a different question and it needs a different kind of judge, which is presumably why it is called a checker rather than another verifier.

The two higher-level loop tiers described alongside this add the containment the verifier does not provide. Actions run inside owned version control worktrees, so a run has its own checkout and its own history rather than mutating the working tree in place. Every run is bounded on four axes at once: attempts, runtime, output size and tokens. Deny-path gates stop writes to locations you have excluded. Checkpoints are resumable, so a run that hits a budget limit does not start over. And the events written locally are described as privacy-safe, which given that the project's other feature is turning runs into training data is a distinction worth pausing on.

That last point is where the two halves of this project meet. The data layer collects activity, cost estimates and run history, and the flywheel turns approved, redacted runs into evaluation cases and training-ready records. The word doing the work in both descriptions is approved and redacted, and the project states that none of it trains a model and none of it sends telemetry. A local loop that writes its own evaluation data is a genuinely interesting idea, and it is also the feature where you would want to know exactly what gets captured before you run a hundred iterations.

Four correctness fixes and no features

The most recent release is a patch, and it contains four fixes with no new capability. Reading them is more instructive about the design than a feature list would be, because each one is a failure mode somebody hit in production.

The first is about memory correctness. Retrieval and rendering now share a single map of what has superseded what, so a recall operation no longer returns stale guidance next to the replacement that replaced it. This is the most important of the four, and the reason is that it is the one that makes the memory layer untrustworthy rather than merely incomplete. A memory system that returns a retracted lesson is worse than one that returns nothing, because nothing is visibly nothing while a retracted lesson reads as current. The fix is described as making the two code paths agree on one map rather than each maintaining its own view, which is the right shape of fix: one source of truth instead of two derived ones.

The second is a path bug in the upgrade command. Newly added loop skills were landing in a doubled subdirectory, one level deeper than intended, so they were being written somewhere the rest of the system does not look. That is the classic failure of a code generator that composes paths from a base and a name and forgets that one of the inputs already contains the other. It shipped, and the next release fixed it, which tells you the skill installation path was not covered by the dry run's reporting.

The third is character encoding. Claims containing non-ASCII characters now print and persist correctly regardless of the host's locale. The specific file named is the one that records new lessons, which means the memory layer was writing text that was correct in one environment and mojibake in another. For a tool whose whole purpose is accumulating durable knowledge, an encoding bug is a data-loss bug, and it is the kind that only appears on a developer's machine with a different locale from CI.

The fourth is a leaked file handle when checking whether a lesson had already been appended. The check was correct and the resource management was not. This is the least interesting of the four and the most predictable: it is a bug that a linter for unclosed resources should have caught, and its presence suggests the check was added late and tested functionally rather than under a resource tracker.

Four fixes, no features, and three of the four are in the memory or skill layer. That is a project whose risk is concentrated in the part that accumulates state, not in the part that runs code.

Thirteen harnesses, one folder, and a manifest that can drift

The portability claim is thirteen harnesses long, and the list is worth reading as a compatibility statement rather than a boast.

It runs from the two largest coding assistants through a terminal-first editor, a terminal-first coding agent, an open-source one, a couple of command line assistants from model providers, a couple of others, and ends with two escape hatches: a standalone Python loop for people who do not want a harness at all, and a standalone option for a specific environment. The last two are the important ones. A tool that only supports other people's applications is at the mercy of their configuration formats, and offering a do-it-yourself loop means the portable format has a reference implementation that is not a vendor product.

The portable unit is a folder, and that is the right choice. A folder is something you can commit, put in a monorepo, copy between machines, and diff. It is also something a human can read, which matters for anything that is going to influence what your agent does. The folder holds memory in four kinds, skills, protocols, and a generated index.

The cost of a folder convention is that nothing enforces it. Each harness has its own configuration format, its own idea of where a skill lives, and its own update schedule, and this project has to write to all thirteen. The generated index is the place where that shows: it is a manifest built from the frontmatter of installed skill files, which means a user who adds a skill by hand has created a file the index does not know about, and a user who edits a file's frontmatter has created a disagreement. The repair command exists for exactly that, and the readme tells you to run it as a fix rather than as a routine.

The state of the installation is a single file, and there is a directory of schemas alongside it. That pairing suggests the state file has a versioned shape and that the thirteen adapters are not all writing the same fields into it. The read-only audit command that reports a green, yellow or red status per adapter is therefore doing more than checking whether files exist. It is checking whether an installation is in a state the current backend understands, which is the mechanism that catches a drift before it becomes an orphan.

Removal deserves one sentence. The command to remove an adapter shows a confirmation prompt and then deletes, with the readme stating in the same line that there is no quarantine and no undo. Paired with the export and import command, which moves memory between machines over a network bridge, the intended workflow is clear: export first if you care, then remove.

The installer notices whether it is talking to a terminal

Two details in the installation section are worth extracting, because they are about behaviour in environments most readmes ignore.

The first is what happens when the install script is run with no arguments. On a project that has never been set up, it opens a multi-select wizard listing the harnesses, and it detects which ones are already installed on the machine and pre-checks those boxes. On a project that already has a state file, the same bare invocation opens the dashboard instead. So one command does two entirely different things depending on state, and the state is detected rather than asked about.

The second is the fallback. In a shell with no terminal attached, which is what a continuous integration job looks like, the script does not attempt to open an interactive interface at all. It prints the list of available subcommands and exits. That single branch prevents the two worst outcomes of a wizard in CI: a job that hangs forever waiting for input, and a job that fails because a prompt could not be drawn.

There is a versioning detail in the same section that matters more than it appears. The state file is required for the current backend to track an installation, and installations from before a particular early version do not have one. Rather than refuse to run, the audit command synthesises the state file by inspecting what is on disk. The readme is explicit that installing on top without that step would orphan the previous installations, which is an unusually direct warning about a destructive failure.

The distribution story is conventional and complete. There is a package manager tap and formula for the two Unix systems, a native installer script for Windows with the same verbs, and a clone-and-run path for people who would rather not install anything. The adapter names are listed three separate times in the readme, in the quickstart, in the clone instructions and in a comment on the subcommand list, which is redundant but reflects how many people will scan rather than read.

And the runtime footprint is worth one line: the requirements file lists two packages, one client library for each of the two model providers. Everything else the project does, including the dashboard, the memory layer, the schema validation and the wizard, is built on the standard library. For something that installs into your shell and writes into your repository, that is a reasonable thing to have found.

Editorial conclusion

agentic-stack is worth trying if you genuinely move between coding assistants and lose your accumulated context every time you do, because the portable unit is a folder in your repository rather than a setting in somebody else's application, and the blast radius of its upgrade command is documented precisely enough to run on a live project. It is a poor fit if you use one harness and never change it, and a poor fit as a security boundary, because the project says itself that the loop supervisor bounds and audits child processes without being a sandbox, so any code it writes and runs is running with your permissions. Read the upgrade dry run before the first real upgrade, keep your own copy of the memory directory under version control, and read the denial of isolation as a hard boundary rather than a caveat.

Frequently asked questions

What does agentic-stack actually do?

It stores agent memory, skills and protocols in a single portable folder in your repository and wires that folder into whichever coding harness you use, so switching between assistants does not reset how the agent behaves. It also includes a local data layer for monitoring agent activity, cost estimates and scheduled runs, and can turn approved, redacted runs into local training-ready records.

Is agentic-stack a sandbox for the code its agents write?

No, and the project says so explicitly. The supervisor bounds and audits child processes, using limits on attempts, runtime, output and tokens, plus deny-path gates, but it is not an operating-system sandbox. The readme directs you to the harness's own sandbox and approval prompts when you need stronger isolation.

What does the agentic-stack upgrade command modify?

It refreshes the harness code, the memory and tool code, the generated skill index and any newly added skill directories. It explicitly does not rewrite your main instruction file, your agent settings file, any of the four kinds of memory, candidate records, or existing skill directories. A dry run previews the change first.

How does the agentic-stack loop contract work?

It has three roles in sequence: a maker that produces a change, a deterministic verifier such as tests or a build, and an independent checker that judges whether the verifier's result actually means the change is correct. Higher tiers add owned version control worktrees, budgets on attempts, runtime, output and tokens, deny-path gates and resumable checkpoints.

How many coding harnesses does agentic-stack support?

Thirteen, including the major desktop and terminal coding assistants, command line assistants from two model providers, and two standalone options: one for a do-it-yourself Python loop and one for a specific environment. The standalone options matter because they remove the dependency on another vendor's configuration format.

Official sources

  1. codejunkie99/agentic-stack on GitHub
  2. Issues
  3. License: Apache-2.0
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/codejunkie99-agentic-stack.svg)](https://hysenlabs.com/projects/codejunkie99-agentic-stack)