# ACE (Agentic Context Engine): a persistent Skillbook for agents that keep repeating mistakes

> ACE is the open source engine behind kayba.ai that gives an agent a learning loop without fine-tuning: the Agent runs, a Reflector reads the trace, and a SkillManager curates a Skillbook of strategies. This is what the repository documents, where it is thin, and what to check before adopting it.

**kayba-ai/agentic-context-engine** — 🧠 Make your agents learn from experience. Now available as a hosted solution at kayba.ai 

- Repository: https://github.com/kayba-ai/agentic-context-engine
- Website: https://www.kayba.ai
- Stars: 2,584 · Forks: 311
- Language: Python
- License: Apache-2.0
- Published: 2026-09-09 · Updated: 2026-09-09 · Language: en
- Canonical page: https://hysenlabs.com/projects/kayba-ai-agentic-context-engine

## The problem ACE targets: an agent that relearns nothing between sessions

The README opens with the claim that AI agents do not learn from experience: they repeat the same mistakes, forget what worked, and ignore what failed. ACE is the project's answer to that, and it is aimed at developers already running an LLM agent in Python who want the agent to improve without a training pipeline. The README is explicit that this involves no fine-tuning, no training data, and no vector database. That last point matters more than it looks. A vector store gives you retrieval over documents; ACE instead maintains a Skillbook, described as a persistent collection of strategies that evolves with every task. The unit of learning is a strategy, not an embedding. The intended user is therefore someone who can already produce a correction signal, either from a human or from a test, and who wants that signal converted into something the agent carries into the next session. The README's seahorse example is the smallest version of this: the agent claims a seahorse emoji exists, a correction is fed in, and on the next attempt the agent answers correctly. The repository also positions ACE as the open source engine behind Kayba, a hosted service, so the same loop exists in two forms and the README points readers to the hosted option rather than hiding it.

## Agent, Reflector, SkillManager: the three roles and the Skillbook they share

The mechanism is a division of labour across three named roles. The Agent executes tasks, enhanced with Skillbook strategies. The Reflector analyzes execution traces to extract what worked and what failed. The SkillManager curates the Skillbook by adding, refining, and removing strategies. The data flow is trace-driven: an execution produces a trace, the Reflector reads it, and the SkillManager writes the result back into the Skillbook that the Agent reads on the next run. What makes this more than a summarization prompt is the Recursive Reflector, which the README calls the key innovation. Instead of summarizing a trace in a single pass, it writes and executes Python code in a sandboxed environment to programmatically search the trace. That is a real architectural choice with a real cost: you are running model-generated code, so the sandbox is load-bearing rather than decorative. The repository layout supports this reading. There are separate top-level test files named test_rr_live.py, test_sm_e2e.py, test_sm_live.py and test_sm_tau_retail.py, which suggests the Recursive Reflector (rr) and SkillManager (sm) are exercised separately from the rest of the suite, including against the Tau retail benchmark. The README's results table cites 15 learned strategies doubling pass^4 on the Tau2 airline benchmark with no reward signals, and a 49 percent token reduction over a 10-run learning curve in browser automation. Those are the project's own numbers from its own README, not independent measurements, and the README does not describe the conditions under which they were produced.

## Installing ace-framework and running the feedback loop once

The README gives uv as the install path. The package name on PyPI is ace-framework, not ace, and pyproject.toml requires Python 3.12 or newer, so an older interpreter will fail at resolution rather than at import.

```bash
uv add ace-framework
```

There are two configuration routes. The interactive one runs a setup wizard that walks through model selection, API keys, and connection validation.

```bash
ace setup
```

The manual route is an environment variable. The README shows OPENAI_API_KEY and notes that ANTHROPIC_API_KEY works, or any of the providers LiteLLM supports; .env.example lists OPENAI_API_KEY, ANTHROPIC_API_KEY, GOOGLE_API_KEY, DEEPSEEK_API_KEY, COHERE_API_KEY, BEDROCK_API_KEY, AZURE_API_KEY with AZURE_API_BASE and AZURE_API_VERSION, plus AWS, Hugging Face, Replicate and Together keys. It also defines DEFAULT_MODEL as gpt-4o-mini, DEFAULT_TEMPERATURE as 0.0, DEFAULT_MAX_TOKENS as 512, and the cost controls TRACK_COSTS and MAX_BUDGET.

```bash
export OPENAI_API_KEY="your-key"
```

The first real use is the seahorse loop from the README. You ask a question, feed a correction, ask again, and inspect what was stored.

```python
from ace import ACELiteLLM

agent = ACELiteLLM(model="gpt-4o-mini")

answer = agent.ask("Is there a seahorse emoji?")

agent.learn_from_feedback("There is no seahorse emoji in Unicode.")

answer = agent.ask("Is there a seahorse emoji?")

print(agent.get_strategies())
```

What you should see is the first answer being whatever the model produces unaided, and the second answer reflecting the correction. The interesting output is get_strategies(), which prints the Skillbook contents. Read that list before trusting the loop: it is the thing that will be injected into future prompts, and the README does not show a sample of what a bad strategy looks like. The import path is ace while the distribution is ace-framework, so a pip install ace would not get you this package.

## Where ACE is the wrong tool, and what the README leaves open

ACE is a poor fit when the failure you are trying to fix is not expressible as a reusable strategy. If your agent fails because a tool endpoint is down, or because a user typed something ambiguous, there is nothing for the Reflector to extract and the Skillbook gains noise. The loop also assumes you can supply a correction. The README's example is a human writing learn_from_feedback("There is no seahorse emoji in Unicode."), and the results table describes the Tau2 result as achieved with no reward signals, but the README does not document an automatic source of corrections beyond traces, so a team without a feedback path is buying a loop with nothing to feed it. Cost is the second constraint. Every learning step is at least one additional model call, and the Recursive Reflector writes and runs code, which is more than one call in practice. The repository ships TRACK_COSTS and MAX_BUDGET in .env.example, which implies the authors expect spend to be a live concern, but the README does not state what happens when the budget is hit. Third, the Skillbook is state. The documentation does not describe rollback, versioning, or what happens when two strategies contradict each other after a model change. The maintenance picture is mixed: the last push was on 2026-08-29, so the repository is not stale, but pyproject.toml still declares Development Status 4 - Beta and the 0.12.0 release note describes a RR/Skillbook v2 rewrite, which is a signal that internals are still moving.

## ACE against a vector-memory or retrieval layer

The obvious alternative for an agent that forgets is a retrieval memory layer: embed past interactions, store them in a vector database, and retrieve the nearest ones into the prompt. The difference is in what gets stored and who decides. A retrieval layer stores raw past content and ranks it by similarity to the current query, so the selection rule is fixed and the model never edits the store. ACE stores curated strategies and puts a SkillManager in charge of adding, refining, and removing them, so the store is rewritten by a model on every learning step. That gives ACE a smaller, more opinionated context, which is where the README's 49 percent token reduction figure would come from if it holds. It also gives ACE a new failure mode: a bad curation step can delete a good strategy, and the README does not describe a review gate before Skillbook writes. Retrieval memory has the opposite profile. It rarely loses information, but it grows without bound and its relevance ranking does not improve with use. If your problem is recall of facts, retrieval is the simpler answer. If your problem is that the agent keeps making the same procedural mistake, a strategy store is the better shape, and ACE's README states plainly that it does not need a vector database to provide one. The repository also carries optional extras named deduplication, which pulls in numpy and sentence-transformers, and langchain, which pulls in langchain-openai, langchain-anthropic, langchain-litellm and langgraph, so the two approaches are not mutually exclusive in practice.

## Licence, packaging and the cost of keeping up

ACE is Apache-2.0, declared in both pyproject.toml and the LICENSE file, and the classifier list includes License :: OSI Approved :: Apache Software License. For most teams that is the permissive case: you can use it in a closed product, and the obligations are around notices and attribution rather than source disclosure. This is a description of what the repository declares, not legal advice, and the Apache-2.0 patent grant is the clause worth having your own counsel read if you are embedding the engine in something you ship. The upgrade cost is the more practical question. The version in pyproject.toml is 0.12.0, and the release note for that version reads "RR/Skillbook v2 rewrite + SM hardening", which tells you the core abstractions were reworked rather than patched. A project on 0.x with a rewrite in the most recent release should expect to pin versions and read CHANGELOG.md before bumping. The dependency list is not trivial either: litellm, pydantic, pydantic-ai-slim with the litellm extra, rank-bm25, tau2 and tenacity are required, and the optional extras add heavier stacks, including browser-use, transformers, logfire, instructor, bedrock via boto3, and the tracing extra which pulls kayba-tracing. The tau2 dependency is notable because it is a benchmark harness, and it is a required rather than optional dependency, which means it lands in your environment even if you never run the benchmarks. There is also a separate TypeScript side of the project: the release list includes @kayba_ai/tracing 0.10.0 and @kayba_ai/openclaw-tracing v0.1.1, so tracing exists outside the Python package.

## Conclusion

Adopt ACE if you run a Python agent that fails the same way across sessions and you can afford an extra reflection call after a correction; the README's own example is a single agent.learn_from_feedback call on top of a LiteLLM-backed model, and the package requires Python 3.12 or newer. Do not adopt it if you need a stable API surface: pyproject.toml still declares Development Status 4 - Beta, and the 0.12.0 release note describes a RR/Skillbook v2 rewrite, so method names and Skillbook internals can move between minor versions. Before wiring it into a production loop, verify three things yourself: that your provider key is picked up by the LiteLLM path, that agent.get_strategies() returns strategies you would actually accept into a prompt, and that the reflection step's token cost fits your budget, since the repository ships TRACK_COSTS and MAX_BUDGET settings but the README does not document what happens when the Skillbook grows without bound.

## FAQ

### What is agentic context engineering?

In ACE it means maintaining the context an agent runs with as a persistent, curated artifact rather than a fixed prompt. The README describes a Skillbook of strategies that the Agent reads, the Reflector updates from traces, and the SkillManager curates by adding, refining and removing entries.

### What is a context engine in the ACE sense?

ACE's engine is the three-role loop plus the Skillbook it maintains: an Agent that executes tasks with strategies, a Reflector that analyzes execution traces, and a SkillManager that writes back to the Skillbook. The README states this runs with no fine-tuning, no training data and no vector database.

### What does agentic context engineering mean for a running agent?

For ACE it means the agent's context changes after a correction instead of staying fixed. The README's example calls agent.learn_from_feedback with a correction, and later calls to agent.ask benefit from the strategy that was extracted and stored.

### What does "agentic" mean in simple terms?

ACE does not define the term, but its README treats an agent as a system that executes tasks and produces execution traces, which the Reflector then analyzes to extract what worked and what failed. The learning loop is built on those traces rather than on model weights.

### Is ChatGPT an agentic AI?

The repository does not address ChatGPT. ACE is a Python package, ace-framework, that wraps models through LiteLLM, and the README's example uses gpt-4o-mini through the ACELiteLLM class rather than any specific chat product.

## Sources

- [kayba-ai/agentic-context-engine on GitHub](https://github.com/kayba-ai/agentic-context-engine)
- [License: Apache-2.0](https://github.com/kayba-ai/agentic-context-engine/blob/main/LICENSE)
- [Project website](https://www.kayba.ai)
- [README](https://github.com/kayba-ai/agentic-context-engine/blob/main/README.md)
- [Releases](https://github.com/kayba-ai/agentic-context-engine/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/kayba-ai-agentic-context-engine
