learn-agent-architecture: the method is diffing adjacent sections
Learn AI agents from scratch.
At a glance
- What is it?
- A teaching repository that explains the harness layer of AI agents across 24 numbered sections grouped into eight layers, using four studied systems as worked examples. Its stated learning method is to read a section's source and diff it against the previous section, and one of its 24 sections continues in a separate repository.
- Who is it for?
- This repository suits a developer who can read Python and wants to understand what sits between a model call and an action, using real systems as reference points rather than invented examples. It does not suit a reader who wants a survey of frameworks, or anyone without an Anthropic key, because every demo loads one.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 10 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The learning method is a diff between two source trees
The How to learn section states its own method, and it is unusual. Sections are meant to be read in order, each building on the layer before it, and for a runnable one you read a file and run a demo. The third instruction is the one that matters:
> Diff a section's `src/` against the section before it. The diff is the one mechanism that section adds.
So the repository is constructed so that its own history between two adjacent folders is the lesson. A section is not a topic explained in prose; it is a delta applied to the previous one. That has consequences for a reader. You cannot jump to section 14 and understand it, because the diff is against section 13 and the point is what changed. It also means the source directories have to be kept deliberately minimal and incrementally different, which is a maintenance constraint the project has accepted. And it explains why there is no code at the repository root at all: every line of runnable code belongs to a numbered section, because the numbering is the curriculum.
Every section is written through the same four-part lens
The method has a fixed shape, and the four parts are named. The opening states what problem the layer solves. The mechanism gives the general design and the control flow. The per-system part shows how real systems implement it. The failure modes say what breaks and how to mitigate it. Every section is self-contained and uses all four. That is the difference between this and a chapter in a book: a chapter can assume the previous one, a self-contained section cannot, so the per-system slot has to do the work that a cross-reference would otherwise do. It is also why the failure-modes slot is mandatory rather than optional. A repository teaching harness design has to be honest that permissions and context budgeting are where things go wrong, and making that a required fourth heading means no section can quietly omit it.
Section 9 is split and the rest of it lives in another repository
The sections table runs to 23, and a note above it says that section 9 continues in a companion repository with ten more stages that scale the memory loop to production. So one of the twenty-four numbered sections is not one section. The repository also points at two further companion projects by the same author: one that takes the memory loop into a production memory subsystem, and one that teaches a specific harness from scratch one plugin seam at a time. A third companion turns real cases from agent applications into reproducible, verifiable eval sets. The reasoning behind splitting is defensible, since memory is the layer most likely to grow past what a single writeup should carry, and the naming makes each piece findable. The cost is that the advertised structure of eight layers and twenty-four self-contained sections is true except at one point, and the exception is the memory layer, which is where a reader is most likely to arrive expecting the full treatment.
Eight layers hold twenty-four sections and the numbering is the curriculum
The table describes eight layers running from the basic loop to a harness that runs itself, and the numbering is zero-based, so sections run from 0 to 23. Layer zero is a single section, the harness thesis, which asks where agency comes from and answers with the model versus harness distinction. Layer one is the core loop and holds four sections: the loop itself, the tool runtime, permissions and sandboxing, and hooks. Layer two is complex work and holds four more: planning and todos, subagents, skills, and context management. Layer three is knowledge and resilience and starts with memory at number 9. Each row names the question it answers and the mechanisms it introduces, so the mechanisms column reads as a vocabulary: messages array, stop reason, registry, schemas, dispatch, permission modes, approvals, lifecycle events with the two named hooks, plan mode, fresh message arrays for delegation, a skill file with progressive disclosure, then budgeting, stubs, compaction and summaries. That vocabulary is reused deliberately, which is what makes the diff-based method possible.
Two dependencies, both exact, and one of them is a vendor client
The dependency list is two lines:
anthropic==0.112.0
python-dotenv==1.2.2Both are pinned exactly rather than by range, which is the right call for a repository whose whole purpose is reproducibility, since a demonstration that silently changed behaviour would undermine the argument. One of the two is a single vendor's model client. The sample environment file names the consequence: an API key, a model, and a commented-out base URL described as optional and intended for an Anthropic-compatible proxy. So running the demos requires a key from one provider, and the escape hatch is a compatible proxy rather than a different client. That is a defensible scope decision for a repository teaching a specific harness layer, and it is worth knowing before you clone expecting to point it at a local model.
The release that renamed the repository is in the version history
Three releases are on record and their titles describe what changed rather than what shipped. One adds typed decision edges. One adds Japanese and Korean translations. The third is the interesting one: it records the rename to the current repository name. So the version history doubles as a changelog for the project's own identity, and a reader can see that the current name is recent. That matters for anyone who arrived from an older link, and it means the earlier name is not obviously wrong so much as superseded. The translations release is similarly informative, because the language switcher at the top of the page offers five languages, so the two that release names are the two most recently added and the three older ones predate them.
Each studied system is pinned to a version, including a release candidate
The systems table has a version column and it is the most methodologically careful thing on the page, because every entry is a moving target. A frontier coding agent is studied at one specific release. A long-term assistant is studied at another, at a version number formatted as a year and a month. A research baseline is studied at a third, and that baseline is described as one bash tool in about 150 lines. The fourth is a plugin-first harness studied at a release candidate, so that entry is pinned to a pre-release build. Each system is also mapped to the sections worth reading for it rather than implying full coverage: the frontier agent is read for everything from 0 to 23, the long-term assistant for seven specific sections, the baseline for the smallest complete loop and the evaluation harness, and the plugin-first one for plugin seams, a durable session log and a protocol layer. The table's last row is a placeholder for more systems to come.
The thesis is that a model call cannot do any of this on its own
The opening claim is a single division. The model reasons; the harness turns that reasoning into controlled action by running tools, keeping state across calls, gating side effects and coordinating loops, and a model call cannot do those things by itself. Everything in the repository is downstream of that sentence, including the control flow summary, which describes the loop as calling the model, running requested tools, appending results and calling again, and then says the loop is small and most engineering is around it. That is the framing that makes a harness worth twenty-four sections: the thing being taught is the part that is not the model. The stated payoff is portability, that a coding tool, a chat assistant and an autonomous runner mostly differ in harness choices, which is the claim that justifies studying four unrelated systems against one set of sections rather than writing four separate guides.
Editorial conclusion
This repository suits a developer who can read Python and wants to understand what sits between a model call and an action, using real systems as reference points rather than invented examples. It does not suit a reader who wants a survey of frameworks, or anyone without an Anthropic key, because every demo loads one. Check four things first. Check that the four-part lens suits how you learn, since every section is structured the same way and that structure is the argument. Check which system you want to read, because the table maps each one to specific section numbers rather than expecting you to read all twenty-four. Check that section 9 is enough for you, since its continuation and ten further stages live in a companion repository. And check your API key and model before starting, since the dependency list is two packages and the sample configuration names one model. The licence is MIT, the default branch was pushed on 2026-09-23, and the repository was renamed in September 2026.
Frequently asked questions
What is the typical architecture of a learning agent?
As this repository frames it, a model call alone cannot run tools, keep state across calls, gate side effects or coordinate loops. The harness supplies those four things around a small loop that calls the model, runs the requested tools, appends the results and calls again, and the rest of the architecture is the layers around it: permissions and sandboxing, hooks, planning, subagents, skills and context management.
Can I learn agentic AI from scratch?
This repository is written for that, and its stated method is a diff rather than a reading order alone. Read the sections in sequence, read a section's source file and run its demo, then diff that section's source against the previous section's, since the diff is the single mechanism the section adds. Every section follows the same four-part structure.
What does learn-agent-architecture teach about the agent loop?
That the loop itself is small and shared: call the model, run the requested tools, append the results, call again. The engineering sits around it rather than inside it, covering tool dispatch, permission gating, context management, state persistence and coordinating other loops, which is what the numbered sections after the loop each take one piece of.
How do I run the demos in learn-agent-architecture?
Copy the sample environment file, fill in an API key and optionally a model, which defaults to a named Anthropic model. The dependencies are the Anthropic client and python-dotenv, both pinned exactly, and the demos load the environment automatically. An optional base URL is provided for pointing at an Anthropic-compatible proxy.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/hardness1020-learn-agent-architecture)