Model or dataset
baby-llm/baby-agent avatar
baby-llm/baby-agent

baby-agent: A Twelve-Chapter Go Tutorial for Building AI Agents from First Principles

AI agent tutorials for backend developers without AI background. 适合后端工程师的零基础 AI Agent 教程

484 stars74 forksGoApache-2.0

At a glance

What is it?
baby-agent is a Go tutorial repository aimed at backend engineers who have no prior LLM background. It walks through agent fundamentals in twelve independently runnable chapters, from raw HTTP calls to the chat completions API through function calling, MCP protocol integration, context management, memory systems, Agentic RAG, safety guardrails, and web service deployment.
Who is it for?
baby-agent is the right choice for backend engineers with Go experience who want to understand agent internals by writing them rather than by configuring a framework. It is explicitly not a production framework, and the README cautions against using it as a starting point for production systems without adding engineering hardening and security boundaries.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 114 days ago.
What is it written in?
Mainly Go, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What baby-agent Is and Who It Is For

Most AI agent frameworks are designed to be used, not read. baby-agent is designed to be read and then run. Each chapter introduces a core agent concept, shows the simplest possible Go implementation of it, and explains why the design is what it is, including what trade-offs were made and what alternative approaches exist.

The target reader is a backend engineer who writes Go and wants to understand how LLM agents work at the protocol level: how function calling actually works in the chat completions protocol, how an agent loop is structured, what MCP is and why it exists, and how context windows are managed under pressure. The README states the project intentionally skips complex mathematical derivations and focuses on engineering implementation.

The README is explicit about what the project is not: it is not an open-source production framework and is not recommended for direct production use. It is positioned as a reference for understanding principles and validating prototypes. A team that wants to ship an agent in production should use baby-agent to understand what is happening inside a framework like LangChain or LlamaIndex, not to replace one.

The project is licensed under Apache-2.0 and uses Go 1.24 or later.

Chapter Structure and Learning Path

The twelve chapters are independent: each chapter lives in its own directory (ch01/ through ch12/) with its own main.go, and each can be run without completing the previous chapters. The README describes the progression as moving from basic LLM calls to progressively more complex agent capabilities.

Chapter 1 covers the chat completions API at the HTTP level, comparing raw HTTP calls to the OpenAI Go SDK. Chapter 2 adds function calling and the agent loop. Chapter 3 adds TUI visualisation using Bubble Tea so the agent's reasoning and tool calls are visible in the terminal. Chapter 4 integrates MCP (Model Context Protocol), adding external tool servers to the agent's toolset. Chapter 5 covers context engineering: truncation, offloading, and summarisation strategies to manage the context window. Chapter 6 implements a two-layer memory system with global (cross-session) and workspace (project-level) memory. Chapter 7 builds Agentic RAG with pgvector-backed semantic search and a reranking step. Chapter 8 adds Docker sandbox isolation and human-in-the-loop confirmation for dangerous commands. Chapter 9 implements a skill plugin system. Chapter 10 converts the CLI agent into an HTTP service with Server-Sent Events streaming.

Chapters 11 (observability with OpenTelemetry) and 12 (LLM evaluation and automated testing) are marked as work in progress.

Setting Up and Running the First Chapters

The setup requires Go 1.24+, an API key from a compatible LLM provider, and a configuration file. Copy the example environment file and fill in your credentials:

bash
cp .env.example .env

The .env.example shows three fields: OPENAI_BASE_URL, OPENAI_API_KEY, and OPENAI_MODEL. The base URL can point to OpenAI.com, DeepSeek, GLM, or any provider that exposes an OpenAI-compatible API.

Each chapter is independently runnable. Chapter 1 streams output from a single query:

bash
go run ./ch01/main --stream -q "用 Go 语言写一个 Hello World"

Chapter 2 demonstrates tool calling by reading a file:

bash
go run ./ch02/main -q "请读取 README.md 并总结项目目标"

Chapter 3 opens the TUI visualisation:

bash
go run ./ch03/main

Chapter 10 starts the web service on port 8080:

bash
go run ./ch10/main

The go.mod file uses module name babyagent and lists dependencies including the OpenAI Go SDK v3, Bubble Tea v2, Gin, tiktoken-go, GORM with a PostgreSQL driver, and the official MCP Go SDK from modelcontextprotocol/go-sdk.

Context Engineering and Memory System Design

Chapters 5 and 6 cover two of the most practically important concepts for production-quality agents: managing the context window and providing persistent memory.

Chapter 5 implements three context management strategies as a composable interface: truncation (dropping old messages safely while keeping the most recent exchanges), offloading (storing long messages to external storage with a recovery pointer), and summarisation (using the LLM to compress older turns into a single entry). The chapter also integrates tiktoken-go for real-time token counting so the strategy fires at the correct threshold. The README notes the strategies are composable and designed to stack together.

Chapter 6 builds a two-layer memory architecture. Global memory persists facts across sessions and is injected into the system prompt. Workspace memory tracks project-specific context. Both are backed by file-system persistence. The LLM itself drives memory updates: the agent calls a memory update function that asks the LLM to extract and store relevant facts from the conversation, rather than storing the raw conversation history.

This design reflects a real engineering constraint: the system prompt has a fixed cost at every turn. Injecting full conversation history into the system prompt is expensive. Distilling relevant facts into a compact memory that the LLM can extract information from on demand reduces that cost while preserving continuity.

MCP Integration and the Tool Namespace Strategy

Chapter 4 integrates MCP (Model Context Protocol), which addresses the problem of connecting many tools to many agents without requiring each combination to be implemented separately. The README describes MCP's three roles: Client, Server, and Tool, and explains the JSON-RPC 2.0 protocol used for tool discovery and invocation.

The go.mod file lists github.com/modelcontextprotocol/go-sdk v1.5.0 as a dependency, which is the official Go SDK for the MCP protocol. Local tools (reading files, executing shell commands) and MCP tools coexist in the same agent session. The chapter implements a namespace strategy to prevent name collisions when a local tool and an MCP server tool share the same name.

The mcp-server.json file at the repository root configures which MCP servers the agent connects to. This is the same configuration format used by Claude Desktop and similar MCP-compatible clients, which means the MCP servers built in this tutorial can be used with external clients that speak the protocol.

Agentic RAG, Safety Guardrails, and Web Deployment

Chapter 7 implements Agentic RAG, where the agent decides autonomously when to retrieve, what to retrieve, and how many results to fetch, rather than always retrieving at every turn. The implementation uses pgvector for vector storage (via GORM's PostgreSQL driver), an embedding model for vectorising code snippets, and a reranking step using a Rerank model. Two chunking strategies are provided: LineChunker and ParagraphChunker.

Chapter 8 adds safety through Docker sandbox isolation for shell command execution. The implementation detects whether Docker is available and falls back gracefully if it is not. Before the agent executes a command classified as dangerous, the TUI shows a confirmation prompt with three options: Allow (once), Reject, or Always Allow for this session.

Chapter 10 converts the CLI agent into an HTTP service. The architecture separates the agent's event output from the transport layer: the agent emits StreamEvent objects through a Go channel, and the HTTP handler converts those events into Server-Sent Events for browser clients. Conversation and message history are stored in SQLite using the libtnb/sqlite dependency. Each message records a parent_message_id so the service can reconstruct the conversation tree by following ancestor references.

Limitations, Missing Chapters, and Maintenance

Two of the twelve chapters are incomplete. Chapter 11 (observability with OpenTelemetry, Jaeger, Prometheus, and Grafana) and Chapter 12 (LLM evaluation, mock LLM interfaces, and automated testing pipelines) are marked as work in progress. Developers who need observability or evaluation infrastructure will need to implement those capabilities outside of what baby-agent provides.

The project is explicitly not a production framework. The README states it is designed for understanding principles and validating prototypes. Specific areas that require engineering hardening before production use include authentication in the Chapter 10 web service, rate limiting, input validation on all LLM inputs, and secure handling of environment variables in container deployments.

The last push was on 2026-06-08. The project does not yet have GitHub releases. The go.mod requires Go 1.25, which at the time of the last push was a pre-release version, so upgrading to a stable Go toolchain version may require updating the module file.

The README does not document how to run pgvector locally for Chapter 7. Teams attempting that chapter will need to provision PostgreSQL with pgvector separately.

Editorial conclusion

baby-agent is the right choice for backend engineers with Go experience who want to understand agent internals by writing them rather than by configuring a framework. It is explicitly not a production framework, and the README cautions against using it as a starting point for production systems without adding engineering hardening and security boundaries. The last push was on 2026-06-08. Chapters 11 and 12 (observability and LLM evaluation) are still planned and not implemented, so teams that need those capabilities will need to look elsewhere for now.

Frequently asked questions

What LLM providers does baby-agent support?

Any provider that exposes an OpenAI-compatible API. The .env.example shows OPENAI_BASE_URL, OPENAI_API_KEY, and OPENAI_MODEL fields; the base URL can point to OpenAI.com, DeepSeek, GLM, or similar providers.

Can each chapter of baby-agent be run independently?

Yes. Each chapter directory contains its own main.go and can be run with `go run ./chNN/main` without completing prior chapters. The README describes the chapters as individually runnable.

Is baby-agent suitable for production use?

The README explicitly states it is not a production framework and not recommended for direct production use. It is designed as a reference for understanding agent principles and validating prototypes, with the expectation that production deployments use it as a learning foundation rather than a base library.

Official sources

  1. baby-llm/baby-agent on GitHub
  2. Issues
  3. License: Apache-2.0
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/baby-llm-baby-agent.svg)](https://hysenlabs.com/projects/baby-llm-baby-agent)