All concepts
Concept

What is an AI agent?

An AI agent is a program that uses a large language model (LLM) to decide on actions, call tools, observe the results and repeat until a task is finished. This page explains the mechanism, when it is worth using, and how the term appears in open-source projects.

Published September 28, 2026

How an AI agent works

A plain LLM call takes a prompt and returns text. An agent wraps that call in a loop. The loop has four parts: a goal, a model, a set of tools, and a stopping condition. The model reads the goal and the conversation so far, then emits either a final answer or a request to call a tool. The runtime executes the tool, appends the result to the conversation, and calls the model again. This repeats until the model returns a final answer or a step limit is reached.

The tools are what separate an agent from a chatbot. A tool is a function with a name, a description and a schema. The model does not run the function itself; it produces a structured request, and the surrounding program runs it. Typical tools include reading a file, running a shell command, searching the web, or calling an HTTP API. google-gemini/gemini-cli follows this pattern: according to its description it brings Gemini models into the terminal with file, shell and search tools. The agent loop is the model choosing among those tools.

Memory is usually just the conversation history plus whatever the tools return. Some agents add a scratchpad file or a vector store, but the core mechanism stays the same: model output becomes an action, the action produces an observation, and the observation goes back into the next prompt. Because the model only sees what is in the context window, long tasks need summarisation or external notes.

The stopping condition matters as much as the loop. A step budget, a token budget, or a human approval gate prevents an agent from running forever. paperclipai/paperclip, described as a workspace for assigning, tracking and reviewing work performed by multiple AI agents, puts roles, budgets and approval gates underneath a task manager interface. That is the same idea at team scale: the agent proposes, a budget or a reviewer disposes.

When you need an agent, and when you do not

You need an agent when the path to the answer is not known in advance. If a task requires trying a command, reading the error, and trying a different command, a fixed script will not cover the branches. An agent can choose the next step from the observed state. Examples include debugging a failing test, filling a web form whose layout varies, or exploring a dataset before writing a query.

You do not need an agent when the steps are fixed. If the workflow is a known sequence of API calls, a normal program is cheaper, faster and easier to test. An LLM call inside that program can still help with classification or extraction, but the control flow should stay in code. Agents also add little when the task fits in one prompt. A single well-written prompt with the relevant context often beats a loop that spends tokens deciding what to do next.

Cost and latency are the practical dividing line. Every loop iteration is another model call, and every tool result is more context. A task that takes one call as a plain prompt may take five to twenty calls as an agent. If the task runs rarely and the input is messy, that trade is often worth it. If it runs on every request, a deterministic path is usually better.

browser-use/browser-use is a case where the agent approach fits the problem. According to its description and analysis, it is a Python library and CLI that drives Chrome through the Chrome DevTools Protocol so an LLM can click, type and fill forms. Website layouts change, so a fixed selector script breaks; an agent can look at the page and decide what to click. The same flexibility is unnecessary for a site you control, where a plain HTTP client would be simpler.

Common pitfalls and limits

The first limit is context. The model only knows what is in the prompt. If a tool returns a large file or a long web page, the agent may lose earlier instructions or waste tokens on irrelevant text. Truncation and summarisation help, but both can drop the detail the task needed. This is why agent quality often depends more on tool design than on model choice.

The second limit is error compounding. A small mistake in step two changes the state for step three. Over a ten-step task, a per-step success rate that looks good can produce a low end-to-end success rate. Step limits and checkpoints exist for this reason, but they only bound the damage; they do not remove it.

The third limit is evaluation. It is hard to say whether an agent is good, because the same goal can be reached by different paths. A test suite with fixed expected outputs does not fit an open-ended loop. Teams often fall back on human review, which does not scale. dair-ai/Prompt-Engineering-Guide, described as guides, papers, lessons and notebooks for prompt engineering, context engineering, RAG and AI agents, is a reading resource first and a runnable app second; its analysis notes the README is the weakest part of it. That is a reminder that documentation and evaluation are separate problems from building the loop.

The fourth limit is permissions. An agent with shell access can delete files. An agent with browser access can submit forms. Sandboxing, read-only defaults and approval gates are not optional extras once the agent can act. microsoft/ai-agents-for-beginners teaches agent construction through Microsoft Agent Framework and Foundry with Python notebooks per lesson; it is free and MIT licensed, but the code path runs through Azure, so the permission and deployment model is tied to that platform.

How the term shows up in open-source projects

Open-source projects use "AI agent" in several distinct ways, and the label does not mean the same thing across them.

Some projects are runnable agents. google-gemini/gemini-cli is described as an open-source AI agent that brings Gemini into the terminal, with file, shell and search tools. browser-use/browser-use is a Python library and CLI that drives Chrome through the Chrome DevTools Protocol so an LLM can click, type and fill forms; it is MIT licensed, and its last push to main was on 2026-08-16. Panniantong/Agent-Reach is an MIT-licensed Python CLI that routes an agent's web lookups through per-platform backends and reports which ones work with a single doctor command. These are programs you install and run.

Some projects are collections of agent examples. Shubhamsaboo/awesome-llm-apps packages over 100 open-source AI agents, agent skills and RAG apps as ready-to-run templates. Its analysis calls it a copy-and-run repository, not a framework, with real constraints around maintenance and vendor lock-in. The value is in the examples, not in a shared runtime.

Some projects are research harnesses. karpathy/autoresearch gives an AI agent a single-GPU nanochat training setup, a fixed 5-minute budget and one metric, then lets it modify train.py and keep whatever improves val_bpb. The analysis describes it as a research harness, not a training framework. The narrow scope is the point: one metric, one file, one budget.

Some projects are infrastructure around agents. paperclipai/paperclip is a Node.js server and React UI that assigns, tracks and reviews work performed by multiple AI agents, with roles, budgets and approval gates underneath. OpenBB-finance/OpenBB is a Python data-integration layer that exposes one consistent interface across Python, REST, Excel and MCP surfaces; it is AGPLv3, and its README is clearer about connecting it than about running it in production. thedaviddias/Front-End-Checklist has grown from a GitHub markdown list into a rule corpus with a hosted MCP server and installable skills, which is how a checklist becomes something an agent can call.

Finally, the term appears in learning material. microsoft/ai-agents-for-beginners is a Microsoft course repository with 18 lessons and Python notebooks, free and MIT licensed, though the code path runs through Azure. dair-ai/Prompt-Engineering-Guide covers prompt engineering, context engineering, RAG and AI agents as a reading resource first. If you are new to the topic, the distinction to hold onto is between a project that is an agent, a project that contains agents, and a project that teaches or supports agents.

In practice

An AI agent is a loop around an LLM: the model picks a tool, the runtime runs it, and the result goes back into the prompt until a stopping condition is met. Use one when the steps cannot be fixed in advance and the task tolerates extra latency and cost; use a plain program when they can. To go further, read the README of google-gemini/gemini-cli for the terminal agent loop, browser-use/browser-use for browser control, and karpathy/autoresearch for a narrow, metric-driven harness. Check each repository's last push date before depending on it.

Shubhamsaboo/awesome-llm-appsGitHub describes it as 100+ AI Agents, Agent Skills and RAG Apps - Free and Open Source.. The repository metadata lists Python as its primary language. The metadata lists the Apache-2.0 license. This article stays within the project description and details documented in the GitHub repository README.140,206 stars · Pythonbrowser-use/browser-use🌐 Make websites accessible for AI agents. Automate tasks online with ease.116,719 stars · Pythongoogle-gemini/gemini-cliAn open-source AI agent that brings the power of Gemini directly into your terminal.107,166 stars · TypeScriptkarpathy/autoresearchGitHub describes it as AI agents running research on single-GPU nanochat training automatically. The repository metadata lists Python as its primary language. This article stays within the project description and details documented in the GitHub repository README.96,997 stars · Pythonpaperclipai/paperclipPaperclip is a workspace for assigning, tracking, and reviewing work performed by multiple AI agents.89,265 stars · TypeScriptPanniantong/Agent-ReachAgent Reach lets an AI agent search and read public content from services such as GitHub, Reddit, YouTube, Bilibili, and Xiaohongshu through one CLI.85,425 stars · Pythondair-ai/Prompt-Engineering-GuideGitHub describes it as 🐙 Guides, papers, lessons, notebooks and resources for prompt engineering, context engineering, RAG, and AI Agents.. The repository metadata lists MDX as its primary language. The metadata lists the MIT license. This article stays within the project description and details documented in the GitHub repository README.78,721 stars · MDXmicrosoft/ai-agents-for-beginnersGitHub describes it as 18 Lessons to Get Started Building AI Agents. The repository metadata lists Jupyter Notebook as its primary language. The metadata lists the MIT license. This article stays within the project description and details documented in the GitHub repository README.76,092 stars · Jupyter Notebookthedaviddias/Front-End-Checklist🗂 The essential checklist for modern web development, for humans and AI agents74,310 stars · MDXOpenBB-finance/OpenBBOpen data platform that pipes proprietary, licensed, and public financial data into Python, Excel, MCP servers, and REST APIs for analysts, quants, and AI agents.73,636 stars · Pythonmvanhorn/last30days-skillAI agent skill that researches any topic across Reddit, X, YouTube, HN, Polymarket, and the web - then synthesizes a grounded summary.62,806 stars · PythoncrewAIInc/crewAICrewAI coordinates role-based AI agents into crews and event-driven flows, with tools for tasks, memory, tracing, and deployment.59,187 stars · Python

Sources

  1. Shubhamsaboo/awesome-llm-apps repository
  2. browser-use/browser-use repository
  3. google-gemini/gemini-cli repository
  4. karpathy/autoresearch repository
  5. Panniantong/Agent-Reach repository