# Microsoft Agent Lightning v1.0: training agents inside their real harnesses

> Agent Lightning is a Python framework that inserts a proxy between an agent and its model so reinforcement learning can train the agent without rewriting its tools or control flow. The v1.0 refactor cut the core to roughly 3,500 lines and added Kubernetes rollout execution.

**microsoft/agent-lightning** — The absolute trainer to light up AI agents.

- Repository: https://github.com/microsoft/agent-lightning
- Website: https://microsoft.github.io/agent-lightning/
- Stars: 18,538 · Forks: 1,649
- Language: Python
- License: MIT
- Published: 2026-08-21 · Updated: 2026-08-21 · Language: en
- Canonical page: https://hysenlabs.com/projects/microsoft-agent-lightning

## The problem Agent Lightning solves: training an agent you already built

Most reinforcement learning code for language models assumes the model is the whole program. You write a training loop, generate completions, score them, update weights. An agent is not that. It calls tools, keeps context across turns, branches on intermediate results, and often runs inside a harness such as an IDE plugin or a sandboxed code executor. Rewriting that harness into a training script is where agent RL projects stall, because the harness is the part that took the longest to get right.

Agent Lightning takes the position that the harness should stay where it is. The README describes the goal as training with real agent harnesses where agents interact with the model through the Agent Lightning v1.0 proxy with "ZERO changes", while tools, context, control flow and environments stay in the loop. The audience is therefore narrow and specific: teams that already have an agent that works, and now want to improve the policy behind it rather than replace the scaffolding around it. The repository ships examples for Calc-X, GSM8K, ScienceWorld, Search-R1, LLM-in-Sandbox and a coding agent, which is a fair map of the intended territory: multi-turn, tool-using, environment-bound tasks.

## Trainer, API Gateway, Rollout Controller: the three components in v1.0

The v1.0 architecture has three components and the README is explicit about the division of labour. The Trainer runs `verl` and vLLM, builds training samples, and updates the policy. The API Gateway proxies model requests and captures training data. The Rollout Controller runs agents locally or as Kubernetes Jobs.

The data flow runs in a loop. The Trainer creates rollouts. The Controller launches agents to serve those rollouts. The agents talk to their model through the Gateway, which records the interactions. The Gateway turns those interactions into training data, and the Trainer consumes it to update the policy. Nothing in that description requires the agent to know it is being trained, which is the whole point of putting the capture point at the model boundary rather than inside the agent.

Two design choices are worth naming. First, the capture point is the proxy, so anything the agent does that does not pass through a model call is invisible to training. That is a real boundary, not a detail. Second, the Controller is pluggable between local processes and Kubernetes Jobs, and the README says the Kubernetes path runs agents directly as Jobs "without relying on external sandbox services". That matters when your agent executes untrusted code, since the isolation story is then your cluster's rather than a vendor's.

## Installing Agent Lightning and running a first rollout

The README gives a CUDA 13.0 example. It assumes you have cloned the repository and are inside it. `uv sync` installs the base environment from `uv.lock`, and the setup script installs the `verl` GPU stack at a pinned version and CUDA tag.

```bash
cd <this-repo>
uv sync
bash scripts/setup_verl.sh 0.8.0 cu130
```

The package itself declares `requires-python = ">=3.12"` in `pyproject.toml`, and its runtime dependencies are modest: FastAPI, uvicorn, pydantic, httpx, hydra-core, omegaconf, structlog, jinja2, kr8s and pyyaml. The heavy machinery lives in the `verl` stack that the setup script pulls in, not in the `agentlightning` wheel. If you only want to read the code or run the lighter examples, that split is useful to know before you provision a GPU machine.

The project installs two console entry points, defined in `pyproject.toml`: `agl-server` for the API Gateway and `agl-controller` for the rollout controller.

```bash
agl-server
agl-controller
```

For an actual first run, the README points at the Quick Start page and at the example directories under `examples/`. The Calc-X example is described as a proof-of-concept math reasoning setup using AutoGen with MCP calculator tools that requires only one GPU, which makes it the cheapest place to confirm that the Gateway is capturing traffic before you scale to the coding agent pipeline. The README does not document the exact command line for each example in the top-level file; those live in the per-example documentation.

## What the coding agent result does and does not tell you

The headline number in the README is a coding agent trained on 6K samples, where an end-to-end Qwen3.5-9B workflow improves SWE-bench Verified from 41.8% to 56.4%, a gain of 14.6 percentage points. The repository says the full pipeline is released, including data cleaning, reward-hacking prevention and training scripts.

That is a strong result and it is also a single result. It is one model, one benchmark, one sample budget. Nothing in the README reports variance across seeds, and nothing reports what the same pipeline does on a smaller or larger base model. The reward-hacking prevention work is mentioned but not described in the top-level file, which is exactly the part a reader would want to inspect before trusting the number, since a coding agent scored against repository tests is a setting where reward hacking is easy to produce accidentally. Treat the figure as evidence that the approach can work at this scale, and read the released pipeline and the technical report before treating it as a number you can expect to reproduce on your own task.

## Where Agent Lightning is the wrong tool

The framework assumes your agent makes model calls you can route through a proxy. If your agent's behaviour is dominated by deterministic logic, retrieval, or a fixed workflow with a single model call at the end, there is little trajectory to learn from and the Trainer will be building samples out of a thin signal. Likewise, if your agent calls a model through a path you cannot redirect, such as a vendor SDK with hardcoded endpoints or a hosted assistant you do not control, the Gateway never sees the traffic and the whole capture mechanism is inert.

The second constraint is operational. The Trainer runs `verl` and vLLM, and the README's own install example is a CUDA machine. This is not a laptop framework despite the small core. The roughly 3,500 lines describe the framework, not the compute it drives. Teams without GPU capacity, or without the appetite to run a Kubernetes cluster for isolated rollouts, will find the setup cost larger than the training benefit for small experiments.

The third is the v1.0 break. The README states that Agent Lightning was completely refactored in v1.0 and points readers to a separate branch for legacy releases earlier than v1.0. Anything you find written against the 0.x API, including third-party blog posts and community forks, may not map onto the current component names. The community list already includes a fork used by Youtu-Agent, which is a sign that some users pin to a modified branch rather than the mainline.

## Agent Lightning against DSPy and other agent training routes

The closest comparison people search for is DSPy, and the difference is in what gets optimized. DSPy compiles prompts and program structure against a metric: you describe modules and a signature, and the optimizer searches over instructions and few-shot demonstrations. The model weights do not move. Agent Lightning does the opposite. It leaves the prompt and the harness alone and updates the policy with reinforcement learning, using `verl` and vLLM under the Trainer.

That distinction decides most adoption questions. If your bottleneck is prompt quality and you have no GPU budget for training, DSPy-style compilation is cheaper and faster to iterate. If your bottleneck is that the model itself cannot do the task no matter how the prompt is written, and you have a harness whose trajectories you can capture, Agent Lightning is aimed at that problem. There is also the option of doing nothing to the weights and instead relying on a hosted fine-tuning service, but that route typically requires your data in a specific format and does not preserve the multi-turn tool loop the way the Gateway capture does. The README's own community list shows a third pattern: Youtu-Agent built on a modified branch of Agent Lightning and reports verified training up to 128 GPUs, which suggests that at large scale teams are willing to fork rather than wait for upstream.

## Licence, maintenance and the cost of upgrading

Agent Lightning is MIT licensed, stated in the README, in `LICENSE` and in the `pyproject.toml` license field. MIT is permissive: it allows commercial use, modification and redistribution provided the copyright notice and permission notice are retained. That is a description of the licence text, not legal advice, and the usual caveat applies that your own legal review decides how it interacts with your dependencies, since the `verl` and vLLM stack the Trainer drives carries its own licensing.

The repository is not archived, and the last push was on 2026-08-24, the same day as the v1.0.1 release. Version 1.0.0 landed on 2026-08-17, a week earlier, so v1.0.1 is a quick follow-up rather than a long-settled release. The jump from v0.3.0 in December 2025 to v1.0.0 in August 2026 is the upgrade cost in one line: the refactor is a breaking change, and the README tells you to use the v0.x branch if you need the old behaviour. Plan for the possibility that v1.0.1 has rough edges that a later patch release smooths out, and pin your dependency version rather than tracking `main` if you are running training you need to reproduce.

## Conclusion

Adopt Agent Lightning if you already have a working agent harness and want to train its policy without rewriting tools, prompts or control flow, and if you can supply the GPU stack that verl and vLLM need. Do not adopt it if you want a hosted training service, if your agent has no traceable model calls, or if you cannot run Kubernetes or a local process to launch rollouts. Before committing, verify three things in your own environment: that your agent's model traffic can be pointed at the gateway, that your reward function can be computed from the captured trajectories rather than from the agent's own final answer, and that the rollout launcher you intend to use (local or Kubernetes) matches the isolation your tools need.

## FAQ

### What is Microsoft Agent Lightning used for?

It is used to train AI agents with reinforcement learning while they keep running in their existing harness. The Trainer runs verl and vLLM, the API Gateway proxies model requests and captures training data, and the Rollout Controller launches the agents locally or as Kubernetes Jobs.

### How do I use Agent Lightning?

You install the base environment and the verl GPU stack, then run the gateway and controller entry points and point your agent's model calls at the gateway. The README's install example is uv sync followed by bash scripts/setup_verl.sh 0.8.0 cu130 on a CUDA 13.0 machine, and the Quick Start page covers the first local run.

### What is Agent Lightning in Microsoft's project list?

It is an open source Python framework from Microsoft, MIT licensed, described in the README as a lightweight agentic RL framework for training agents with real harnesses. The current line is v1.0, a complete refactor of the earlier 0.x releases.

### How does Agent Lightning differ from DSPy?

DSPy compiles prompts and program structure against a metric without changing model weights, while Agent Lightning updates the policy with reinforcement learning through verl and vLLM. Agent Lightning leaves the prompt and harness alone and captures trajectories at the model proxy instead.

### What are the alternatives to Agent Lightning?

The README does not name a direct alternative, but it lists community projects built on or alongside it, including DeepWerewolf, AgentFlow and Youtu-Agent, the last of which uses a modified branch of Agent Lightning. For prompt-level optimization rather than weight updates, DSPy is the approach people most often compare it against.

## Sources

- [Official documentation](https://microsoft.github.io/agent-lightning/)
- [Official README](https://github.com/microsoft/agent-lightning#readme)
- [Project repository](https://github.com/microsoft/agent-lightning)
- [Release notes](https://github.com/microsoft/agent-lightning/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/microsoft-agent-lightning
