# microsoft/Trace: end-to-end generative optimization for AI agents

> Trace is a PyTorch-like Python library that turns prompt, code and parameter tuning into a backward pass over an execution trace. It fits teams already comfortable with autodiff who want one optimizer over a whole agent pipeline, and it expects you to write the feedback function yourself.

**microsoft/Trace** — End-to-end Generative Optimization for AI Agents

- Repository: https://github.com/microsoft/Trace
- Website: https://microsoft.github.io/Trace/
- Stars: 760 · Forks: 63
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/microsoft-trace

## What problem microsoft/Trace solves, and for whom

Most agent stacks are tuned by hand. Someone edits a system prompt, reruns an evaluation, reads the failures, edits again. Trace exists to replace that loop with something closer to training. The README describes it as "a new AutoDiff-like tool for training AI systems end-to-end with general feedback (like numerical rewards or losses, natural language text, compiler errors, etc.)". The feedback does not have to be a scalar. That is the part that separates it from classical gradient descent: the signal can be a sentence, a stack trace, or a failed test name.

The intended user is a Python developer who is comfortable with autodiff concepts and wants to optimize more than one artifact at once. In the sales-agent example in the README, a system prompt sits alongside two trainable instruction nodes, and the optimizer sees all of them. A prompt-only tool would touch one string. Trace is aimed at the case where the prompt, the routing logic and a helper function all contribute to the final answer, and where you would rather optimize the graph than each piece in isolation.

## How the trace graph, node and bundle primitives work

Trace has two primitives, `node` and `bundle`. A node wraps a value and marks whether it is trainable. A bundle wraps a Python function so its body becomes part of the graph. The README states that once a node is declared, "all the following operations on the node object will be automatically traced", including arithmetic and list methods, and that plain integers used in an operation are converted to nodes automatically.

The data flow is the interesting part. When you call `backward` on a node, Trace walks the recorded execution trace rather than a symbolic expression tree. The README shows `z.backward("maximize z", visualize=True, print_limit=25)`, which both requests an optimization direction in natural language and renders the graph. Because the trace records what actually ran, a bundled function that branches or calls an LLM is captured the same way a division is. The optimizer then proposes new content for the trainable nodes using the feedback attached to the trace.

That design has a cost. The trace is only as informative as the feedback you attach to it. Trace does not invent a reward; the README's example writes a `get_feedback` function by hand that returns either "test case passed!" or "test case failed!". The library propagates judgement, it does not supply it.

## Installing trace-opt and running a first optimization

The README gives two install paths. The published package is `trace-opt`, and the repository also supports an editable install for development. Python 3.9 or newer is required, and LiteLLM is the default LLM backend from v0.1.3.5 onward.

```bash
pip install trace-opt
```

If you need the older AutoGen backend, the README documents an extra tag rather than a separate package:

```bash
pip install trace-opt[autogen]
```

For a development checkout, clone the repository and install it in place. The README notes that Git Large File Storage may be needed if git cannot clone the repository.

```bash
pip install -e .
```

A first real use is the strange-sort example from the README. You declare a bundled function as trainable, give it a docstring, and let the optimizer rewrite the body. The docstring matters: it is the specification the optimizer reads when proposing a change.

```python
from opto.trace import node, bundle
from opto.optimizers import OptoPrime

@bundle(trainable=True)
def strange_sort_list(lst):
    '''Given list of integers, return list in strange order.'''
    lst = sorted(lst)
    return lst

optimizer = OptoPrime(strange_sort_list.parameters())
```

The training loop then follows PyTorch conventions almost exactly. You compute an output, compare it to a ground truth, and call `zero_feedback`, `backward` and `step` in sequence. The README's loop runs for two epochs and breaks early when `correctness` is true. If you run it, expect the printed output to show the function's return value changing between epochs as the optimizer rewrites the body; the README does not print the intermediate code, so the only visible signal is the output and the feedback string.

## Where Trace gets in the way

The trace graph is built from Python execution, so anything that runs outside the traced path is invisible to the optimizer. If your agent shells out to a subprocess, calls a remote service that returns a string you never wrap, or stores state in a global that the traced function reads later, the optimizer cannot attribute the outcome to a trainable node. The README does not describe a mechanism for attaching feedback to untraced side effects, so the practical rule is that the thing you want optimized has to be inside the graph.

Feedback quality is the second constraint. Natural-language feedback is flexible, but the optimizer's step is only as good as the sentence you write. A vague string such as "this is bad" gives the optimizer little to work with, while "the extracted name is missing the surname" points at a specific edit. Nothing in the repository enforces a feedback format, so this becomes a discipline problem rather than a library feature.

The project also labels itself Beta. The `pyproject.toml` classifiers include "Development Status :: 4 - Beta", and the most recent release tag is v0.1.3.8, published on 2025-03-23. The last push to the repository was on 2026-06-17. Anyone planning to build a long-lived internal platform on top of Trace should read that version history as an indication that interfaces may still move.

## Trace compared with TextGrad and with prompt-only optimizers

TextGrad is the closest reference point, and the relationship is documented rather than implied. The README's update list for 2024.9.14 states that "TextGrad is available as an optimizer in Trace", and the repository ships an `examples/textgrad_examples/` directory. The difference in approach is scope. TextGrad treats textual gradients as the central object and optimizes text variables. Trace treats the execution trace as the central object and optimizes anything the trace touches, which can include code bodies and numeric parameters alongside text. If your problem is purely prompt wording, TextGrad's narrower model is simpler to reason about. If your problem spans a prompt, a routing decision and a helper function, Trace's graph is the one that can express all three in a single backward pass.

The other comparison is with prompt-only optimizers that never see your code. Those tools cannot propose a change to a sorting function or a mapper, which is exactly the case the README highlights when it cites the 2024.10.21 paper applying Trace to mapper code for parallel programming. That paper is external to this repository, but it is listed in the README's update section as an application of the same optimizer.

## Maintenance, releases and what the MIT licence leaves you

Trace is MIT licensed, and the `setup.py` file declares `license='MIT LICENSE'` while the README carries an MIT badge. For most adopters that means you can read, modify and redistribute the library, including in commercial products, provided the licence text travels with it. This is a summary of what the repository states, not legal advice; if you are embedding Trace in a shipped product, have your own counsel read the LICENSE file rather than this paragraph.

Upgrade cost is shaped by the dependency list in `setup.py`: `graphviz`, `scikit-learn`, `xgboost`, `litellm` and `black`. LiteLLM is the one most likely to move under you, since it tracks provider APIs, and it is also the component the README points to as the default backend. The optional AutoGen path is pinned tightly at `autogen-agentchat==0.2.40` in `pyproject.toml`, which suggests the compatibility shim is maintained against one specific version rather than a range. If you rely on that backend, pin it deliberately.

The repository has a Makefile with two targets, `doc` and `doc-deploy`, which build the Jupyter Book documentation and publish it to GitHub Pages. That is the documented path for local documentation builds, and it is the only build tooling the repository files show.

## Conclusion

Adopt Trace if you already think in computation graphs and want a single optimizer spanning prompts, code and parameters inside one agent pipeline; skip it if you need a hosted service, a GUI, or a stable 1.0 API, since the package classifiers mark it Beta and the last release tag is v0.1.3.8 from 2025-03-23. Verify first that your feedback function can return a comparable signal on every run, and that your LLM backend works through LiteLLM, because the optimizer's step depends on both.

## FAQ

### What is microsoft/Trace?

It is a PyTorch-like Python library for training AI systems end-to-end with general feedback such as numerical rewards, natural language text or compiler errors. The README describes it as generalizing back-propagation by capturing and propagating an AI system's execution trace.

### How do I install microsoft/Trace?

The README gives `pip install trace-opt` for the published package, or `pip install -e .` after cloning for development. Python 3.9 or newer is required, and the AutoGen backend is available through the `trace-opt[autogen]` extra.

### What are the node and bundle primitives in microsoft/Trace?

`node` defines a value in the computation graph and takes a `trainable` flag, while `bundle` wraps a Python function so its body can be optimized. The README states that operations on a node are traced automatically once it is declared.

### Which LLM backend does microsoft/Trace use?

Starting with v0.1.3.5 the README says LiteLLM is the default backend. For backward compatibility an AutoGen backend is offered through the `[autogen]` install tag, pinned at `autogen-agentchat==0.2.40` in the project metadata.

### What licence does microsoft/Trace use?

The repository is MIT licensed, declared in `setup.py` and shown as a badge in the README. That permits modification and redistribution as long as the licence text is included, though the repository does not itself give legal guidance.

## Sources

- [License: MIT](https://github.com/microsoft/Trace/blob/main/LICENSE)
- [microsoft/Trace on GitHub](https://github.com/microsoft/Trace)
- [Project website](https://microsoft.github.io/Trace/)
- [README](https://github.com/microsoft/Trace/blob/main/README.md)
- [Releases](https://github.com/microsoft/Trace/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/microsoft-trace
