# NVIDIA NeMo Agent Toolkit: Profiling and Optimization for Production AI Agents

> NVIDIA NeMo Agent Toolkit (NAT) is an Apache-licensed Python library that adds profiling, observability, evaluation, and fine-tuning capabilities to AI agents built with LangChain, CrewAI, LlamaIndex, and other frameworks. It is framework-agnostic and works alongside existing orchestration tools rather than replacing them, with version 1.9.0 released in September 2026.

**NVIDIA/NeMo-Agent-Toolkit** — The NVIDIA NeMo Agent toolkit is an open-source library for efficiently connecting and optimizing teams of AI agents.

- Repository: https://github.com/NVIDIA/NeMo-Agent-Toolkit
- Website: https://docs.nvidia.com/nemo/agent-toolkit/latest/
- Stars: 2,650 · Forks: 773
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/nvidia-nemo-agent-toolkit

## What NAT Adds to Existing Agent Frameworks

Frameworks like LangChain, CrewAI, LlamaIndex, and Microsoft Semantic Kernel give engineers the tools to build and coordinate AI agents. They do not answer what happens to performance at scale, which prompt configuration produces the best results, or how to fine-tune the underlying model for a specific workflow. NVIDIA NeMo Agent Toolkit addresses those gaps.

The README positions NAT as working side-by-side with these frameworks, adding the instrumentation necessary for observing, profiling, and optimizing agents. It does not replace LangChain or CrewAI: it wraps around them. According to the README, it works with popular frameworks including LangChain, LlamaIndex, CrewAI, Microsoft Semantic Kernel, and Google ADK, as well as custom enterprise frameworks and simple Python agents. The package is installed as `nvidia-nat` and requires Python 3.11 through 3.13.

## Installing NAT and Getting Started

The pyproject.toml names the package `nvidia-nat`. The project requires Python 3.11 or newer (with an upper bound of Python 3.14, exclusive). The README links to official documentation at docs.nvidia.com/nemo/agent-toolkit/latest/ and to a Colab notebook badge for quick exploration.

The repository has a `pyproject.toml` and a `uv.lock` file, indicating both uv and pip installation paths are available. The latest release at the time of the last push was v1.9.0, published on 2026-09-10, with v1.8.0 preceding it on 2026-06-16 and v1.7.0 on 2026-05-21.

A migration notice in the README flags version 1.5.0 as simplifying package installation and dependency management. Engineers upgrading from before 1.5.0 should consult the Migration Guide at `docs/source/resources/migration-guide.md`. The built-in chat UI can be launched to interact with agents, visualize output, and debug workflows, as described in `docs/source/run-workflows/launching-ui.md`.

## Profiling Agents Down to Individual Tokens

The profiler is described in the README as covering entire workflows from the agent level all the way down to individual tokens, identifying bottlenecks, analyzing token efficiency, and guiding developers in optimizing their agents. This granularity is unusual: most production monitoring tools profile at the request or function call level, not at the token level.

Token-level profiling matters for agentic workflows specifically because agent steps vary significantly in how much reasoning they require. A step that calls a tool may produce very few output tokens; a planning step may produce hundreds. Understanding which steps consume the most tokens and whether those tokens contribute to task success is the kind of analysis the profiler supports.

The observability integration at `docs/source/run-workflows/observe/observe.md` covers tracing execution flows and tracking performance in production. Native LangSmith tracing is also supported, allowing teams who already use LangSmith to observe NeMo Agent Toolkit workflows within their existing LangSmith account, run evaluation experiments, and compare outcomes across prompt versions.

## Evaluation, Prompt Optimization, and Fine-Tuning

The evaluation system at `docs/source/improve-workflows/evaluate.md` validates and maintains agent accuracy through offline evaluation. This is not a benchmark against a fixed dataset: it is designed for evaluating the specific workflows an organization is running.

The Hyper-Parameter and Prompt Optimizer automatically identifies the best configuration and prompts, documented at `docs/source/improve-workflows/optimizer.md`. This is a practical feature: prompt sensitivity is one of the hardest problems in production agent deployment, and manual prompt tuning does not scale.

Beyond prompt optimization, NAT supports fine-tuning LLMs with reinforcement learning, documented at `docs/source/improve-workflows/finetuning/index.md`. The README describes this as fine-tuning specifically for an agent and training intrinsic information about a workflow directly into the model. This is a longer-cycle improvement path compared to prompt optimization but addresses root-cause accuracy problems that prompting alone cannot fix.

## New in v1.9.0: APP, Dynamo, FastMCP, and the Plugin API

The README's new features section for the current release covers four additions. Agent Performance Primitives (APP) introduce framework-agnostic performance acceleration for graph-based frameworks including LangChain, CrewAI, and Agno, using parallel execution, speculative branching, and node-level priority routing.

Dynamo Runtime Intelligence is listed as experimental. It automatically infers per-request latency sensitivity from agent profiles and applies runtime hints for cache control, load-aware routing, and priority-aware serving. This is relevant for teams running NAT-instrumented agents on NVIDIA Dynamo infrastructure.

FastMCP Workflow Publishing allows NAT workflows to be published as MCP servers using the FastMCP server runtime. The Public Plugin API enables third-party maintainers to build integrations against a published interface rather than internal packages: current external plugins include NeMo-Agent-Toolkit-Tavily, NeMo-Agent-Toolkit-Redis, and NeMo-Agent-Toolkit-ATR, all maintained outside the main repository.

## Limitations and What NAT Is Not

NAT does not replace an orchestration framework. Teams who want a single library that both builds and optimizes agents will need both NAT and a framework like LangChain. NAT is specifically the measurement and improvement layer, not the construction layer.

The Dynamo Runtime Intelligence feature is explicitly marked experimental in the README. Teams who want to rely on it for production workloads should track its status in the CHANGELOG.md before building on it.

The plugin API for third-party integrations is explicitly noted in the README as one where those packages are managed outside the NAT repository and may release, change, or run CI on independent schedules. This means the reliability of a third-party plugin depends on that plugin's own maintenance, not on NVIDIA's.

The full dependency set is a monorepo-style structure with multiple packages under the packages/ directory. This gives fine-grained control over what gets installed but requires understanding the package structure before adding NAT to a project with strict dependency constraints.

## NAT Compared to Framework-Native Observability

LangChain provides LangSmith for tracing and monitoring. CrewAI includes its own telemetry. These framework-native tools are designed for their specific framework and do not provide cross-framework consistency. NAT's framework-agnostic observability layer means a team that uses both LangChain and a custom Python agent in the same system can observe both through the same NAT instrumentation.

The deeper difference is that NAT adds optimization capabilities: prompt optimization, reinforcement learning fine-tuning, and the profiler that goes to the token level. LangSmith, by comparison, is primarily an observability and evaluation platform. For teams whose primary concern is observability, the combination of LangSmith and a framework's native tooling may be sufficient. For teams who need the optimization and fine-tuning pipeline, NAT provides capabilities that framework-native tools do not.

## Conclusion

NVIDIA NeMo Agent Toolkit is the right choice for engineering teams who have built AI agents on LangChain, CrewAI, LlamaIndex, or a custom framework and now need to understand and improve their production behavior. The profiling, evaluation, and fine-tuning pipeline answers the specific question of why an agent behaves a certain way and what to do about it, which is not answered by the orchestration frameworks themselves. Teams that are still building an initial agent prototype and have not yet encountered the performance or accuracy problems NAT is designed to solve will find the setup overhead heavier than their current needs. Before adopting it, check whether the capabilities you need, such as Dynamo Runtime Intelligence and Agent Performance Primitives, are still experimental or fully released, since the README marks some features with that distinction explicitly.

## FAQ

### What is the NeMo Agent Toolkit?

NVIDIA NeMo Agent Toolkit is an Apache-licensed Python library that adds profiling, observability, evaluation, prompt optimization, and fine-tuning to AI agents built with any framework including LangChain, CrewAI, LlamaIndex, and custom Python agents. It is installed as the package nvidia-nat and requires Python 3.11 through 3.13.

### What is NVIDIA NeMo Agent Toolkit and how does it differ from LangChain?

NVIDIA NeMo Agent Toolkit is not an orchestration framework: it works alongside frameworks like LangChain, CrewAI, and LlamaIndex to add profiling, observability, evaluation, and fine-tuning capabilities. LangChain handles agent construction and orchestration. NAT handles measuring and improving what those agents do in production.

### How do I install NeMo Agent Toolkit?

The package is named nvidia-nat and installs via pip or uv. It requires Python 3.11 or newer. The repository includes a uv.lock file. The official documentation at docs.nvidia.com/nemo/agent-toolkit/latest/ has the full setup instructions, and the Migration Guide at docs/source/resources/migration-guide.md covers breaking changes introduced in version 1.5.0.

## Sources

- [License: Apache-2.0](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/develop/LICENSE)
- [NVIDIA/NeMo-Agent-Toolkit on GitHub](https://github.com/NVIDIA/NeMo-Agent-Toolkit)
- [Project website](https://docs.nvidia.com/nemo/agent-toolkit/latest/)
- [README](https://github.com/NVIDIA/NeMo-Agent-Toolkit/blob/develop/README.md)
- [Releases](https://github.com/NVIDIA/NeMo-Agent-Toolkit/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/nvidia-nemo-agent-toolkit
