Model or dataset
NVIDIA-NeMo/Guardrails avatar
NVIDIA-NeMo/Guardrails

NVIDIA NeMo Guardrails: Programmable Rails for LLM Chat Applications

NeMo Guardrails is an open-source toolkit for easily adding programmable guardrails to LLM-based conversational systems.

7,218 stars855 forksPythonNOASSERTION

At a glance

What is it?
NeMo Guardrails wraps an LLM call in a Python layer that can enforce dialog flows, topic boundaries and input/output checks. It is a Colang-driven toolkit for teams that already run their own model endpoint, not a drop-in hosted filter.
Who is it for?
Adopt NeMo Guardrails if you already control the model endpoint and want dialog flows and topic rails expressed as reviewable files rather than prompt text. Skip it if you need a hosted filter in front of a vendor API you do not control, or if you cannot accept a Python service in the request path.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem NeMo Guardrails addresses in an LLM chat stack

A chat application that calls a model directly has exactly one control point: the prompt. Everything else, such as refusing a topic, forcing an authentication step before answering, or returning structured data instead of prose, has to be argued for in natural language and hoped for at sampling time. NeMo Guardrails inserts a layer between application code and the model so those behaviours become configuration instead of persuasion.

The README frames the toolkit as a way to add programmable guardrails to LLM-based conversational applications, and lists the intended outcomes: keeping an assistant off unwanted topics, steering it along predefined dialog paths, enforcing standard operating procedures such as authentication and support flows, controlling language style, and extracting structured data. The repository's use-case list is narrower and more concrete than a general safety pitch: question answering over documents with fact-checking and output moderation, domain-specific assistants that must stay on topic, custom LLM endpoints that need safer customer interaction, and LangChain chains wrapped with a rails layer.

The audience is therefore a developer who owns the serving path. If your application calls an external chat API and you cannot put Python in between, this toolkit has nowhere to sit. The README's own integration story assumes you replace the call to the LLM with a call to the rails object, which is a code change in your application, not a network hop you can add at the edge.

How the rails layer sits between your code and the model

The mechanism is a configuration directory plus a wrapper object. RailsConfig.from_path reads a config from a filesystem path, and LLMRails wraps that config. Calls that would have gone to the model go to rails.generate or rails.generate_async instead, using a message list shaped like the OpenAI Chat Completions API. The README describes this as two steps and claims the change to an existing codebase is minimal.

Underneath, the toolkit is async-first: the README states the core mechanics are implemented with the Python async model and that public methods have both sync and async versions. That matters for deployment, because a synchronous call inside an async web framework is a common source of blocking behaviour, and the async variant is the one that fits a FastAPI or aiohttp service.

Configuration is not plain YAML alone. The repository carries a CHANGELOG-Colang.md alongside the main changelog, and the examples tree has both examples/configs and a separate examples/v2_x directory, which tells you the authoring format has its own release history and that older configurations may not be shaped like newer ones. There is also a .railsignore at the repository root, implying config directories can accumulate files that the loader is expected to skip.

One design consequence worth naming: because rails live in a config directory and are loaded by path, the guardrail definition becomes a reviewable artifact in version control. That is a real advantage over prompt-embedded rules, and it also means a stale config directory is a runtime behaviour change with no code diff to review.

Installing NeMo Guardrails and running a first generate call

The README gives a single install command and states Python 3.10, 3.11, 3.12 or 3.13 as the requirement. The packaging metadata in pyproject.toml expresses the same range as requires-python = ">=3.10,<3.14", so 3.14 is outside it on both counts.

bash
pip install nemoguardrails

After installing, the first real use is loading a config directory and calling generate. The README uses a placeholder path, so you supply your own config directory; the repository ships example bots under examples/bots and example configurations under examples/configs that you can point at while you learn the layout.

python
from nemoguardrails import LLMRails, RailsConfig

# Load a guardrails configuration from the specified path.
config = RailsConfig.from_path("PATH/TO/CONFIG")
rails = LLMRails(config)

completion = rails.generate(
    messages=[{"role": "user", "content": "Hello world!"}]
)

The README shows the corresponding sample output as a JSON object with role assistant and content "Hi! How can I help you?". If you get that shape back, the config loaded and the model connector answered. If the config path is wrong or the config does not parse, the failure happens at RailsConfig.from_path, before any model call, which is the cheapest place to catch a mistake.

For a long-running service rather than a script, the packaging metadata defines an optional server extra pulling in FastAPI, Starlette, Uvicorn, aiofiles, watchdog and the openai client, and the repository Dockerfile exposes port 8000 and copies examples/bots into a /config working directory. That Dockerfile is the closest thing to a reference deployment in the repository layout; the README itself points to a separate Server Guide on docs.nvidia.com for the server path.

Where NeMo Guardrails is the wrong tool

The toolkit is classified in pyproject.toml as Development Status :: 4 - Beta. That is the project's own label, and it should set expectations for anyone planning to make it the only thing standing between an end user and a model.

Version skew is the second constraint. The README states the develop branch tracks the latest top of tree development and that the latest released version is 0.24.0, while pyproject.toml on the default branch declares version 0.25.0.dev0. A configuration written against one of those is not guaranteed to behave identically on the other, and the existence of a dedicated Colang changelog suggests the config language itself moves. Pin the package version and treat config changes as migration work.

The third constraint is dependency weight. The base install pulls fastembed and onnxruntime for local embedding, plus aiohttp, httpx, lark, pydantic, numpy and others. That is a meaningful install for what may look like a thin wrapper, and it means the rails process carries an ONNX runtime even if your own application never embeds anything.

Finally, consider what the toolkit cannot do. It cannot stop a user from reaching the model through another path, and it cannot enforce anything about a vendor-hosted assistant you do not own. It is a library inside your process. If your requirement is a network-level control on traffic you do not terminate, this is the wrong layer.

NeMo Guardrails compared with the Guardrails AI package

The two projects share a name fragment and a problem statement, and search traffic for guardrails ai and guardrails ai pypi is easy to confuse with this repository. The difference in approach is visible in how each one is invoked.

NeMo Guardrails is built around a configuration directory loaded by RailsConfig.from_path and a wrapper object whose generate method replaces your model call. Its unit of design is the conversation: dialog paths, topic boundaries, standard operating procedures. The README's use cases are conversational, and the configuration format, Colang, is expressive enough to carry its own changelog.

Guardrails AI is oriented around validators applied to outputs, which is a per-field, per-schema unit of design rather than a per-turn one. If your problem is "this JSON must match this schema and these fields must pass these checks", a validator library is a shorter path than a dialog-flow engine. If your problem is "the assistant must not discuss pricing and must complete identity verification before answering account questions", a validator has no natural place to express the ordering.

Neither is a superset. A team running a retrieval-augmented question answering bot may want output validation on the generated answer and conversational rails on the flow that produced it, and the two concerns do not collapse into one tool. What NeMo Guardrails offers that a validator library does not is the dialog-path concept; what it asks in return is that you adopt its config format and load a config directory at startup.

Maintenance, release cadence and licence status

The repository is not archived, and the last push was on 2026-09-10, so development activity is current rather than historical. The release list shows v0.24.0 on 2026-08-26, v0.23.0 on 2026-07-01 and v0.22.0 on 2026-05-22, a cadence of roughly six to eight weeks between tagged releases. The default branch is develop, which the README explicitly describes as tracking the latest top of tree development, so a dependency on the default branch is a dependency on unreleased code.

Upgrade cost has two parts. The Python package itself is versioned and installable from PyPI, so pinning is straightforward. The configuration is the harder half: Colang has its own changelog file, and the repository keeps a v2_x examples directory separate from the main configs, which is the kind of layout that appears when a format has been revised. Budget for re-reading your configs at each minor upgrade rather than assuming they carry forward.

Licensing is mixed and worth reading carefully. The GitHub metadata reports NOASSERTION, but pyproject.toml declares license = "Apache-2.0" and lists license-files = ["LICENSE*", "LICENCES*"], and the repository root contains LICENSE-Apache-2.0.txt, LICENSE.md and a separate LICENCES-3rd-party file. The presence of a third-party licence file means some bundled or vendored components carry terms other than Apache 2.0. The README's badge also points at Apache 2.0. This is not legal advice: if you redistribute the package or ship it inside a product, have someone read LICENCES-3rd-party and the individual dependency licences rather than relying on the badge alone.

Editorial conclusion

Adopt NeMo Guardrails if you already control the model endpoint and want dialog flows and topic rails expressed as reviewable files rather than prompt text. Skip it if you need a hosted filter in front of a vendor API you do not control, or if you cannot accept a Python service in the request path. Before committing, verify three things against your own setup: that your Python version is in the supported range, that your chosen LLM connector is one the docs list, and that your Colang config loads under the version you pin, since the repository ships a develop branch whose pyproject.toml reads 0.25.0.dev0 while the latest release is 0.24.0.

Frequently asked questions

How do I install NeMo Guardrails?

The README gives one command, pip install nemoguardrails, and states Python 3.10 through 3.13 as the requirement. The packaging metadata expresses the same range as requires-python >=3.10,<3.14.

How do I use NeMo Guardrails in an LLM application?

Load a configuration directory with RailsConfig.from_path, wrap it in LLMRails, and call rails.generate with a message list shaped like the Chat Completions API. The README describes this as two steps and notes both sync and async versions of the public methods exist.

How do I use NeMo Guardrails with LangChain?

The README lists LangChain chains as an optional use case and says to enable the integration by setting the NEMOGUARDRAILS_LLM_FRAMEWORK=langchain environment variable or by calling set_default_framework("langchain").

What is a guardrail in AI?

In this project's terms, rails are specific ways of controlling the output of a large language model, such as not talking about politics, responding in a particular way to specific requests, following a predefined dialog path, or extracting structured data.

How do I apply guardrails in an LLM?

The README describes adding rails by loading a configuration with RailsConfig.from_path, creating an LLMRails instance, and routing model calls through rails.generate or rails.generate_async instead of calling the model directly.

Official sources

  1. Issues
  2. NVIDIA-NeMo/Guardrails on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/nvidia-nemo-guardrails.svg)](https://hysenlabs.com/projects/nvidia-nemo-guardrails)