NVIDIA AI-Q Blueprint: a self-hosted deep research backend built on NeMo Agent Toolkit
The AI-Q NVIDIA Blueprint is an open reference example for building intelligent AI agents that connect to your enterprise data, reason using state-of-the-art models, and deliver trusted business insights.
At a glance
- What is it?
- AI-Q is an Apache-2.0 reference deployment for cited, report-style research agents, with a CLI, web UI, async jobs and an MCP server. It is a research backend, not a coding-agent harness, and the README leaves rollback and upgrade paths undocumented.
- Who is it for?
- Adopt AI-Q if you need a self-hosted research backend with citation-backed reports, an async job API and evaluation harnesses, and you have the GPU or API budget plus the operational capacity to run NeMo Agent Toolkit, Redis and a sandbox provider. Do not adopt it as a coding agent or as a drop-in RAG service; the README states it is focused on governed research workflows and is not a general-purpose coding-agent harness.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What AI-Q actually is, and who the README is written for
The README describes AI-Q as a deployable research backend built on the NVIDIA NeMo Agent Toolkit and LangChain Deep Agents. That phrasing matters more than the marketing around it. This is not a library you import and call; it is an application boundary plus a set of configuration files, deployment assets and evaluation harnesses that a platform team stands up and then points at its own models, data sources, authentication and storage. The stated target audience is teams that want to self-host that boundary, connect deployment-owned models and policy controls, and measure quality with benchmarks rather than vibes.
The scope is deliberately narrow. The README says AI-Q is focused on governed research workflows and is not a general-purpose coding-agent harness. So the intended user is an engineer or platform owner building an internal research assistant over enterprise data, not someone who wants an agent to edit their repository. The repository layout supports that reading: there are separate top-level directories for configs, deploy, frontends, mcp, skills, sources and src, and a pyproject.toml whose package name is aiq-agent. The Python requirement is >=3.11,<3.14, and the dependency list pins nvidia-nat-core==1.8.0 along with the langchain, async_endpoints, phoenix, mcp, eval, profiler, redis and security extras at the same version.
Orchestration, shallow research and structured deep research
The mechanism the README describes is a pipeline with distinct roles rather than one monolithic agent. An orchestration node classifies intent as meta or research, produces meta responses such as greetings and capability statements, and sets research depth to shallow or deep. Shallow research is a bounded, faster researcher with tool calling and source citation. Deep research is more elaborate: advisory source routing, structured planning, concurrent researcher workers, bounded source-tool batching, and a dedicated writer that produces citation-backed reports and other requested output shapes.
The v2.2.0 release notes describe a change here. The structured, concurrent deep research flow replaced an earlier three-role design and moved plan ownership out of the clarifier. That is the kind of refactor that changes prompt behaviour, so anyone upgrading from an earlier version should expect their tuned prompts and configs to behave differently even if the YAML still parses. A separate feature, report follow-up, lets a user ask questions about a completed report, create a child job that performs a cosmetic rewrite, or run delta research using the parent report as context. The README does not document how delta research reconciles contradictory new findings with the parent report, which is the question I would want answered before relying on it for anything audited.
Install and first run: clone, configure, invoke the CLI
The README's getting-started path is clone the repository, run automated setup, obtain API keys, then set environment variables. It does not print a pip install line for the application itself, so the repository is the distribution channel. Python must be 3.11, 3.12 or 3.13.
Start by cloning the default branch and running the setup script the README references. The repository ships a scripts directory and a uv.lock, so the setup path is uv-based.
git clone https://github.com/NVIDIA-AI-Blueprints/aiq.git
cd aiq
./scripts/setup.shAfter setup, the README says to obtain API keys for the model and search providers you intend to use and to set them as environment variables. The exact variable names are not in the portion of the README available here, so read the environment variables section of the README or the docs site before inventing names. The docs site is at https://docs.nvidia.com/aiq-blueprint/latest/index.html.
Once configured, the README lists several ways to run the agents: a command-line interface, a web UI, async deep research jobs, an MCP server, benchmarks and Jupyter notebooks. The CLI is the fastest way to confirm the pipeline is wired correctly, because it exercises the orchestration node and the researchers without the frontend in the way. What you should see is a classified intent, and for a research question, a cited answer from the shallow path or a longer report from the deep path.
The YAML config layer is where the real work happens
Workflow configuration is the part of AI-Q that determines whether the deployment is useful. The README states that YAML configs define agents, tools, LLMs and routing behavior so workflows can be tuned without code changes, and the repository keeps a top-level configs directory for exactly that. All agents, including the orchestration node, shallow researcher, deep researcher and clarifier, are composable and can run standalone or as part of the full pipeline.
That composability is genuinely useful for debugging, and it is also the source of the maintenance burden. Every config is a contract with the pinned nvidia-nat-core==1.8.0 stack and with the prompt templates shipped as Jinja files inside the package, for example the prompts/*.j2 files declared for aiq_agent.agents.chat_researcher, clarifier, data_science, deep_researcher, report_rewriter and shallow_researcher. If you fork prompts, you are now maintaining a fork of behaviour that upstream is actively changing; the v2.2.0 notes show plan ownership moving between components, which is precisely the kind of change that invalidates a local prompt patch. The README does not describe a supported override mechanism for prompts separate from editing them, so treat prompt customisation as a maintenance commitment rather than a configuration toggle.
Sources, skills and the sandbox contract
Source coverage is broader than a single web search tool. The README lists paper search through Serper, SerpAPI and SearchAPI; You.com tools for web, contents, general research and finance research; Nimble for configurable web search; and focused profiles demonstrating DuckDuckGo news, Polymarket, OpenSearch and Azure AI Search knowledge retrieval. A data source registry lets UI toggles and request payloads select web, paper, enterprise, collaboration and knowledge-layer sources per message.
Skills are the more interesting design decision. Built-in research and synthesis skills are exposed as host-side, read-only definitions, and skills that invoke code use what the README calls a provider-neutral sandbox contract. Modal and OpenShell each create one physical sandbox per deep-research job. An opt-in rich-file capture feature checkpoints manifest-declared files after successful sandbox commands, finalizes on success or failure, stores bytes in SQL or S3-compatible storage, and delivers metadata to the Files tab live and on replay. The trade-off is visible: one sandbox per job gives isolation and reproducibility, and it also means job cost and startup latency scale with the number of concurrent deep-research jobs. If your workload is hundreds of small parallel queries, this architecture is heavier than a shared-worker design, and the README offers no pooling option.
Production API, authentication and the security surface you own
AI-Q ships REST endpoints, async job ownership, per-user OAuth-protected MCP sources, token validator entry points and provider lifecycle hooks for authenticated deployments. A separate public MCP server exposes stateless research tools and the README scopes it to trusted networks. Policy controls are opt-in: NeMo Guardrails middleware covers selected workflow and agent boundaries, and application-level encryption can protect final async output plus selected artifact-event content.
The word selected is doing a lot of work in that sentence. Guardrails cover some boundaries, not all, and encryption covers final async output and some artifact-event content, not the whole datastore. Anyone treating these as a complete security posture will be disappointed. The SECURITY.md file in the repository is the place to check for the project's own vulnerability reporting process, and the LICENSE-THIRD-PARTY file lists bundled third-party components. Observability is better specified: NeMo Relay preserves task, named-agent, LLM and tool hierarchy across interactive turns and async researchers, ATOF feeds local debugging and tokenomics reports, and OTEL exports traces to external backends. If your organisation already runs an OTEL collector, that is the least friction path to production visibility.
Where AI-Q is the wrong tool, and what to use instead
AI-Q is the wrong choice if you want a general coding agent, a lightweight RAG endpoint, or a single-process script you can run on a laptop with no external services. The dependency list pulls in Redis, a profiler, an evaluation package and a security extra at pinned versions, and deep research expects a sandbox provider. That is a lot of moving parts for a question-answering feature.
A real alternative in the same problem space is a plain LangGraph or LangChain agent you assemble yourself against your own retrieval layer. The difference in approach is significant. AI-Q hands you an opinionated pipeline: intent classification, a clarifier, advisory source routing, concurrent researcher workers, bounded tool batching and a dedicated writer, all driven by YAML and measured by built-in benchmarks such as FreshQA and DeepResearch. A hand-rolled LangGraph agent gives you none of that structure, but it also imposes no orchestration model, no sandbox contract, no job-ownership semantics and no pinned NeMo Agent Toolkit version. If your research questions are narrow and you already know which sources matter, the hand-rolled path is less code to own. If you need citation-backed reports, per-user source authorisation and an evaluation harness you can point at prompt changes, AI-Q is doing work you would otherwise write and maintain yourself.
Maintenance, licensing and what the README does not tell you
The repository is not archived, and the last push was on 2026-09-02. The most recent release is v2.2.1 from 2026-08-22, following v2.2.0 on 2026-08-17 and a release candidate on 2026-08-21. The default branch is develop, and the README points anyone reproducing published benchmark numbers to the drb1 and drb2 branches instead, which means the code you run in production and the code behind the leaderboard results are not the same tree. That is a reproducibility caveat worth knowing before you compare your own numbers to a published figure.
Licensing is Apache-2.0 for the project itself, with a separate LICENSE-THIRD-PARTY file for bundled components. Apache-2.0 is permissive and includes a patent grant, but the third-party file is the one your legal review will actually want, because the model and search providers you connect are governed by their own terms, not by this repository's licence. Nothing here is legal advice; the practical point is that the project licence is the easy part and the provider contracts are not.
Upgrade cost is the open question. The README does not document a rollback procedure, a migration path between minor versions, or a compatibility matrix for configs across releases. Given that v2.2.0 restructured the deep research flow and moved plan ownership, and that nvidia-nat-core is pinned exactly rather than ranged, an upgrade is a coordinated change across the package pin, your YAML configs and any prompt overrides you have made. Verify what changed in CHANGELOG.md before you move a working deployment.
Editorial conclusion
Adopt AI-Q if you need a self-hosted research backend with citation-backed reports, an async job API and evaluation harnesses, and you have the GPU or API budget plus the operational capacity to run NeMo Agent Toolkit, Redis and a sandbox provider. Do not adopt it as a coding agent or as a drop-in RAG service; the README states it is focused on governed research workflows and is not a general-purpose coding-agent harness. Before committing, verify three things in your own environment: that your chosen sandbox provider (Modal or OpenShell) can create one sandbox per deep-research job, that your OAuth token validator and MCP source configuration satisfy your security review, and that your Python version falls inside the requires-python range of >=3.11,<3.14, since the pinned nvidia-nat-core==1.8.0 stack will not move independently of the rest of the project.
Frequently asked questions
What are Nvidia AI blueprints?
AI-Q is described in its README as an open reference example for building AI agents that connect to enterprise data and produce trusted business insights, and it is published under the NVIDIA-AI-Blueprints organisation. It is a deployable research backend rather than a general-purpose agent framework.
Is the NVIDIA AI-Q Blueprint available on GitHub?
Yes. The repository is NVIDIA-AI-Blueprints/aiq, with develop as the default branch, and the README's getting-started path begins by cloning it. Benchmark reproduction uses the separate drb1 and drb2 branches rather than the default branch.
What Python version does the AI-Q agent require?
The pyproject.toml sets requires-python to >=3.11,<3.14, so Python 3.11, 3.12 and 3.13 are in range. The same file pins nvidia-nat-core==1.8.0 and the related NeMo Agent Toolkit extras at the same version.
Does AI-Q need a sandbox provider to run deep research?
The README states that skills invoking code use a provider-neutral sandbox contract, and that Modal and OpenShell create one physical sandbox per deep-research job. Shallow research is described as a bounded researcher with tool calling and source citation, which is the lighter path.
Can AI-Q be used as a coding agent?
No. The README states that AI-Q is focused on governed research workflows and is not a general-purpose coding-agent harness. Its agents are an orchestration node, a clarifier, a shallow researcher, a deep researcher and a report rewriter.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/nvidia-ai-blueprints-aiq)