# NLWeb: Adding Conversational Natural Language Interfaces to Websites Using Schema.org and MCP

> NLWeb is an open-source Python implementation and protocol specification that lets websites expose natural language query endpoints backed by Schema.org structured data, with every instance also acting as an MCP server accessible to AI agents.

**nlweb-ai/NLWeb** — Main reference implementation for NLWeb, implemented in Python.

- Repository: https://github.com/nlweb-ai/NLWeb
- Stars: 6,257 · Forks: 699
- Language: Python
- License: MIT
- Published: 2026-09-22 · Updated: 2026-09-22 · Language: en
- Canonical page: https://hysenlabs.com/projects/nlweb-ai-nlweb

## What NLWeb Solves and Who It Is For

NLWeb addresses the gap between Schema.org structured data, which over 100 million websites already publish for SEO and feed syndication, and conversational query interfaces that AI assistants and users expect. Most websites that publish Schema.org markup for search engines have no mechanism for a user or an AI agent to ask a natural language question against that structured content.

NLWeb provides a simple protocol and a reference implementation that turns the existing Schema.org markup into a queryable endpoint. The `ask` method accepts a natural language question and returns a JSON response using Schema.org vocabulary. Every NLWeb instance also registers as an MCP server, so AI agents that use the Model Context Protocol can query it using the same `ask` interface.

The primary audience is web developers and platform engineers who maintain content-heavy sites and want to add AI-accessible natural language query capabilities without building a custom retrieval-augmented generation (RAG) pipeline from scratch.

## Protocol and Architecture

NLWeb has two components. The first is a protocol specification, documented at https://nlweb.ai/spec, that defines how a natural language question is posed to a website and what the response format looks like. Responses use Schema.org as the vocabulary for returned data.

The second is the reference implementation, organized into five modules in this repository:

- AskAgent: the core query agent that handles natural language queries against websites using Schema.org structured data, with connectors for LLMs and vector databases, ingestion tools, and a sample web UI
- AgentFinder: a discovery service for finding and routing to NLWeb agents across the web
- DataFinder: a natural language to SQL translator for enterprise data sources such as HubSpot, Dynamics 365, and Jira, using schema.org-based ontology mappings
- ModelRouter: LLM model routing and scoring, selecting cost-effective models that meet quality thresholds
- NLWebScorer: neural scorer models for ranking and evaluating search result quality

Supporting directories include `config/` for YAML configuration files for LLM providers, embedding models, retrieval backends, and web server settings, and `static/` for the frontend web UI assets.

## Platform Support and Backend Options

NLWeb is designed to be backend-agnostic on both the vector store and the LLM side.

Supported vector stores: Qdrant, Snowflake, Milvus, Azure AI Search, Elasticsearch, Postgres, and Cloudflare AutoRAG. Setup documentation for each is linked in the README under `docs/setup-<name>.md`.

Supported LLMs: OpenAI, DeepSeek, Gemini, Anthropic, Inception, and HuggingFace. Configuration for the LLM provider goes in the YAML files under `config/`.

Platform support covers Windows, macOS, and Linux. The README notes that mobile device support is planned but not yet available.

## Getting Started with the Reference Implementation

The repository provides a setup script for initial configuration:

```bash
bash setup.sh
```

The `requirements.txt` at the root delegates to the actual requirements file:

```bash
-r code/python/requirements.txt
```

The `.env.template` file at the root contains the template for API keys and configuration variables. Copy it to `.env` and fill in the values for your chosen LLM provider and vector store.

For Azure deployments, the repository includes `azure-web-config.txt`, `azure-app-service-config.txt`, and a deployment script `deploy_azure_webapp.sh`. OAuth setup is handled by `setup_oauth.sh`.

The getting started guide for a local setup is at `docs/nlweb-hello-world.md`. The demo directory contains example data sources and extraction scripts, including `demo/extract_github_data.py` and `demo/import_clinical_trials.py`, which show how to ingest structured content into the vector store. The demo's `OtherDataSources.txt` lists additional structured data sources you can use to test the pipeline beyond the bundled examples. Documentation for customizing prompts is at `docs/nlweb-prompts.md`, and control flow modification is covered in `docs/nlweb-control-flow.md`.

## MCP Integration and the Agent Web Vision

Every NLWeb instance exposes an MCP server endpoint in addition to the REST API documented at `docs/nlweb-rest-api.md`. This means AI agents that support MCP can discover and query NLWeb instances through the same protocol they use for other tools.

The README frames NLWeb's relationship to MCP using an analogy: NLWeb is to MCP/A2A what HTML is to HTTP. The protocol provides a semantic layer for agents to interact with web content through natural language rather than HTML scraping or custom API integrations.

AgentFinder, one of the five modules, handles discovery: it provides a service for finding NLWeb agents across the web, similar to how a DNS-based service discovery works but for natural language query endpoints. This module is intended to support a future where many sites expose NLWeb endpoints and agents need to route queries to the appropriate one.

## Limitations and When NLWeb Is Not the Right Fit

NLWeb's reference implementation assumes that the site's content can be represented as structured lists using Schema.org types. Products, recipes, attractions, and reviews are the examples given in the README. Sites whose primary content is long-form editorial prose without Schema.org markup are not well served by the current implementation.

The repository explicitly notes that CI/CD pipelines are not yet included and that contributions to add automated testing or deployment workflows are welcome. This means the reference implementation does not have the production guardrails that an enterprise deployment would require.

For enterprise data sources, DataFinder handles HubSpot, Dynamics 365, and Jira through natural language to SQL translation. Other enterprise systems are not covered and would require custom connector development.

A standard RAG pipeline built with LangChain or LlamaIndex is a direct alternative. The key difference is that NLWeb relies on pre-existing Schema.org markup rather than embedding arbitrary text chunks. Teams whose content is already Schema.org-typed benefit from NLWeb's approach; teams with unstructured content benefit more from a generic RAG pipeline.

## License, Maintenance, and Attribution

The repository is MIT licensed and is not archived. The last push was on 2026-08-11. There are no GitHub releases. The contact email for project questions is NLWebSup@microsoft.com, indicating Microsoft organizational backing, though the repository is at the nlweb-ai GitHub organization rather than microsoft.

The `RAI_TRANSPARENCY.md` file addresses responsible AI considerations for the NLWeb system. The repository includes a `CLAUDE.md` file at the root, suggesting contributors use Claude for development tooling within the project. Trademark guidance in the README covers authorized use of Microsoft trademarks appearing in the project.

The top-level directory also includes `check_dependencies.py`, a utility script for verifying that the required LLM and vector store dependencies are reachable before running the main AskAgent. This is a practical starting point for debugging configuration problems. The `ruff.toml` linting configuration and the `.vscode/` directory settings indicate the project has a defined development environment and code style standard.

## Conclusion

NLWeb suits teams who operate content-rich websites with structured Schema.org markup and want to add a conversational query layer that serves both human users and AI agents through MCP without building a custom RAG pipeline from scratch. It is not a good fit for sites whose content is unstructured prose with no Schema.org markup: the implementation's core assumption is that structured lists of products, recipes, reviews, or similar Schema.org-typed records are already present. The CI/CD pipeline is absent by the README's own admission, so treat the reference implementation as a starting point for integration rather than a production-hardened service. The last push was on 2026-08-11.

## FAQ

### What is NLWeb?

NLWeb is an open-source protocol and Python reference implementation for adding natural language query endpoints to websites. It uses Schema.org structured data as the query target and returns responses in Schema.org JSON format. Every NLWeb instance also acts as an MCP server, making it accessible to AI agents.

### What does NLWeb do?

NLWeb takes natural language questions posed to a website and returns structured JSON responses based on the site's Schema.org data. It handles query routing, vector store retrieval, LLM reasoning, and response formatting through its AskAgent module. It also exposes an MCP endpoint for AI agent access.

### NLWeb vs RAG: what is the difference?

A standard RAG pipeline embeds arbitrary text chunks into a vector store and retrieves them based on semantic similarity. NLWeb is designed around Schema.org structured data: it assumes site content is already typed with Schema.org vocabulary (products, recipes, reviews, etc.) and uses that structure for retrieval and response formatting. NLWeb also natively exposes an MCP server interface, which generic RAG pipelines do not.

## Sources

- [Issues](https://github.com/nlweb-ai/NLWeb/issues)
- [License: MIT](https://github.com/nlweb-ai/NLWeb/blob/main/LICENSE)
- [nlweb-ai/NLWeb on GitHub](https://github.com/nlweb-ai/NLWeb)
- [README](https://github.com/nlweb-ai/NLWeb/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/nlweb-ai-nlweb
