# neuml/rag: a Streamlit RAG app on top of txtai, with vector and graph retrieval

> neuml/rag is a Streamlit application that pairs txtai's embeddings index with an LLM, exposing vector RAG and graph RAG through one query box. The repository is small, the configuration is environment variables, and the interesting part is the graph path syntax.

**neuml/rag** — 🚀 Retrieval Augmented Generation (RAG) with txtai. Combine search and LLMs to find insights with your own data.

- Repository: https://github.com/neuml/rag
- Stars: 465 · Forks: 40
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/neuml-rag

## What neuml/rag actually is, and who it is for

The README describes this project as a Retrieval Augmented Generation (RAG) Streamlit application backed by txtai. It is not a library and not a framework. It is one Python file, rag.py, plus a requirements list, a Dockerfile and an images directory. That shape tells you most of what you need to know: the project is an application you run, not a component you import.

The problem it addresses is the gap between having a vector index and having something a person can query. txtai already handles embeddings, indexing and LLM pipelines. What it does not ship is a chat-style front end with a query box, an upload path for new data, and a way to visualise why a particular answer was produced. neuml/rag supplies that layer.

The intended user is someone who already has text they want to ask questions of, and who is willing to run a Python process or a container rather than integrate an SDK. The README's demo shows a basic vector RAG query and a Graph RAG query with uploaded data, which is the workflow the project optimises for: upload, ask, inspect the graph.

It is a poor fit for anyone who needs an HTTP API to call from another service. There is no API surface documented in the README. The interface is Streamlit, which means a browser session, and the extension point is editing rag.py.

## Vector RAG and Graph RAG: two retrieval paths behind one query box

The README splits the application into two categories. Vector RAG runs a vector search for the top N matches to the user's input, places those matches into a prompt, and returns the LLM's answer. The example given is `Who created Linux?`, which runs against the Embeddings index and hydrates a prompt with the matching documents.

Graph RAG replaces the vector search with graph path queries. The README lists three forms. A graph query prefixed with `gq: ` starts with a vector search for the top n results and then expands them through a graph network stored alongside the vector database. A graph path query takes a list of concepts separated by arrows, finds the nodes closest to those concepts, and traverses the graph to build a context of related nodes. The third form combines them: a path traversal first, then a graph query scoped to that traversal.

The node model is worth stating plainly, because it constrains how you should chunk your data. According to the README, every node in the graph is a section, meaning a paragraph, and the node labels are generated at upload time by an LLM prompt that applies a topic label. So the graph is not a hand-curated ontology. It is derived from your text at ingestion, which means ingestion is slower and more expensive than a plain embedding pass, and the quality of the graph depends on the topic-labelling prompt.

Every Graph RAG response also renders the corresponding graph, which is the feature that distinguishes this from a generic chat-over-documents app. You can see which nodes the traversal reached.

## Installing neuml/rag and running a first graph query

The README gives two installation routes and recommends Docker, at least to get an idea of the application's capabilities. The image is published as neuml/rag on Docker Hub. The command below runs it detached, requests all GPUs, and maps the Streamlit port.

```bash
docker run -d --gpus=all -it -p 8501:8501 neuml/rag
```

After this, port 8501 serves the application. Drop the `--gpus=all` flag if you have no GPU, since the flag is passed to the container runtime and will fail on a host without the NVIDIA container toolkit.

The second route is a Python virtual environment, which the README recommends you use for a direct install. The requirements file pulls matplotlib, streamlit and txtai with the graph, pipeline-data and pipeline-llm extras, plus a pinned grand-cypher.

```bash
pip install -r requirements.txt
streamlit run rag.py
```

Once the app is up, the README's own example queries are the fastest way to see both retrieval paths. The default EXAMPLES variable lists four, and the graph path one is the most instructive because it exercises traversal rather than similarity search.

```text
linux -> macos -> microsoft windows
```

Typing that into the query box runs a graph path query. You should see an answer plus a rendered graph. The README notes that when a query begins with `#`, the value is read as a URL or file and loaded into the index, and that plain text works too.

```text
# txtai is an all-in-one AI framework
```

That creates a new entry in the Embeddings database. If you want uploads to survive a restart, set the PERSIST environment variable to a directory, since the README documents it as the optional directory to save index updates to.

## The configuration surface is environment variables, and it is thin in places

Everything tunable goes through environment variables, and the README tabulates them. TITLE changes the header. EXAMPLES is a semicolon-separated list of starter queries. LLM points at a model path and defaults to Qwen3-4B-Instruct-2507. EMBEDDINGS points at an embeddings database and defaults to neuml/txtai-wikipedia-slim. MAXLENGTH caps generation at 2048 for topics and 4096 for RAG. STRIPTHINK controls whether thinking text is stripped from responses and defaults to False. CONTEXT sets the RAG context size and defaults to 10. TEXTBACKEND selects the text extraction backend. DATA indexes a directory at startup. PERSIST saves index updates. TOPICSBATCH sets a batch size for LLM topic queries. SAFEOPEN enables Textractor safeopen mode and defaults to True.

Two of these deserve scrutiny. STRIPTHINK defaulting to False means a reasoning model's thinking text can appear in responses unless you change it. That is a deliberate default rather than an oversight, but it will surprise anyone who swaps in a reasoning model and wonders why the answers look verbose. TOPICSBATCH has no default, which the README marks as None; the variable exists to bound LLM topic queries during ingestion, and leaving it unset means no batching is applied.

There is no documented way to change the prompt template, the number of graph hops, or the chunking strategy from the environment. Those live in rag.py. If your requirements include any of them, you are forking the application, and the README does not describe an extension mechanism for doing so cleanly.

## Where neuml/rag is the wrong tool

The most concrete limitation comes from the Dockerfile. It installs default-jre-headless via apt because Apache Tika needs a JRE for file extraction. If your deployment target is a minimal container or an environment where you cannot add a Java runtime, the provided image will not be your base, and you will be rebuilding the image or running the Python route on a host that has Java available.

The second limitation is persistence. PERSIST is described as an optional directory to save index updates to, and its default is None. Nothing in the README suggests the application manages index state for you otherwise. If you upload documents and restart without PERSIST set, you should assume the uploads are gone. The README does not document a rollback or versioning story for the index either.

The third is the interface itself. Because it is a Streamlit app, there is no documented API, no authentication, and no notion of separate users or collections. Pointing it at a shared index means everyone who can reach port 8501 sees the same data and can add to it. For an internal prototype that is fine. For anything multi-tenant it is disqualifying.

Finally, the graph path syntax is a query language embedded in a text box. Users have to learn `gq: ` prefixes and arrow-separated concept lists. The README documents the forms, but there is no autocomplete or validation described, so a malformed path is a user problem rather than a caught error.

## How it compares to wiring txtai yourself

The obvious alternative is txtai directly, which is the dependency this project is built on. The difference is scope rather than capability. txtai gives you the Embeddings index, the graph component and the LLM pipeline as a Python library, and you decide how queries are routed, how prompts are assembled and what the user sees. neuml/rag makes those decisions for you and hands you a running application.

That trade is visible in the repository layout. There is no package directory, no CLI entry point beyond streamlit run rag.py, and no plugin interface. Choosing neuml/rag means accepting its query routing, its prompt construction, its context size default of 10, and its topic-labelling step at ingestion. Choosing txtai directly means writing the equivalent of rag.py yourself, which is the honest cost of the alternative.

A second alternative worth naming is any RAG stack that treats ingestion as a batch pipeline with an explicit schema, rather than as a side effect of a `#` query. neuml/rag deliberately makes ingestion conversational: you type `# some text` and it lands in the index. That is a genuinely useful property for exploration and a bad property for reproducibility, because the index becomes a function of who typed what.

## Licence, maintenance and upgrade cost

The repository is licensed Apache-2.0, which permits commercial use and modification with the usual attribution and notice requirements. The dependencies are not all under the same terms, and the requirements file pins grand-cypher to 0.13.0 with a comment explaining that the pin exists until an upstream pull request is released. That pin is the clearest upgrade signal in the repository: when the upstream change lands, someone has to verify the pin can be lifted. This is not legal advice; check the licences of txtai, streamlit and grand-cypher against your own distribution model.

The last push to the default branch was on 2026-06-05, the same day v0.14.0 was released. The release before that, v0.13.0, came on 2025-12-01, and v0.12.0 on 2025-11-12. So the cadence is irregular: two releases close together in late 2025, then a gap of roughly six months to v0.14.0. That pattern is consistent with a project that follows its upstream dependency rather than one that ships on a schedule.

Upgrade cost concentrates in two places. The txtai requirement is a floor, `>=9.10.0`, not a pin, so a fresh install can pull a newer txtai than the application was last exercised against. And because the application is a single file rather than a package, there is no versioned import boundary to absorb upstream changes. You upgrade by pulling the repository and re-reading rag.py.

## Conclusion

Adopt neuml/rag if you want a working RAG interface over txtai without writing the retrieval and prompt plumbing yourself, and if your data fits txtai's indexing pipeline. Do not adopt it if you need a documented API, a multi-tenant service, or a retrieval stack you intend to extend in a language other than Python. Before committing, verify three things: that your LLM and embeddings choices can be swapped through the LLM and EMBEDDINGS variables without breaking topic labelling, that PERSIST is set if you care about keeping uploaded data, and that Apache Tika plus a JRE is acceptable in your deployment, since the Dockerfile installs default-jre-headless for file extraction.

## FAQ

### What is RAG and why is it used in neuml/rag?

The README states that Retrieval Augmented Generation helps generate factually correct content by limiting the context in which an LLM can generate answers, typically through a search query that hydrates a prompt with relevant context. In this application that context comes either from a vector search or from a graph path traversal.

### How do I install neuml/rag?

The README gives two routes: a Docker image published as neuml/rag on Docker Hub, which it recommends at least to get an idea of the application's capabilities, and a Python virtual environment where you run pip install -r requirements.txt followed by streamlit run rag.py.

### How do I add my own data to the neuml/rag index?

When a query begins with a `#`, the README says the URL or file is read and loaded into the index, and the same method accepts plain text, so `# txtai is an all-in-one AI framework` creates a new entry in the Embeddings database. To keep those updates across restarts, set the PERSIST environment variable to a directory.

### What is the difference between vector RAG and Graph RAG in neuml/rag?

Vector RAG runs a vector search for the top N matches and passes them to the LLM. Graph RAG instead runs graph path queries, either expanding a vector search through a graph network with the `gq: ` prefix, traversing a path between concepts, or combining both.

### Does neuml/rag expose an API I can call from another service?

The README does not document an API. The application is a Streamlit front end started with streamlit run rag.py or the container entrypoint, and the documented interface is the query box in the browser.

## Sources

- [Issues](https://github.com/neuml/rag/issues)
- [License: Apache-2.0](https://github.com/neuml/rag/blob/master/LICENSE)
- [neuml/rag on GitHub](https://github.com/neuml/rag)
- [README](https://github.com/neuml/rag/blob/master/README.md)
- [Releases](https://github.com/neuml/rag/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/neuml-rag
