Model or dataset
pathwaycom/llm-app avatar
pathwaycom/llm-app

pathwaycom/llm-app: ready-to-run RAG templates that stay in sync with live data

Ready-to-run cloud templates for RAG, AI pipelines, and enterprise search with live data. Docker-friendly.Always in sync with Sharepoint, Google Drive, S3, Kafka, PostgreSQL, real-time data APIs, and more.

58,897 stars1,500 forksJupyter NotebookMIT

At a glance

What is it?
The repository is a collection of Docker-friendly RAG and enterprise search templates built on the Pathway Live Data Framework. It is aimed at developers who want live indexing over Sharepoint, Google Drive, S3, Kafka or PostgreSQL without assembling a separate vector database, cache and API layer.
Who is it for?
Adopt it if you need a RAG service whose index follows additions, deletions and updates in Sharepoint, Google Drive, S3, Kafka or PostgreSQL, and you are willing to build on the Pathway framework: the pyproject.toml pins pathway = "^0.12.0" and Python ">=3.10,<3.13".
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 86 days ago.
What is it written in?
Mainly Jupyter Notebook, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

What pathwaycom/llm-app actually ships

This repository is not a library you import. It is a set of application templates, each in its own directory under templates/, and the pyproject.toml sets package-mode = false, so there is no llm-app wheel to install. The README lists eight templates: a basic question-answering RAG app, a live document indexing service that can act as a retriever backend for Langchain or Llamaindex, a multimodal RAG pipeline using GPT-4o in the parsing stage, an unstructured-to-SQL pipeline that loads financial PDFs into a PostgreSQL table and answers questions by generating SQL, an Adaptive RAG app, a fully private RAG app using Mistral and Ollama, a slides search app, and a video RAG pipeline built on TwelveLabs Pegasus and Marengo.

The intended audience is a developer who has documents in a source that changes, and who does not want to wire together a vector database, a cache and an API framework as separate services. The README states the templates scale up to millions of pages of documents. Treat that as a design target from the project, not a measured result. What is verifiable from the repository layout is that each template carries its own README.md with run instructions, which is the only place per-template detail lives.

How the live indexing pipeline is put together

The mechanism is the Pathway Live Data Framework, which the README describes as a standalone Python library with a Rust engine built into it. The templates depend on pathway = "^0.12.0" rather than on a vector database client. Data source synchronization and API request serving both come from that framework.

Indexing is the part worth understanding. The README says indexing is built in and done in-memory with cache, covering vector search, hybrid search and full-text search. The default vector index is based on the usearch library, and the hybrid full-text indexes use Tantivy. That means the retrieval layer is not a separate process you deploy and monitor; it lives inside the same application process that reads from the data source.

The README frames the contrast explicitly, listing a vector database such as Pinecone, Weaviate or Qdrant, a cache such as Redis, and an API framework such as FastAPI as the modules you would otherwise integrate and maintain separately. The templates expose an HTTP API for a frontend, run as Docker containers, and some include an optional Streamlit UI for testing. The README also says swapping a vector index for a hybrid index is a one-line change, though it does not show that line, so verify it in the template source before planning around it.

Running a template: install and first use

The README does not give a single install path for the whole repository. It states that each app template contains a README.md with instructions on how to run it, and points to the Pathway website for more templates. So the first step is choosing a template directory, for example templates/question_answering_rag/ or templates/private_rag/.

The only dependency information available at the repository root comes from pyproject.toml. It declares the Python range and the Pathway version, and it sets package-mode = false, which is why there is no repository-level install command to give here.

toml
python = ">=3.10,<3.13"
pathway = "^0.12.0"

Those are the two dependency lines as they appear in pyproject.toml. The caret constraint allows any 0.12.x release, so pinning an exact patch version is a deliberate choice rather than something the repository mandates. Python must satisfy ">=3.10,<3.13"; version 3.13 is outside it.

The README says the apps can run as Docker containers and expose an HTTP API, but it does not publish a shared image name, a docker run command or a port at the root level. Any container invocation has to come from the individual template README. That is a real gap for anyone evaluating the project as a whole rather than one template at a time.

Reading templates/question_answering_rag/README.md is the actual first use. It is where the data source connector, the model provider credentials and the HTTP API details are defined for that template. The root README does not document those values, and inventing them here would be wrong.

Where the template model breaks down

The first limitation is packaging. Because package-mode = false, you cannot pip install this repository and import a stable API from it. Upgrading means re-reading a template's README and reapplying your changes, not bumping a version. If you fork a template and modify the pipeline, you own the merge.

The second is maturity signalling. The pyproject.toml classifiers include Development Status :: 3 - Alpha. The version is 0.3.6, and the Pathway dependency is on a 0.x series. Any team with a policy against alpha dependencies has its answer already.

The third is that the in-memory indexing design is a trade-off, not a free win. The README presents the absence of a separate vector database as an advantage, and for a single service over a bounded corpus that is reasonable. It also means the index is rebuilt and held by the application process rather than being an independent service you can scale, snapshot or query from other clients. The README does not document index persistence, snapshotting or rollback behaviour, so if your recovery plan assumed a vector database you can restore, that plan does not carry over.

Finally, the repository is a template collection, so there is no single upgrade path or changelog covering all eight apps. A fix in one template tells you nothing about the others.

How this differs from LangChain plus a vector database

The closest familiar alternative is assembling the same application from LangChain or LlamaIndex with an external vector store. The difference is where synchronization lives. In that stack, you write and schedule the ingestion job that notices a changed document, re-embeds it and upserts it. The retriever only sees what your job pushed.

In these templates, synchronization is the framework's job. The README states the apps connect and sync all new data additions, deletions and updates from the file system, Google Drive, Sharepoint, S3, Kafka, PostgreSQL and real-time data APIs. Deletions matter here: a scheduled re-index job often leaves stale chunks behind unless you handle removal explicitly, and the README claims deletion handling as part of the sync.

The second difference is component count. The LangChain route gives you interchangeable pieces, which is worth having when you need a specific embedding model or a managed vector store with its own SLA. The Pathway route trades that interchangeability for one process that owns ingestion, indexing and serving. Note that the integration is not entirely closed: the README says the document indexing template can serve as a retriever backend for a Langchain or Llamaindex application, so you can keep your existing orchestration and use this only for live indexing.

Maintenance cost and the MIT licence

The repository is not archived. The last push date is not stated in the repository information, so no claim about how recently it was updated can be made either way. Judge activity by looking at the commit history yourself rather than by any summary.

Upgrade cost is concentrated in the Pathway dependency and in Python version support. The declared range is ">=3.10,<3.13", so Python 3.13 is excluded at version 0.3.6. The dependency is pathway = "^0.12.0", which permits any 0.12.x release. Because the templates are copied rather than installed, a Pathway upgrade can change behaviour inside a template you have already modified, and there is no version pin at the template level to protect you.

The licence is MIT, declared in both LICENSE and the pyproject.toml license field. MIT is permissive and permits commercial use and modification. That covers this repository, not the third-party services the templates call: OpenAI, TwelveLabs, Mistral and Ollama each have their own terms, and the README does not discuss them. This is a description of what the repository states, not legal advice.

Editorial conclusion

Adopt it if you need a RAG service whose index follows additions, deletions and updates in Sharepoint, Google Drive, S3, Kafka or PostgreSQL, and you are willing to build on the Pathway framework: the pyproject.toml pins pathway = "^0.12.0" and Python ">=3.10,<3.13". Do not adopt it if you want a stable 1.x library or a drop-in replacement for an existing Pinecone or Weaviate deployment, since the package metadata classifies the project as Development Status :: 3 - Alpha and package-mode = false means there is no installable llm-app package to import. Before committing, open the README.md inside the specific template directory you intend to run and confirm the data source connector, the model provider and the exposed HTTP API match your deployment.

Frequently asked questions

What is pathwaycom/llm-app?

It is a repository of ready-to-deploy LLM application templates for RAG and enterprise search, built on the Pathway Live Data Framework and kept in sync with sources such as Sharepoint, Google Drive, S3, Kafka and PostgreSQL. The pyproject.toml sets package-mode = false, so it is a template collection rather than an installable library.

How do I install pathwaycom/llm-app?

There is no single install command. The README states that each template under templates/ contains its own README.md with run instructions, and the root pyproject.toml declares the Python range ">=3.10,<3.13" and the dependency pathway = "^0.12.0" that the templates rely on.

Can pathwaycom/llm-app run without a separate vector database?

Yes. The README says indexing is built in and done in-memory with cache, using the usearch library for the default vector index and Tantivy for hybrid full-text indexes, so a separate vector database, cache and API framework are not required.

Is pathwaycom/llm-app production ready?

The pyproject.toml classifies the project as Development Status :: 3 - Alpha at version 0.3.6, and the Pathway dependency is on a 0.x series. The README does not make a production-readiness claim, so treat the alpha classifier as the project's own statement about maturity.

What is the licence for pathwaycom/llm-app?

The repository is MIT licensed, stated in the LICENSE file and in the pyproject.toml license field. That covers the repository itself; the third-party model and embedding providers the templates call have separate terms that the README does not cover.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/pathwaycom-llm-app.svg)](https://hysenlabs.com/projects/pathwaycom-llm-app)
Community notes

Community notes