Library / SDK
tonykipkemboi/ollama_pdf_rag avatar
tonykipkemboi/ollama_pdf_rag

ollama_pdf_rag: a local RAG stack with a Next.js UI and FastAPI backend

A full-stack demo showcasing a local RAG (Retrieval Augmented Generation) pipeline to chat with your PDFs.

544 stars198 forksTypeScriptMIT

At a glance

What is it?
A demo-grade but structured way to chat with PDFs entirely on your own machine, using Ollama for models, LangChain for the pipeline, and ChromaDB for storage. The Next.js plus FastAPI path is the one the repository recommends, and it is also the one with the most moving parts.
Who is it for?
Adopt it if you want a working local RAG pipeline you can read end to end, and you are comfortable running Python, Node and Ollama on the same machine. Skip it if you need a hardened multi-user service or a managed document store, because the PDFs and vectors live in local folders and there is no authentication layer described anywhere in the README.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 167 days ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What ollama_pdf_rag actually solves, and for whom

The problem is narrow and specific. You have PDFs you cannot send to a hosted API, either because of confidentiality or because you do not want to pay per token, and you want question answering over them with citations back to the source text. The README states the pitch directly: "100% Local - All processing happens on your machine, no data leaves." That is the whole value proposition, and everything else in the repository exists to make that usable.

The intended reader is not a production platform team. It is someone learning how retrieval augmented generation fits together, or a developer who wants a reference implementation to fork. The repository ships three interfaces for three levels of patience: a Next.js web app described as the primary one, a Streamlit app, and Jupyter notebooks under notebooks/experiments/ for people who want to see the pipeline without a UI in the way. The presence of a video tutorial and a MkDocs config (mkdocs.yml) points at the same audience.

If you already have a hosted RAG service and are happy with it, this project has nothing to offer you. Its reason to exist is the constraint that nothing leaves the machine.

The pipeline: pdfplumber in, ChromaDB stored, Ollama answering

The repository layout in the README shows the core split clearly. src/core/ holds document.py for PDF processing, embeddings.py for vector embeddings, llm.py for LLM configuration, and rag.py for the pipeline itself. src/api/ holds the FastAPI service with routers/ and services/ subfolders, and src/app/ holds the Streamlit application. The Next.js frontend lives separately in web-ui/, with its own app/, components/ and lib/ directories.

Data flow, as far as the documentation describes it: a PDF is uploaded, processed into text, split into chunks, embedded, and written into a ChromaDB collection under data/vectors/. At query time the question is embedded, matching chunks are retrieved, and the chat model produces an answer with source citations. The README advertises "Multi-Query RAG - Intelligent retrieval with source citations" and lists ChromaDB as the vector store, with langchain-chroma in requirements.txt. The Next.js interface is described as showing retrieved chunks and reasoning steps alongside the answer, which is the useful part for debugging retrieval quality.

Two details in requirements.txt are worth noting because they shape the install. The stack pins langchain==1.0.0 alongside langchain-core>=1.0.0, langchain-ollama==1.0.1, langchain_text_splitters==1.0.0 and langchain-classic==1.0.0. It also pulls unstructured[all-docs], which is a large dependency tree, and onnx. The ONNX entry explains the Windows troubleshooting note about a DLL load failure for onnx_copy2py_export, which the README says is fixed by installing the Microsoft Visual C++ Redistributable.

Installing ollama_pdf_rag and asking your first question

Prerequisites come first. Install Ollama from ollama.ai, then pull the two models the README names. The chat model and the embedding model are separate, and the embedding one is not optional.

bash
ollama pull llama3.2
ollama pull nomic-embed-text

Clone the repository and set up the Python environment. The README gives the venv activation line for both platforms.

bash
git clone https://github.com/tonykipkemboi/ollama_pdf_rag.git
cd ollama_pdf_rag
python -m venv venv
source venv/bin/activate
pip install -r requirements.txt

If you want the Next.js interface, install the frontend dependencies and run the database migration. The project uses pnpm, not npm, and the migration step is listed before you start anything.

bash
cd web-ui
pnpm install
pnpm db:migrate
cd ..

Now start both services. The FastAPI backend runs on port 8001 and the Next.js frontend on port 3000. The repository also provides a convenience script, start_all.sh, that starts everything at once.

bash
python run_api.py
cd web-ui && pnpm dev

With both running, open http://localhost:3000, upload a PDF through the attachment button or by dragging a file in, pick a model from the list of locally available Ollama models, and ask a question. The sidebar shows uploaded PDFs with chunk counts; the answer comes back with citations and the retrieved chunks. Swagger documentation for the API is at http://localhost:8001/docs.

There is also a .env.example at the repository root with three settings: OLLAMA_HOST=http://localhost:11434, VECTORDB_PATH=data/vectors, and LOG_LEVEL=INFO. If your Ollama instance is not on the default host, that is the value to change.

If you prefer the simpler path, python run.py starts the Streamlit interface on port 8501 instead, and the README lists that as its own option rather than a fallback.

Where the local-first design starts to hurt

The troubleshooting section is the honest part of the README, and it doubles as a list of failure modes. "No chunks retrieved" is answered with "Re-upload PDFs to rebuild the vector database." That tells you the vector store is not treated as durable state you can trust across changes. If you swap embedding models, change chunking, or move the data/vectors directory, the sensible move is to rebuild rather than assume the old collection still matches.

The CPU-only note is the second constraint. The README says to reduce chunk_size to 500-1000 in src/core/document.py if you hit memory issues on a CPU-only system. That is a code edit, not a configuration toggle, which is a sign of the project's demo origins. There is no documented environment variable for chunk size.

The third issue is dependency weight. Pulling unstructured[all-docs] and onnx to parse PDFs is a lot of surface area for what is fundamentally text extraction, and it is where the Windows DLL error comes from. A leaner deployment would use a narrower parser, but the repository does not offer that as an option.

Finally, the README does not document authentication, multi-user separation, or rollback of an indexed document set. The API endpoints are open on localhost. That is fine for a laptop and wrong for anything exposed to a network.

How this differs from a generic LangChain RAG example

The obvious alternative is a bare LangChain notebook: chain a PDF loader into a splitter, an embedding model and a vector store, then call the model. That is roughly what notebooks/experiments/updated_rag_notebook.ipynb contains, and it is genuinely less work if you only want to understand the mechanics. The difference is that a notebook gives you no upload path, no persistent PDF list, and no API. You re-run cells to change a document.

The second alternative is a hosted document-chat product. Those handle storage, scaling and access control, and they do it without you maintaining Python and Node toolchains. The trade is exactly the one this project refuses: your documents go to someone else's servers. If that is acceptable to you, ollama_pdf_rag's main selling point evaporates and its operational overhead becomes pure cost.

The third comparison is against building the FastAPI service yourself on top of LangChain. That is a real option, and the repository is essentially a worked answer to it. What you inherit by using this project instead is the endpoint set: POST /api/v1/pdfs/upload, GET /api/v1/pdfs, DELETE /api/v1/pdfs/{pdf_id}, POST /api/v1/query, GET /api/v1/models, and GET /api/v1/health, plus a Next.js UI already wired to them. Whether that is worth the extra dependencies depends on how much of the UI you would otherwise write.

Maintenance, licensing and what the release history shows

The repository is not archived, and the last push was on 2026-04-16. That is roughly five months before the date of this article, so treat it as a project that has shipped in bursts rather than one with continuous activity. The release history supports that reading: v2.0 in November 2024, v2.1.0 in January 2025, then a long gap until v3.0.0 in December 2025, which the release notes describe as "Next.js UI, FastAPI Backend & Enhanced RAG." The v3 line is the architecture the README now presents as primary, so anything you find written about the earlier Streamlit-only version describes a different shape of the project.

Upgrade cost is dominated by the pinned dependencies. langchain==1.0.0, langchain_text_splitters==1.0.0, langchain-classic==1.0.0 and pydantic==2.10.4 are exact pins, while numpy==1.26.4 and Pillow==10.4.0 are also fixed. Moving any one of those forward means re-testing the whole chain, and the repository does include a test suite (python -m pytest tests/ -v, with coverage via --cov=src) plus pre-commit hooks, which is more than most demos offer. The GitHub Actions badge in the README points at a Python test workflow.

Licensing is MIT, which is permissive and imposes no obligation on how you deploy a modified version. That is a statement about the licence text, not advice about your situation; if you plan to redistribute, read the LICENSE file in the repository root yourself.

Editorial conclusion

Adopt it if you want a working local RAG pipeline you can read end to end, and you are comfortable running Python, Node and Ollama on the same machine. Skip it if you need a hardened multi-user service or a managed document store, because the PDFs and vectors live in local folders and there is no authentication layer described anywhere in the README. Before you commit, verify two things yourself: that your Ollama install serves both a chat model and nomic-embed-text, and that pnpm db:migrate completes in web-ui, since the SQLite database it creates is what the FastAPI backend reads from.

Frequently asked questions

How do I create a RAG pipeline from a PDF with ollama_pdf_rag?

Install Ollama, pull llama3.2 and nomic-embed-text, then install the Python requirements and the web-ui dependencies with pnpm install followed by pnpm db:migrate. Start the backend with python run_api.py and the frontend with pnpm dev, then upload a PDF through the Next.js interface and ask a question.

Can a local LLM read PDFs in ollama_pdf_rag?

The project processes PDFs locally with pdfplumber and unstructured, splits them into chunks, and stores embeddings in ChromaDB under data/vectors. The model itself answers from retrieved chunks rather than reading the whole document, and the README states that all processing happens on your machine.

How can I build a local RAG system with ollama_pdf_rag?

The README describes three interfaces for the same pipeline: the Next.js app with the FastAPI backend on ports 3000 and 8001, a Streamlit app started with python run.py on port 8501, and Jupyter notebooks under notebooks/experiments/. All three use the same core RAG code in src/core/.

Is Ollama a library used by ollama_pdf_rag?

Ollama is a separate application you install and run, and this project talks to it over HTTP at OLLAMA_HOST, which the .env.example sets to http://localhost:11434. The Python side depends on the ollama package and langchain-ollama, but the models themselves are served by the Ollama process, not by the Python code.

Official sources

  1. License: MIT
  2. Project website
  3. README
  4. Releases
  5. tonykipkemboi/ollama_pdf_rag on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/tonykipkemboi-ollama-pdf-rag.svg)](https://hysenlabs.com/projects/tonykipkemboi-ollama-pdf-rag)