# pdfGPT: Chat with a PDF Using OpenAI and Sentence Embeddings

> pdfGPT is an open-source Python tool that lets you ask questions about a PDF file by chunking its text, generating embeddings with a Deep Averaging Network encoder, and retrieving the most relevant chunks via KNN before sending them to an OpenAI model for an answer.

**bhaskatripathi/pdfGPT** — PDF GPT allows you to chat with the contents of your PDF file by using GPT capabilities. The most effective open source solution to turn your pdf files in a chatbot!

- Repository: https://github.com/bhaskatripathi/pdfGPT
- Website: https://huggingface.co/spaces/bhaskartripathi/pdfChatter
- Stars: 7,163 · Forks: 833
- Language: Python
- License: MIT
- Published: 2026-09-22 · Updated: 2026-09-22 · Language: en
- Canonical page: https://hysenlabs.com/projects/bhaskatripathi-pdfgpt

## The Problem pdfGPT Was Built to Solve

Asking an LLM a question about a specific PDF has two technical obstacles. First, a large PDF contains far more text than an LLM's context window can hold in a single request. Second, passing an entire document to OpenAI at once, even when it fits, causes the model to return vague or unfocused answers because the relevant information is buried.

pdfGPT addresses both issues with a retrieval approach that predates more complex orchestration frameworks. The author notes in the README that the first version was developed in 2021 as one of the earliest open-source retrieval-augmented generation solutions. The design goal was accuracy over a polished interface: the README states that despite many RAG solutions available since, pdfGPT remains competitive in response accuracy due to its simple and specific architecture.

The tool targets developers and researchers who want to interrogate a PDF document without sending the entire file to an external service and who prefer a lightweight implementation with no vector database or indexing layer.

## How pdfGPT Retrieves and Answers Questions

The architecture follows a sequence described in the README's flowchart. When a user uploads a PDF or provides a URL, pdfGPT parses the file to text using PyMuPDF, then splits the text into chunks of approximately 150 words. These chunks are kept together with their citation history (the page number they came from).

Each chunk is passed through a Deep Averaging Network encoder from the Universal Sentence Encoder family, loaded via TensorFlow Hub. This produces a dense vector embedding for each chunk. When a user asks a question, the question is embedded using the same model, and KNN (K-Nearest Neighbor) search finds the chunks whose embeddings are closest to the question embedding. The author cites Andrej Karpathy's observation that KNN is appropriate for this kind of retrieval problem.

The most relevant chunks, together with their page references, are assembled into a prompt and sent to the OpenAI API. The returned answer can include page numbers in square brackets, citing which pages contain the supporting information.

This approach avoids a vector database entirely: embeddings are computed and stored in memory (or in a local file if previously generated) and searched with scikit-learn's KNN implementation. The trade-off is that scaling to very large document sets is limited by memory.

The README includes a specific note about model choice: Turbo models such as gpt-3.5-turbo are chat completion models and perform poorly when embedding similarity is low. For precise answers, the recommendation is to use text-DaVinci-003 or GPT-4 and above.

## Installation and Running pdfGPT

The repository provides two deployment paths: Docker Compose and direct Python execution.

For Docker Compose, the stack has two services:

```yaml
version: '3'
services:
  langchain-serve:
    build:
      context: .
      target: langchain-serve-img
    ports:
      - '8080:8080'
  pdf-gpt:
    build:
      context: .
      target: pdf-gpt-img
    ports:
      - '7860:7860'
```

The pdf-gpt service builds from the pdf-gpt-img target in the Dockerfile:

```dockerfile
FROM python:3.8-slim-buster as pdf-gpt-img
WORKDIR /app
COPY requirements.txt requirements.txt
RUN pip3 install -r requirements.txt
COPY . .
CMD [ "python3", "app.py" ]
```

The application runs a Gradio interface on port 7860. A live demo is available at https://huggingface.co/spaces/bhaskartripathi/pdfChatter for testing without local installation.

The requirements.txt pins specific versions: PyMuPDF 1.22.1 for PDF parsing, numpy 1.23.5, scikit-learn 1.2.2 for KNN, tensorflow for the encoder, tensorflow_hub 0.13.0, openai 0.27.4, gradio 4.11.0, and langchain-serve for production API deployment. These pins were current in 2023 and may conflict with more recent package versions.

For running via langchain-serve (the langchain-serve-img target), the tool exposes a REST API at port 8080.

## What the Embeddings Actually Do

The choice of the Universal Sentence Encoder family (Deep Averaging Network encoder) instead of OpenAI's own embedding API is a deliberate design decision. The README argues that OpenAI embeddings produce poor results for certain Q&A tasks and that the USE encoder produces better semantic similarity for this retrieval use case.

The Deep Averaging Network works by averaging word vectors across the tokens in a sentence, then applying a deep network on top. It is a relatively compact model that runs locally, which means the embedding step does not require an API call or internet access. Once the embeddings file exists for a document, subsequent questions reuse it without re-processing the PDF.

The embedding and KNN approach has a specific limitation: it works well for factual questions where the answer is localized to one or two chunks. Questions that require reasoning across many parts of the document, or questions that ask for a synthesis of information spread throughout the file, receive answers only as good as the top-K retrieved chunks allow. If the relevant chunks are not in the top results, the answer will be incomplete or incorrect, and the model will not indicate what was missed.

## Limitations and Maintenance Status

The last push to this repository was on 2026-03-06. The author's own README states: "The library documentation that you see below is a bit outdated as I do not get enough time to maintain it."

The openai package is pinned at version 0.27.4. OpenAI migrated to a new client API (the openai Python library v1.0+) in late 2023, which is incompatible with the 0.27.4 call signatures used in this project. Running the code against a current OpenAI account with the pinned version may require API compatibility patches.

The README lists several planned features that have not been delivered: OCR support, multiple PDF file support, Falcon/Vicuna/Meta Llama integration, and a Node.js web application. These are listed under an upcoming release pipeline with no dates.

There is also no vectorDB or index: all embeddings are held in memory or a local file. For a document set larger than a few hundred pages, memory usage and KNN search time grow proportionally with the number of chunks. The 150-word chunk size and the absence of chunking configuration options limit how users can tune retrieval quality for their specific documents.

## pdfGPT versus ChatPDF

ChatPDF is a commercial web service that the README names as one of the prior solutions in this space. The functional difference is where the document lives: ChatPDF uploads your PDF to its servers, processes it there, and charges for usage beyond a free tier. pdfGPT runs locally, so the PDF file never leaves your machine. Your OpenAI API key is what makes the final generation call, but the document parsing and embedding happen on your hardware.

For users with sensitive documents (legal filings, medical records, confidential reports), the local processing model of pdfGPT is a meaningful distinction. ChatPDF offers a polished web interface and handles OCR, multiple documents, and maintained infrastructure. pdfGPT offers source code you can inspect and modify, at the cost of setup effort and ongoing maintenance.

A more current alternative in the open-source space is LlamaIndex, which provides a maintained retrieval framework that supports multiple document types, vector stores, and LLM providers. Unlike pdfGPT, LlamaIndex is actively maintained and supports the current OpenAI API.

## License and Citation

pdfGPT is licensed under the MIT License, which permits commercial use without restriction. The repository includes a CITATION section with a BibTeX entry for researchers who use the project in academic work.

The repository structure is minimal: Dockerfile, LICENSE.txt, README.md, api.py, app.py, docker-compose.yaml, and requirements.txt. The app.py file contains the Gradio interface and the main pipeline. The api.py file provides the langchain-serve API wrapper.

The author has noted openness to contributors for maintaining the project jointly. A companion project, LLM_Quantization, is linked from the README for users interested in quantizing large language models for consumer hardware.

## Conclusion

pdfGPT is a reasonable starting point for developers who want to understand how a minimal RAG pipeline works without third-party orchestration libraries. The architecture is simple enough to read in a day. It is not the right choice for production applications: the last push was on 2026-03-06, the openai package is pinned at version 0.27.4 (an outdated API format), and the author's own README notes that the library documentation is out of date. Teams building PDF chat into a product should treat this as a reference implementation, verify that the pinned dependencies still work with current OpenAI API versions, and plan to replace the embedding and retrieval components with maintained alternatives.

## FAQ

### Is PDF GPT free?

The pdfGPT software is MIT licensed and free to use. However, it requires your own OpenAI API key to generate answers, and OpenAI charges for API usage. The embedding step uses a local Universal Sentence Encoder model and does not require an API call.

### Is pdfGPT free?

The pdfGPT source code is free and open-source under the MIT license. Generating answers requires an OpenAI API key, which has usage-based costs. The Hugging Face Spaces demo at bhaskartripathi/pdfChatter can be used without a local installation.

### Does pdfGPT support PDFs that contain scanned images rather than text?

The README lists OCR support as an upcoming feature that has not yet been delivered. The current version uses PyMuPDF for text extraction, which works on PDFs with embedded text but not on scanned image-only PDFs.

## Sources

- [bhaskatripathi/pdfGPT on GitHub](https://github.com/bhaskatripathi/pdfGPT)
- [Issues](https://github.com/bhaskatripathi/pdfGPT/issues)
- [License: MIT](https://github.com/bhaskatripathi/pdfGPT/blob/main/LICENSE)
- [Project website](https://huggingface.co/spaces/bhaskartripathi/pdfChatter)
- [README](https://github.com/bhaskatripathi/pdfGPT/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/bhaskatripathi-pdfgpt
