# pdftochat: a Next.js RAG template where the retrieval layer is the interesting part

> Nutlope's pdftochat puts a chat box on a PDF using LangChain.js, Mixtral through Together AI and Chroma Cloud hybrid search. It reads as a demo template, and its own task list says so more honestly than the README does.

**Nutlope/pdftochat** — Chat with your PDFs with AI

- Repository: https://github.com/Nutlope/pdftochat
- Website: https://www.pdftochat.com/
- Stars: 1,441 · Forks: 300
- Language: TypeScript
- License: MIT
- Published: 2026-10-07 · Updated: 2026-10-07 · Language: en
- Canonical page: https://hysenlabs.com/projects/nutlope-pdftochat

## Six services before the first upload

The tech stack section lists eight technologies, and five of them are paid services you must sign up for before anything renders. Together AI supplies the model through Mixtral. Chroma Cloud handles vector search. Bytescale stores the uploaded PDFs. Clerk handles authentication. Vercel hosts the app and provides Postgres, with Neon named as an alternative. LangSmith is optional and only for tracing.

That is the real adoption cost of this repository, and the README states it as a bulleted checklist rather than burying it. What you are signing up for is a reference implementation of a shape, not an application. The `.env.example` file makes the same point in a more honest form, since most of it is empty keys waiting for values: `TOGETHER_AI_API_KEY`, `NEXT_PUBLIC_BYTESCALE_API_KEY`, `CLERK_SECRET_KEY`, and the three Postgres connection strings.

One detail in that file deserves a flag rather than a mention. The Bytescale key is named `NEXT_PUBLIC_BYTESCALE_API_KEY`, and the `NEXT_PUBLIC_` prefix is the Next.js convention for variables exposed to the browser bundle. A storage credential carrying that prefix is worth understanding before you deploy, because the README does not explain what the client is expected to do with it and the package manifest includes both `uploader` and `react-uploader`, which are Bytescale's browser and React packages. If you fork this, verify where that key is used rather than assuming the naming is a mistake.

## Chroma Cloud does the embedding work for you

This is the part of the project that has moved on from the usual tutorial, and it is worth understanding before reading any other code. The README states that embeddings are generated automatically by Chroma Cloud using Qwen for dense vectors and SPLADE for sparse vectors, combined at query time with Reciprocal Rank Fusion. No local embedding model is required.

That matters more than it first appears. A conventional RAG demo makes you pick an embedding model, run it yourself, and store vectors in whatever database you configured. Here the vector database is also the embedding service, and it does both a dense and a sparse pass over the same document. Hybrid retrieval with RRF is what makes keyword-matching failures less likely on a document full of exact terms, which for PDFs of contracts and specs is the usual failure mode.

The setup is four environment variables and one sign-up:

```ini
NEXT_PUBLIC_VECTORSTORE=chroma

CHROMA_API_KEY=       # Your Chroma Cloud API key
CHROMA_TENANT=        # Your tenant ID
CHROMA_DATABASE=      # Your database name
```

Collections are created per document automatically, with no manual index setup. The `package.json` agrees with the README here, listing `@chroma-core/default-embed`, `@chroma-core/chroma-cloud-qwen` and `@chroma-core/chroma-cloud-splade` as separate dependencies, which is a reasonable confirmation that both embedding paths are wired rather than aspirational.

A caution about that file. `.env.example` still carries the full set of alternative vector store credentials, `PINECONE_API_KEY`, `PINECONE_INDEX_NAME`, `MONGODB_ATLAS_URI` and four more MongoDB fields, and `NEXT_PUBLIC_VECTORSTORE` is commented as accepting pinecone, mongodb or chroma. The README documents only the Chroma path. The Pinecone and MongoDB LangChain packages remain in the dependencies, so the code likely still supports them, but that path is undocumented and untested by the write-up.

## Getting the Postgres schema to exist

One command, and it is the step most likely to be missed:

```bash
npx prisma db push
```

That creates the `Document` table in Postgres, which the README names explicitly in its troubleshooting list. The `prisma/` directory is at the top level of the tree, and `package.json` wires the ORM into the build with a `vercel-build` script that runs `prisma generate` before `next build`, plus a separate `prisma:generate` script.

The troubleshooting section is short and specific, which is a good sign about the project's honesty. Four things to check: that your `.env` holds working API keys, that `NEXT_PUBLIC_VECTORSTORE` is `chroma` and your Chroma credentials are correct, that you have run `npx prisma db push`, and that Together AI has a credit card on file if you are hitting rate limits. That last one is not a bug, it is the free tier.

Everything else is Next.js convention: `app/` for the App Router routes, `components/` for the interface, `utils/` for the RAG logic, `middleware.ts` for route handling, `tailwind.config.js` and `styles/` for styling, `postcss.config.js` for the build. Local development is `next dev`, and `package.json` pins Next.js at `^14.2.0`.

## The future task list is the honest documentation

Most READMEs of this kind stop at the demo. This one ships a task list of nineteen unchecked items, and reading it tells you more about the state of the code than the feature description does.

Three entries are about missing functionality a user would assume is present. There is a task to add a trash icon and implement delete functionality, which means uploaded PDFs cannot currently be removed. There is a task to protect API routes by making sure users are signed in before executing chats, which means the chat endpoint is not gated behind authentication even though Clerk is wired up. And there is a task to save chats for each user in the Postgres database, which means conversation history does not persist.

Two more are about quality nobody has measured. One asks for an initial benchmark on how accurate chunking and retrieval are, and another asks to research best practices for chunking and retrieval and run benchmarks. Nothing in the repository shows a retrieval evaluation, and there is no test directory in the tree at all.

The rest is polish: sample questions on the page, markdown output options, better auto scrolling, SWR revalidation, toasts on failure, clickable sources, analytics, a message suggesting people compress PDFs beyond 10MB. That last item is the only size constraint mentioned anywhere, and it is a planned message rather than an enforced limit.

One inconsistency is worth flagging. The task list includes an item to upgrade to Next.js 14, while `package.json` already depends on `^14.2.0`. Either the task is stale or it refers to work that has since been done and never ticked. Also worth noting: `package.json` defines an `evals` script pointing at `./scripts/runEvaluation.ts`, but there is no `scripts/` directory in the repository tree, so that command would fail as written.

## Where it sits against a hosted PDF chat product

A reader's honest alternative is not another template, it is a hosted service. There are commercial PDF chat products that do exactly this with none of the setup, and for reading a document once and asking three questions, they are the better buy. The template earns its keep in the other direction: as a place to see how a RAG pipeline is wired when you intend to change the wiring.

The specific things worth copying are the hybrid retrieval setup through Chroma Cloud, since having the vector database own both Qwen dense and SPLADE sparse embeddings removes a whole service from a typical deployment, and the RRF combination at query time. Those two decisions are the reason the README frames Chroma as the approach rather than an option.

The specific reasons not to build on it as-is are the three missing features from the task list. Delete, authentication on the chat route, and persisted history are not enhancements, they are the minimum for handling anyone's documents. Add to that five services to configure, no tests, no retrieval benchmark, and a package manifest that has drifted since the last commit on 2026-06-26, and the picture is clear: this is a well-chosen stack assembled into a demo, not a maintained application.

Licensing is the easy part. MIT, stated in both the repository metadata and the `license` field in `package.json`, with no CLA in the tree and no `LICENSE-STATUS` file to complicate it. Forking is unrestricted; the work you are signing up for is making it safe.

## Conclusion

pdftochat is a good read for the shape of a retrieval pipeline and a poor foundation to build a product on without work. What you get is a working upload, chunk, embed and answer loop in about a dozen dependencies, with Chroma Cloud handling both dense and sparse embeddings so no local model is needed. What you do not get is persistence, deletion, protected chat routes, tested retrieval quality, or an evaluation harness, since the repo has no tests and its `evals` script points at a `scripts/` directory the tree does not contain. Running it needs five separate services configured before the first upload. Clone it, set `NEXT_PUBLIC_VECTORSTORE=chroma` with your three Chroma credentials, run `npx prisma db push`, and read `utils/` before you decide whether the RAG wiring is worth keeping or whether you would rather write against LangChain directly.

## FAQ

### Is PDF AI free?

The code is free, the services are not. pdftochat is MIT licensed, but it calls Mixtral through Together AI, Chroma Cloud, Bytescale and Clerk, and the README's troubleshooting section explicitly tells you to add a credit card to Together AI if free tier rate limits are a problem. Deployed on Vercel's free tier, the paid dependencies still apply.

### Can I upload a PDF and ask questions directly?

Yes, that is the core of the app. You upload a PDF through a Bytescale-backed uploader, and questions go to Mixtral through Together AI with context retrieved from Chroma Cloud. The catch is that uploaded files cannot be deleted from the dashboard yet, and the chat API route is not yet gated behind sign-in, both of which are open items in the repository's task list.

### Do I need a local embedding model to run pdftochat?

No. Chroma Cloud generates the embeddings for you, using Qwen for dense vectors and SPLADE for sparse ones, combined at query time with Reciprocal Rank Fusion. You supply an API key, a tenant ID and a database name, and collections are created per document automatically, so there is no local model to download or serve.

## Sources

- [Issues](https://github.com/Nutlope/pdftochat/issues)
- [License: MIT](https://github.com/Nutlope/pdftochat/blob/main/LICENSE)
- [Nutlope/pdftochat on GitHub](https://github.com/Nutlope/pdftochat)
- [Project website](https://www.pdftochat.com/)
- [README](https://github.com/Nutlope/pdftochat/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/nutlope-pdftochat
