# danny-avila/rag_api: an ID-based RAG FastAPI service for LibreChat

> rag_api is a FastAPI and Langchain service that stores embeddings in PostgreSQL with pgvector and keys every chunk to a file or entity id. Its latest release reworks authorization so callers only read and delete what they own, which changes upgrade order and delete semantics.

**danny-avila/rag_api** — ID-based RAG FastAPI: Integration with Langchain and PostgreSQL/pgvector

- Repository: https://github.com/danny-avila/rag_api
- Website: https://librechat.ai/
- Stars: 909 · Forks: 407
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/danny-avila-rag-api

## The file-id problem rag_api was built around

Most retrieval stacks assume you search one flat pool of chunks and filter afterwards. rag_api takes the opposite position: files are organized into embeddings by file_id, and the primary use case is integration with LibreChat. That matters because LibreChat already stores file metadata in its own database, so the API does not need to invent a document model. It needs to answer questions about a known set of files. The README states the main reason for the ID approach is to work with embeddings on a file level, which makes for targeted queries when combined with file metadata stored in a database.

The audience is therefore narrower than the topic list suggests. If you are building a product where users upload documents and then chat over them, and you already have a table mapping uploads to owners, this service slots in underneath. If you want semantic search across an entire corpus with no notion of who owns what, the ID model is overhead, not help.

## How chunks are stored, scoped and queried

The stack is Langchain for the vector store abstraction, FastAPI for the HTTP surface, and PostgreSQL with pgvector for storage. Embeddings live in the langchain_pg_embedding table, with chunk metadata carried in a cmetadata JSONB column. That column is where ownership is recorded: the README's upgrade SQL reads and writes cmetadata->>'user_id'.

The v0.9.0 release moved authorization into the query itself. Every route that reads or removes stored content resolves the caller's owner set from the verified token and puts it into the store query before ranking, so a chunk outside that set is never read into the process. The owner set is built in one place, app/scope.py, rather than re-derived per route. That is a real design decision: a single resolution point is easier to audit than per-route checks, and it means a new route cannot accidentally skip scoping by forgetting to call a helper.

The release notes are unusually candid about what came before. GET /ids listed every file id in the deployment. POST /query_multiple performed no authorization at all, so pairing it with GET /ids disclosed the content of every file to any authenticated caller. POST /query authorized the whole result set from documents[0], so any hit behind the first was never checked. Because a file_id is chosen by whoever uploads, an attacker's own row ranking first authorized the rows behind it. GET /documents, GET /documents/{id}/context and DELETE /documents read or deleted the chunks of any file id the caller could name. On the synchronous store path, a failed ingestion rolled back by file_id alone, so an upload under someone else's file id destroyed their chunks. The async pgvector pipeline already scoped its rollback to the ingestion attempt.

One behavioural detail is worth internalizing: a file id outside the caller's scope answers not found rather than found but refused, so none of these routes is an existence oracle. That is the right call for disclosure, but it makes 404 ambiguous.

## Installing rag_api and running a first query

The README gives two paths. Docker is the short one: docker compose up starts both PSQL/pgvector and the RAG API. If you want only the database, use the db-compose.yaml variant; if you want only the API against an existing database, use api-compose.yaml. The compose file builds from the repository Dockerfile, which is python:3.12, installs pandoc and netcat-openbsd, installs requirements.txt, and downloads NLTK data at build time so unstructured does not fetch packages at runtime. It also sets SCARF_NO_ANALYTICS=true to disable Unstructured analytics.

The bundled database service is ankane/pgvector:latest, with POSTGRES_DB mydatabase, POSTGRES_USER myuser and POSTGRES_PASSWORD mypassword, and it publishes 5433 on the host mapped to 5432 in the container. The API publishes 8000. Note that the compose file passes DB_HOST=db and DB_PORT=5432 into the fastapi service, so the API talks to the container port, not the host port.

```bash
docker compose up
```

That single command brings up the database and the API together. The README also documents running the API locally, in a virtual environment, after pointing DB_HOST at your database:

```bash
pip install -r requirements.txt
uvicorn main:app
```

Before any of this works you need a .env file, and the README points to its environment variables section rather than listing every key inline. The one key the upgrade notes discuss by name is JWT_SECRET. With no signing key configured there is no caller identity to record, so chunks written in that period are owned by the literal string public. Set it before you ingest anything you care about, or you will be writing rows that later become unreadable.

The delete route has a parameter you should know about before you build a client. Chunks embedded under an entity_id, an agent knowledge base for instance, are owned by that entity rather than by the uploading user, so DELETE /documents needs the same entity_id that the upload used, as a query parameter alongside the JSON body of file ids. A delete that omits it resolves to the caller's own scope, matches nothing, and answers 404 with the chunks left in place.

## The upgrade order trap and the public-owner migration

This is the part of the release that will bite people who skim. Deploy order matters in one direction only. A client that sends entity_id against an older build is inert, because the parameter is simply undeclared there, so the request behaves exactly as before. An older client against this build orphans every agent knowledge-base file it tries to delete. The README's instruction is to upgrade the client first, or both together, never this service first. LibreChat carries the matching change: it records the owner each embed was made under and sends it on delete, with npm run migrate:embed-owners to backfill files embedded before that.

If your deployment holds chunks with no user_id, they are owned by nobody and are no longer readable. The README provides a stamping statement for that case:

```sql
UPDATE langchain_pg_embedding
SET cmetadata = jsonb_set(cmetadata, '{user_id}', '"<owner>"')
WHERE cmetadata->>'user_id' IS NULL;
```

If the deployment ever ran without JWT_SECRET, the README says to check for public too, and gives an inspection query rather than a blind rewrite, which is the correct instinct because that content was never identified as belonging to anyone:

```sql
-- inspect before rewriting: this is content nobody was ever identified as owning
SELECT count(*) FROM langchain_pg_embedding WHERE cmetadata->>'user_id' = 'public';
```

Deployments that never set JWT_SECRET are unaffected by this migration: with no key configured the read scope is public as well, so what was written is what is read. Atlas MongoDB deployments have a separate prerequisite: they must add user_id to the vector search index first.

## Where rag_api is the wrong tool: entity_id is still asserted by the caller

The honest limitation is stated plainly in the README, and it is not a small one. Agent knowledge bases are owned by an agent id rather than a user id, so a caller reading one names it via entity_id. That id now widens the owner set rather than replacing the caller's identity, and the caller's own scope always remains. But nothing in a token minted today proves the caller may act for the entity it names. A caller that knows another owner's id can still name it, on read to reach that owner's chunks, and on the ingestion routes, where entity_id is what gets stamped as the owner, to write into that owner's namespace.

The README's conclusion is direct: deployments exposing this API to untrusted callers must continue to authorize entity access upstream, and closing the gap requires the token to carry entity authorization, which is a coordinated change with the callers that mint those tokens and is tracked separately from this release. If you were hoping v0.9.0 made the service safe to put on the open internet, it did not. It closed the file_id holes. It did not close entity impersonation, and it says so.

The second limitation is operational rather than security-shaped. Because a missing entity_id on delete returns 404 with the chunks left in place, and a 404 is indistinguishable from already deleted, a caller that treats 404 as success will orphan those chunks silently. There is no error that distinguishes the two cases. That is a design consequence of the not-found-rather-than-refused rule, and it means your client has to be correct about entity_id rather than relying on the server to tell it when it is wrong.

## How it compares with a general-purpose vector database service

The obvious alternative is to run a vector store directly and skip the API layer: pgvector with your own SQL, or a hosted vector database with its own SDK. The difference is where identity lives. With a raw pgvector table you write the WHERE clause yourself on every query, and nothing forces you to put the owner filter before ranking. rag_api's contribution is that the filter is resolved once in app/scope.py and injected into the store query before ranking, so the default is scoped rather than unscoped.

A second alternative is a full RAG framework with an HTTP server attached. Those tend to model documents as first-class objects with their own ids and collections. rag_api deliberately does not. It keys on file_id and entity_id, which are ids that already exist in the caller's world. That is less flexible and considerably less work to integrate if your system already speaks those ids. If it does not, you will spend your time mapping your identifiers onto theirs.

The dependency list is worth reading before you commit. requirements.txt pins langchain 1.3.10, langchain-community 0.4.2, langchain-openai 1.3.2, langchain-core 1.4.8, fastapi 0.115.12, pgvector 0.2.5, sqlalchemy 2.0.41 and asyncpg 0.29.0, among others. It also pulls in a wide set of document parsers: pypdf, unstructured, docx2txt, pypandoc, openpyxl, xlrd, python-pptx, msoffcrypto-tool, and rapidocr-onnxruntime with opencv-python-headless. There are provider integrations for AWS, Ollama, HuggingFace and Google GenAI, plus sentence_transformers and langchain-mongodb. That is a heavy image. The repository does ship a requirements.lite.txt and a Dockerfile.lite, so a smaller build exists, but the README excerpt does not document what the lite variant omits.

## Licence, maintenance and what an upgrade costs

The licence is MIT, which places few restrictions on reuse and modification. That is a permissive licence, not a statement about the project's security posture or support commitments, and nothing here is legal advice for your particular distribution model.

On maintenance: the repository is not archived, and the last push was on 2026-08-15. Releases are tagged regularly, with v0.9.0 on 2026-07-31, v0.8.0 on 2026-04-21 and v0.7.3 on 2026-03-20. The README states the API will evolve over time to employ different querying and re-ranking methods, embedding models and vector stores, which is a fair warning that the surface is not frozen.

The upgrade cost for v0.9.0 is concentrated in three places. Your client must send entity_id on delete before the service is upgraded, or agent knowledge-base files get orphaned. Your database may need a one-time cmetadata update for chunks with no user_id, and possibly a separate review for rows owned by public. And your Atlas MongoDB vector search index needs user_id added if you use that backend. Everything else about the request shapes is unchanged, and a client sending entity_id to an older build is simply inert.

## Conclusion

Adopt rag_api if you run LibreChat or another system that already keys documents by file id and you want retrieval without building a vector pipeline. Do not adopt it as a general-purpose multi-tenant retrieval service if you expose it to untrusted callers, because entity_id remains caller-asserted and the README says entity authorization must be handled upstream. Before upgrading, verify whether your deployment ever ran without JWT_SECRET, since chunks written then are owned by the literal string public and stop being readable once a key is set, and check that your client sends entity_id before the service is upgraded.

## FAQ

### What is rag_api used for?

It is a FastAPI service that indexes documents into PostgreSQL with pgvector and serves retrieval queries, with files organized into embeddings by file_id. The README names LibreChat as the primary use case, while noting the API can be used for any ID-based use case.

### How do I install rag_api?

Configure a .env file, then run docker compose up to start both PSQL/pgvector and the RAG API, or use api-compose.yaml to run only the API against an existing database. The README also documents a local path with pip install -r requirements.txt followed by uvicorn main:app.

### Why does deleting a file return 404 with rag_api?

If the chunks were embedded under an entity_id, DELETE /documents needs that same entity_id as a query parameter alongside the JSON body of file ids. A delete that omits it resolves to the caller's own scope, matches nothing, and answers 404 with the chunks left in place, which is indistinguishable from already deleted.

## Sources

- [danny-avila/rag_api on GitHub](https://github.com/danny-avila/rag_api)
- [License: MIT](https://github.com/danny-avila/rag_api/blob/main/LICENSE)
- [Project website](https://librechat.ai/)
- [README](https://github.com/danny-avila/rag_api/blob/main/README.md)
- [Releases](https://github.com/danny-avila/rag_api/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/danny-avila-rag-api
