ray-project/llm-applications: A RAG Reference Implementation Built on Ray
A comprehensive guide to building RAG-based LLM applications for production.
At a glance
- What is it?
- The repository is a notebook-driven guide to building, evaluating and serving a retrieval augmented generation pipeline, with Ray handling scale and Postgres with pgvector handling retrieval. The code is real, the setup assumes Anyscale or comparable GPU infrastructure, and the last push was on 2026-08-15.
- Who is it for?
- Adopt this repository if you already run Ray or have access to a g3.8xlarge-class machine and want a working RAG pipeline to modify rather than a library to import. Skip it if you need a pip-installable package with a stable API, or if you cannot supply OpenAI and Anyscale credentials plus a Postgres instance with pgvector.
- Can I use it commercially?
- Yes, with credit. CC-BY-4.0 allows commercial use as long as you credit the authors and indicate what you changed. It is written for creative content, so check how it applies to any code.
- Is it still maintained?
- Yes. The repository last received commits 46 days ago.
- What is it written in?
- Mainly Jupyter Notebook, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What ray-project/llm-applications actually is
This is not a library you install and import. It is a guide repository: a long interactive notebook, notebooks/rag.ipynb, plus supporting Python under rag/, deploy/, experiments/ and migrations/, that walks through building a retrieval augmented generation application end to end. The README frames the scope explicitly, listing development of a RAG application from scratch, scaling of its components (load, chunk, embed, index, serve), evaluation of configurations against per-component and overall scores, hybrid routing between open and closed models, and serving. The audience is an engineer who wants to see the whole pipeline assembled and then change it, not someone looking for a drop-in RAG framework. The primary language is Jupyter Notebook, which tells you where the narrative lives: the prose, the decisions and the ordering are in the notebook, and the Python files are the pieces it calls.
The pipeline: load, chunk, embed, index, serve
The architecture follows the component list in the README. Documents are loaded and chunked, chunks are embedded, embeddings go into a vector index, and a serving layer answers queries by retrieving chunks and passing them to an LLM. Retrieval is backed by Postgres with pgvector: requirements.txt pulls in asyncpg, pgvector, psycopg with binary and pool extras, psycopg2-binary and sqlalchemy[asyncio], and the repo ships setup-pgvector.sh and a migrations/ directory. That is a deliberate choice with a real consequence: you need a running Postgres instance before any of the interesting code works, and the DB_CONNECTION_STRING in the README points at localhost by default. Ray sits underneath the expensive stages so that embedding and indexing can be distributed across workers, which is why the README recommends a g3.8xlarge head node with 2 GPUs and 32 CPUs and mentions adding GPU worker nodes. Evaluation is treated as a first-class stage rather than an afterthought, with experiments/ holding configuration runs and metrics named in the README such as retrieval_score and quality_score. The hybrid routing piece is the most opinionated part: rather than committing to one model, the guide shows routing between OSS models served through Anyscale Endpoints (the README names Llama-2-70b) and OpenAI models such as gpt-3.5-turbo and gpt-4. Two providers means two sets of credentials and two failure modes in the same request path.
Installing it and running the first query
Clone the repository, then install the dependency set. The README uses pip with the --user flag and also sets PYTHONPATH so that the rag/ package is importable from the repository root.
git clone https://github.com/ray-project/llm-applications.git .
pip install --user -r requirements.txt
export PYTHONPATH=$PYTHONPATH:$PWDCredentials go in a .env file at the repository root. The README lists four variables plus the Postgres connection string, and each key has a comment pointing at where to obtain it.
touch .env
# Add environment variables to .env
OPENAI_API_BASE="https://api.openai.com/v1"
OPENAI_API_KEY=""
ANYSCALE_API_BASE="https://api.endpoints.anyscale.com/v1"
ANYSCALE_API_KEY=""
DB_CONNECTION_STRING="dbname=postgres user=postgres host=localhost password=postgres"
source .envThe README also runs pre-commit install and pre-commit autoupdate as part of environment setup, which matters if you plan to contribute rather than just read. After that, the README's next instruction is to open notebooks/rag.ipynb, and the repository provides update-index.sh and test.py at the top level for index refreshes and checks. What you should see is a notebook that walks the stages in order and, once the vector index is populated, returns retrieved chunks and a generated answer for a query. The README does not document a single command that runs the whole pipeline headlessly, so expect to execute cells rather than invoke a CLI.
Where the setup will fight you
The README's compute section is honest about the target: a g3.8xlarge head node with 2 GPUs and 32 CPUs, the default_cluster_env_2.6.2_py39 cluster environment, and the us-west-2 region if you want the shared storage artifacts. The data path is hardcoded to /efs/shared_storage/goku/docs.ray.io/en/master/, and the README notes that if you load the data yourself, the output directory must be on shared storage so workers can reach it. That is an Anyscale-shaped assumption baked into the instructions. On a laptop, or on a cloud account without a shared filesystem mounted at that path, you will be rewriting paths before the first embedding runs. The second constraint is the two-provider credential requirement: OpenAI and Anyscale accounts are both needed for the full routing example, and the README gives no offline or single-provider fallback for the routing section. Third, the release history is thin and dated. The most recent release listed is v0.0.9 from 2023-09-17, described as "Update Ray Assistant Service", and the two before it are from August and September 2023. The repository itself received a push on 2026-08-15, so work continues on main, but the tagged releases do not reflect it. If you pin to a release tag, you are pinning to something from 2023. There is also no README documentation of rollback, of upgrade steps between versions, or of a supported Python version beyond the py39 in the cluster environment name and the py39plus setting in pyproject.toml.
How it differs from LangChain and LlamaIndex
LangChain and LlamaIndex are libraries: you add them as dependencies and compose their abstractions inside your own application. This repository inverts that relationship. It is an application you read and adapt, and it happens to depend on langchain (listed in requirements.txt) rather than competing with it. The practical difference shows up in three places. First, orchestration: here the pipeline is explicit notebook code with Ray distributing the heavy stages, so you can see and edit each step; in a framework the steps are hidden behind chain or index abstractions you configure. Second, retrieval storage: this guide commits to Postgres with pgvector via SQLAlchemy and asyncpg, which means your vector store is a database you already know how to operate; frameworks typically abstract the store behind an interface and leave the operational work to you anyway. Third, evaluation: experiments/ and the retrieval_score and quality_score metrics are part of the repository, so configuration comparison is a stage in the pipeline rather than a separate tool you bolt on. If you want a maintained abstraction layer with a stable API and broad integrations, a framework is the better fit. If you want to understand and control every stage, including where Ray parallelizes, this repository is the more direct teacher.
Licence and maintenance cost
The repository is licensed CC-BY-4.0, which is a content licence rather than a software licence. That distinction matters for adoption: CC-BY-4.0 is designed for creative works and requires attribution when you redistribute or adapt the material. It is not the Apache-2.0 or MIT licence that most engineering teams expect to see on a code repository, and it does not include the patent grant that Apache-2.0 carries. This is not legal advice, and the practical question of how CC-BY-4.0 applies to code you copy into a proprietary product is one for your own counsel. On maintenance: the last push was on 2026-08-15, so the repository is not archived and work on main is recent, but the release tags stop in September 2023 and the README still references Ray Summit 2024 and a staging console URL (console.anyscale-staging.com), which suggests parts of the documentation have not been revisited in a while. Budget for reading the notebook rather than the release notes when you upgrade, because the release notes do not cover the gap.
Editorial conclusion
Adopt this repository if you already run Ray or have access to a g3.8xlarge-class machine and want a working RAG pipeline to modify rather than a library to import. Skip it if you need a pip-installable package with a stable API, or if you cannot supply OpenAI and Anyscale credentials plus a Postgres instance with pgvector. Before committing, open notebooks/rag.ipynb and check whether the Anyscale-specific paths such as /efs/shared_storage/goku/docs.ray.io/en/master/ and the default_cluster_env_2.6.2_py39 environment match your infrastructure, because the README does not document a portable fallback for either.
Frequently asked questions
What are LLM apps?
In this repository the term covers an application that combines a language model with retrieval, evaluation and serving infrastructure. The README describes building a retrieval augmented generation application that loads, chunks, embeds and indexes documents, then serves answers from them.
What are examples of LLM?
The README names the models used in the guide: gpt-3.5-turbo and gpt-4 through OpenAI, and Llama-2-70b through Anyscale Endpoints. The hybrid routing section exists specifically to route between these open and closed options.
What are the top LLM apps?
The repository does not rank applications. It is a single guide for building one RAG application, and the README does not compare it against other LLM applications or list any.
Is ChatGPT LLM or nlp?
The repository does not address this distinction. It treats ChatGPT models as API endpoints accessed through OPENAI_API_BASE and OPENAI_API_KEY, and does not classify them against NLP terminology.
what is llm applications
The README describes the guide as covering development of a retrieval augmented generation application from scratch, scaling of its components, evaluation of configurations, hybrid routing between open and closed models, and serving it. That is the definition the repository works from.
what is building llm applications
According to the README, it means loading, chunking, embedding, indexing and serving documents, then evaluating the result with per-component metrics such as retrieval_score and an overall quality_score. The repository's notebooks/rag.ipynb walks through each of those stages.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/ray-project-llm-applications)