pguso/rag-from-scratch: Building a Local RAG Pipeline in JavaScript, One Example at a Time
Demystify RAG by building it from scratch. Local LLMs, no black boxes - real understanding of embeddings, vector search, retrieval, and context-augmented generation.
At a glance
- What is it?
- A teaching repository that walks through embeddings, vector search, retrieval and augmentation using node-llama-cpp instead of cloud APIs. It is a learning path, not a library to drop into production.
- Who is it for?
- Use pguso/rag-from-scratch if you want to read and run a RAG pipeline end to end in JavaScript before choosing a framework, and if you accept that the examples are numbered folders rather than a supported package. Do not adopt it as the retrieval layer of a product: the vector store example is in-memory and the README documents no persistence, no rollback and no upgrade path.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Activity is slowing. The repository last received commits 6 months ago.
- What is it written in?
- Mainly JavaScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Who pguso/rag-from-scratch is written for
The README states the goal plainly: demystify Retrieval-Augmented Generation by building it yourself, with local code and no cloud APIs. That framing tells you the audience. This is for a developer who already writes JavaScript and wants to see the machinery of retrieval, chunking and prompt augmentation without a framework hiding the steps. The repository follows the same philosophy as the author's earlier ai-agents-from-scratch project, which the README names as a sibling.
The problem it addresses is not that RAG is hard to call. It is that most introductions start from a framework import, so the reader never sees where a document becomes a vector or why a retrieved chunk was ranked above another. Here the pipeline is broken into numbered folders, each with an example.js, a CODE.md walkthrough and a CONCEPT.md explanation. If you learn by reading short files in order, that layout is the point. If you want a dependency to install and call, this is the wrong shape of project.
The ten-stage pipeline the examples build
The README lists the stages in order: knowledge requirements, data loading, splitting and chunking, embedding, vector store, retrieval, post-retrieval re-ranking, query preprocessing and embedding normalization, augmentation, then generation. The examples folder mirrors that sequence, from 00_how_rag_works through 01_intro_to_llms, 02_data_loading, 03_text_splitting_and_chunking, 04_intro_to_embeddings, 05_building_vector_store and 06_retrieval_strategies.
Retrieval is where the repository spends most of its effort. Under 06_retrieval_strategies there is basic retrieval, query preprocessing, hybrid search and multi-query retrieval. The hybrid example combines vector similarity with keyword signals, and the README names BM25 plus embeddings as the concept. The multi-query example covers decomposing a query into sub-queries, running them in parallel, and fusing result lists with reciprocal rank fusion or weighted fusion, then deduplicating. Those are real retrieval engineering topics, and they appear as separate runnable folders rather than as options on a client object.
The first example is worth noting for its size: the README says it shows a minimal end-to-end RAG flow in under 70 lines. That constraint is deliberate. It gives you the whole loop before the later folders add ranking and normalization.
Installing it and running the first example
The README does not contain an install section, and the repository has no releases. What it does have is a DOWNLOAD.md file at the top level, alongside .env_example, package.json and a models/ directory. DOWNLOAD.md is where the project tells you how to obtain the model files the examples need; read it before running anything, because the local generation path depends on those files being present.
The package.json declares "type": "module", so the examples are ES modules. Dependencies include node-llama-cpp for local inference, embedded-vector-db for the store, pdf-parse for document loading, dotenv for configuration, and chalk, ora and boxen for terminal output. There is no start script; the only scripts are an empty test entry and a husky prepare hook. You run an example file directly.
A first run looks like this, from the repository root:
npm install
node examples/00_how_rag_works/example.jsAfter the install finishes, the second command runs the smallest example in the repository. According to the README, that file shows a minimal end-to-end RAG flow in under 70 lines, so expect terminal output from the chalk, ora and boxen dependencies rather than a web interface.
The .env_example file and what it configures
Because dotenv is a dependency and .env_example sits at the top level, configuration is expected to live in a .env file. The README does not document the individual keys, so the file itself is the source of truth for names. The presence of the openai package alongside node-llama-cpp suggests the examples can point at either a local model or a hosted endpoint, even though the README's stated philosophy is local and cloud-free. Treat that as an option in the code, not as the recommended path.
If you only want the local route, the pieces you need are the model files from DOWNLOAD.md and whatever key names .env_example defines. The README does not describe environment variable names, so do not guess them; open the file.
Where this approach breaks down
The vector store example is an in-memory store, and the folder name says so: 05_building_vector_store/01_in_memory_store. Nothing in the README describes persistence, an index format on disk, or reloading embeddings after a restart. For a tutorial that is fine, because the point is nearest-neighbor search, not durability. For anything that has to survive a process restart, it means re-embedding your corpus each time unless you write the persistence layer yourself.
The second limitation is the native dependency. node-llama-cpp compiles against your platform, and the README gives no guidance on supported Node versions or troubleshooting a failed build. That is the step most likely to stop a new user, and the repository is silent on it.
The third is scope. The README documents no evaluation harness, no accuracy measurement and no failure analysis for retrieval quality. You can follow every example and still have no way to tell whether your chunk size or your fusion weights are better than the defaults you copied. Maintenance is also worth stating as a fact rather than an impression: the last push to the repository was on 2026-03-11, and there are no releases.
How it differs from a framework-based RAG tutorial
The obvious alternative is a framework such as LangChain, which people search for alongside this project. The difference is where the abstraction sits. A framework gives you a retriever object, a vector store interface and a chain, and you configure them. This repository gives you the loop itself: chunk the text, embed it, score it, take the top k, paste the context into a prompt, generate. The README's hybrid and multi-query folders then implement fusion by hand, which is exactly the code a framework would hide.
That makes the trade-off straightforward. Reading this project teaches you what a retriever is doing internally, which makes debugging a framework later easier. Adopting it does not give you the framework's connectors, its persistence options or its maintenance. The two are complements, not substitutes, and the README's own framing supports that reading.
Licence and what upgrading looks like
The repository is MIT licensed, with the LICENSE file at the top level. MIT permits commercial use and modification provided the copyright notice and permission notice are retained. That is a description of the licence text, not legal advice; if you plan to redistribute the examples inside a product, have your own counsel read the file.
Upgrade cost is unusual here because there is no versioned release to upgrade to. The package.json declares version 1.0.0, but the README lists no releases, so the practical unit of change is a commit on main. You are copying example code, and any later change to the examples does not reach you unless you pull it. The dependency set is small but not trivial: node-llama-cpp is a native module, and embedded-vector-db is at 0.0.4, a pre-1.0 version. Both are the kind of dependency that can change behaviour between installs. Pin them in your own lockfile if you build on top.
Frequently asked questions
The README does not answer every question a new user brings. The ones below are limited to what the repository and its files actually state.
Editorial conclusion
Use pguso/rag-from-scratch if you want to read and run a RAG pipeline end to end in JavaScript before choosing a framework, and if you accept that the examples are numbered folders rather than a supported package. Do not adopt it as the retrieval layer of a product: the vector store example is in-memory and the README documents no persistence, no rollback and no upgrade path. Before you start, open DOWNLOAD.md and confirm which model files the examples expect, then check node-llama-cpp against your Node version, because that native dependency is the part most likely to fail first.
Frequently asked questions
What is RAG vs LLM?
The README describes RAG as giving a language model access to external knowledge by retrieving relevant context before it generates a response, rather than asking the model to remember everything. The LLM is the generation stage at the end of the pipeline.
Can you explain RAG to a beginner?
The README's own explanation is that you retrieve relevant context and feed it into the model's prompt, then generate a grounded answer. The first example folder, 00_how_rag_works, is written to show that loop in under 70 lines.
What does RAG stand for?
Retrieval-Augmented Generation, which is the title the README uses and the pipeline the examples build from data loading through to generation.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/pguso-rag-from-scratch)