# LLM-Zero-to-Hundred: a monorepo of RAG, agent and fine-tuning notebooks

> Farzad-R/LLM-Zero-to-Hundred collects eight LLM chatbot projects and two tutorials in one repository, each in its own folder with its own configs, data and source. It is a reference implementation set for engineers who want to read working RAG and agent code, not a library you install.

**Farzad-R/LLM-Zero-to-Hundred** — This repository contains different LLM chatbot projects (RAG, LLM agents, etc.) and well-known techniques for training and fine tuning LLMs.

- Repository: https://github.com/Farzad-R/LLM-Zero-to-Hundred
- Stars: 559 · Forks: 225
- Language: Jupyter Notebook
- License: not declared
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/farzad-r-llm-zero-to-hundred

## What LLM-Zero-to-Hundred actually is, and who it is for

This is not a framework, an SDK or a pip package. It is a collection repository: one GitHub project holding eight chatbot and training projects plus two tutorials, each in a separate top-level folder. The README lists them as Hidden-Technical-Debt-Behind-AI-Agents, WebGPT, RAG-GPT, WebRAGQuery, LLM Full Finetuning, RAG-Master (LlamaIndex vs LangChain), open-source-RAG-GEMMA, and HUMAIN, an advanced multimodal chatbot.

The audience is narrow and specific. If you are an engineer who learns by reading a complete, runnable project rather than a minimal snippet, this layout suits you. Each folder is self-contained: configs, data, src and images. If you want a maintained library with semver releases, this is the wrong shape of thing. The repository has no releases, and the README does not state a licence.

One structural detail matters more than it looks. Every project folder is described as containing a HELPER.md for extra execution information, plus a .here marker file for the project root. That marker convention tells you the author expects you to run code from inside the project folder, not from the repository root. Treat each folder as its own small repository.

## The shared folder layout that every project follows

The README prints the general structure once and applies it across the projects:

```
Project-folder
  ├── README.md
  ├── HELPER.md
  ├── .env
  ├── .here
  ├── configs
  ├── data
  ├── src
  |   └── utils
  └── images
```

The README adds a caveat: this is the general structure, and individual projects may differ slightly because of their own needs. Read that caveat literally. The configs folder holds yml files, data holds sample data, src holds the executable code with a utils subfolder for shared modules. The .env file is described as local configuration via dotenv.

Two consequences follow. First, secrets and endpoints live per project, so you will edit a .env in each folder you try, not one at the repository root. Second, because the same layout repeats, once you have run one project the next one is faster to set up. That repetition is the main design decision here, and it is a reasonable one for a teaching repository.

## Installing and running the RAG-GPT project

The README does not give a single install command for the repository as a whole, and there is no root-level dependency file in the top-level listing. Installation is per project. The README points to the RAG-GPT folder as the base that HUMAIN and Open-Source-RAG-GEMMA were built on top of, which makes it the sensible first project to try.

Start by cloning and entering that folder. The .here file marks the project root, so commands belong inside it:

```bash
git clone https://github.com/Farzad-R/LLM-Zero-to-Hundred.git
cd LLM-Zero-to-Hundred/RAG-GPT
```

Before running anything, read HELPER.md in that folder. The README describes it as containing extra information useful for executing the project, which in practice is where the author puts the steps that do not fit the top-level README. Then open the .env file and fill in the configuration it expects. The README names OpenAI as one of the libraries used across these projects, and the fine-tuning project lists huggingface, OpenAI and chainlit as its libraries. The exact variable names are not printed in the README, so open the file and read the keys it already contains.

Finally, run the project entry point from src. The README says src contains the source code for executing the project and utils holds the supporting modules. The exact script name differs per project and is not listed in the README, so check the folder listing before you run it.

## What the individual projects cover, and where the code is thin

The range is wide. WebGPT handles questions that need internet searches and, per the README, identifies and executes the most relevant given Python functions in response to a user query. That is function calling as a routing mechanism, and the repository also ships a separate LLM Function Calling tutorial. RAGMaster compares five RAG techniques from LangChain and LlamaIndex on 40 questions across five documents, and provides two separate RAG chatbots covering eight techniques from the two frameworks. That is the most evaluation-oriented folder in the set, and the README is explicit about the test size, which is small enough that you should treat the comparison as illustrative rather than conclusive.

open-source-RAG-GEMMA takes RAG-GPT and converts it into a fully open source chatbot using Google Gemma 7B as the LLM and BAAI/bge-large-en as the embedding model, with the stated goal of on-prem deployment. HUMAIN is the largest: a ChatGPT-like assistant with RAG in three modes (preprocessed documents, user-uploaded documents, and any website the user requests), image generation through a stable diffusion model, image understanding through the LLava model, DuckDuckGo search integration, summarization, text and voice input, and session memory. The README states plainly that HUMAIN was built on top of RAG-GPT and WebRAGQuery.

The fine-tuning project is the odd one out. It uses a fictional company called Cubetriangle and lays out a pipeline to process raw data, fine-tune three LLMs on it, and build a chatbot with the best model. That is a training workflow, not an inference app, and it will not fit a reader who only wants to wire up a retrieval chatbot.

## Where this repository will let you down

The first limitation is the licence. The repository listing gives no licence, and the README does not mention one. Without a stated licence, you have no granted permission to reuse the code, and the safe assumption is that default copyright applies. For a repository aimed at teaching, that is an odd omission, and it is the single biggest obstacle to using any of this in a product.

The second is reproducibility. Projects pull in OpenAI, Hugging Face models, chainlit, LangChain and LlamaIndex, and no version pins are visible anywhere in the repository layout or the README. A notebook written against one LangChain release will not run unchanged against a later one, and the RAGMaster comparison in particular depends on framework behaviour that changes between releases. Expect to debug imports before you see output.

The third is that this is a notebook-first repository. The primary language is Jupyter Notebook, which is excellent for explaining a pipeline step by step and poor for shipping it. There is no packaging, no test suite mentioned, and no CI visible in the top-level listing.

Use it as a wrong tool when you need a supported component in a production service, when you need a documented API surface, or when legal review requires a clear licence.

## RAGMaster versus building the same comparison yourself

The obvious alternative to RAGMaster is doing the comparison in your own repository, and the difference in approach is worth stating precisely. RAGMaster is a fixed study: five techniques, 40 questions, five documents, two frameworks, with the results already recorded in the notebook. Building it yourself means you choose the documents, write the questions, and control the embedding model and the chunking parameters.

That control matters because retrieval quality is sensitive to exactly those choices, and a fixed study cannot tell you how the techniques behave on your corpus. The trade is time against relevance. RAGMaster gives you a working harness and eight implemented techniques you can read and adapt; a self-built comparison gives you results that apply to your own data but costs you the setup work that RAGMaster has already done.

A middle path the repository itself suggests: start from RAGMaster's two chatbots, swap in your own documents from the data folder, and keep the evaluation questions in a file you can rerun. You inherit the implementation and still get numbers that mean something for your use case.

## Maintenance, upgrades and what the licence silence means

The last push to the repository was on 2026-04-29. It is not archived. That is recent enough that the code has not been abandoned, but there are no releases, so there is nothing to pin to and no changelog to read. Upgrades here mean tracking upstream changes in LangChain, LlamaIndex, chainlit, Hugging Face and the OpenAI API yourself, because the repository does not declare versions. Budget for that: a project that ran six months ago may need import fixes today, and the fix will be yours to make.

On licensing, the repository shows no licence file in the top-level listing and no licence section in the README. That is not legal advice, and it is not a statement that the code is unusable. It is a gap you should resolve with the author before you copy anything into a commercial codebase. For personal study, running the projects locally carries no such question.

## Conclusion

Adopt it as a reading and reference repository: pick one folder, read its README and HELPER.md, and run that project on its own. Do not adopt it as a dependency, because there is no package, no versioning and no licence file in the repository listing. Before you spend time on any single project, check whether its folder ships a .env.example or only a .env, and confirm which model provider and API key that project expects.

## FAQ

### Is LLM-Zero-to-Hundred free to use?

The repository does not state a licence and the top-level listing contains no licence file, so there is no granted permission to reuse the code. The projects themselves call paid services such as the OpenAI API, which you would pay for separately. Treat it as free to read and free to run locally, with reuse unresolved.

### How do I install LLM-Zero-to-Hundred from GitHub?

There is no repository-wide install. Clone the repository, then enter the folder of the project you want, such as RAG-GPT, read its HELPER.md, fill in its .env, and run the code in its src folder. The .here file in each project marks the root you should run from.

### Which project in LLM-Zero-to-Hundred should I start with?

RAG-GPT is the base that the README says HUMAIN and Open-Source-RAG-GEMMA were built on top of, which makes it the natural entry point. If you care about comparing retrieval techniques, RAGMaster covers eight techniques from LangChain and LlamaIndex. If you want a fully open source, on-prem setup, start with Open-Source-RAG-GEMMA instead.

### Does LLM-Zero-to-Hundred work with open source models instead of OpenAI?

Yes, in one project. The README states that open-source-RAG-GEMMA converts RAG-GPT into a fully open source RAG chatbot using Google Gemma 7B as the LLM and BAAI/bge-large-en as the embedding model, with on-prem deployment as the goal. The other projects list OpenAI among their libraries.

## Sources

- [Farzad-R/LLM-Zero-to-Hundred on GitHub](https://github.com/Farzad-R/LLM-Zero-to-Hundred)
- [Issues](https://github.com/Farzad-R/LLM-Zero-to-Hundred/issues)
- [README](https://github.com/Farzad-R/LLM-Zero-to-Hundred/blob/master/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/farzad-r-llm-zero-to-hundred
