Model or dataset
RMA-MUN/RAGNotebook avatar
RMA-MUN/RAGNotebook

RAGNotebook: a self-hosted knowledge graph notebook built on FastAPI and LangChain

基于LangChain、FastAPI和React的RAG项目,主分支为基于知识图谱的知识管理平台,base-rag分支为开箱即用的基础RAG项目供学习使用

459 stars81 forksPythonMIT

At a glance

What is it?
RAGNotebook is a self-hosted note manager that writes every note into a Neo4j knowledge graph and answers questions with graph-guided retrieval. Its Docker Compose stack is easy to start, but the LLM keys, the Neo4j dependency and the missing rollback story are the parts to check before you commit.
Who is it for?
Adopt RAGNotebook if you are an individual who wants notes, a knowledge graph and a RAG chat in one self-hosted stack, or a developer who wants to read the base-rag branch as a working LangChain example. Do not adopt it if you need a multi-tenant service with a documented upgrade path, or if you cannot run Neo4j, MySQL and Redis alongside it.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 16 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem RAGNotebook targets: notes that are written and never read again

The README states the problem directly: notes get written and never revisited, and knowledge ends up scattered into isolated islands. That is a personal knowledge management complaint, not an enterprise search complaint, and the feature list follows from it. Notes are written in a Markdown editor, saved, and then pushed into a Neo4j graph by an LLM extraction step. Tags and categories (work, study, life, project) are generated asynchronously after saving, so the user does not classify anything by hand. A spaced-repetition review module uses an Ebbinghaus forgetting curve algorithm to bring notes back at intervals. Chat answers cite both graph nodes and note sources.

The intended user is one person running the stack for themselves. The README also names job seekers who need a portfolio project, which is a candid thing to put in a README and tells you something about the project's centre of gravity: it is built to be read and run by individuals, not operated by a platform team. If you want a pure retrieval service to embed in another application, the README points you at the base-rag branch instead, which it describes as a document upload, vector retrieval and question answering pipeline without GraphRAG.

How the graph-guided retrieval pipeline actually fits together

Four storage systems do different jobs. Neo4j holds both the knowledge graph and the chunk-level vector and full-text indexes, so semantic search runs as a hybrid of vector and keyword retrieval fused with RRF. MySQL stores chat history, notes and reviews through SQLAlchemy's async ORM. Redis handles caching, and the repository layout shows a cache directory with Redis decorators. The backend is FastAPI, and the agent layer is described as LangChain's create_agent plus tool definitions, which means the chat path is an agent choosing tools rather than a fixed retrieve-then-generate chain.

The interesting design choice is that retrieval is graph-guided rather than purely top-k. The README describes expansion along graph edges to gather evidence, and the answer carries citations back to graph nodes and notes. That is a real difference from a flat vector store: a note about one topic can pull in a related note that shares no vocabulary, because the LLM extracted a relationship between them at write time. The cost is that graph quality depends entirely on extraction quality. If the extraction step misses a relationship, the expansion has nothing to follow, and the system quietly degrades to ordinary vector search. The README does not document an evaluation harness or a way to inspect extraction accuracy, which is the gap I would want closed before trusting the graph path in daily use.

Isolation is handled at the user level: the README states that RAG retrieval can only reach the current user's own data, and the frontend uses route guards with JWT validation. That is a per-user boundary inside one deployment, not tenant isolation for many organisations.

Installing RAGNotebook with Docker Compose and opening the first note

The README recommends Docker for the fastest path and says you do not need to install Python, Node, MySQL, Redis or Neo4j locally. The prerequisites it lists are Docker Desktop with the engine running, at least 4GB of memory, and about ten minutes for the first build because the backend dependencies are large.

Start by cloning the repository and entering it:

bash
git clone https://github.com/RMA-MUN/RAGNotebook.git
cd RAGNotebook

On Windows the README says you can run start.bat in the repository root. It generates backend/.env if it does not exist, builds and starts all five containers, waits for the backend to be ready, and opens the browser. On Linux and macOS, where start.bat does not apply, create the environment file and fill in the model credentials:

bash
cp backend/.env.example backend/.env
# edit backend/.env: at minimum OPENAI_BASE_URL, OPENAI_API_KEY, OPENAI_MODEL_NAME
docker compose up -d --build

The README notes that you can override MYSQL_ROOT_PASSWORD and NEO4J_PASSWORD in a root .env file if you want different database passwords. When the stack is up, the frontend is at http://localhost:3000 with the default account admin / admin1234, which the README says the backend creates on startup, and the API documentation is at http://localhost:8000/docs. Day-to-day commands are the usual Compose ones:

bash
docker compose up -d          # start again after a reboot
docker compose down           # stop, keeping data
docker compose logs -f backend
docker compose restart backend

One detail matters more than it looks. The README warns that model keys are injected through backend/.env, that the file is gitignored, and that the container addresses and passwords for MySQL, Redis and Neo4j are taken over by docker-compose.yml, so the localhost values in .env should not be changed. The compose file confirms this with an anchor that forces MYSQL_HOST to mysql, REDIS_HOST to redis and NEO4J_URI to bolt://neo4j:7687 for the backend container. If you edit those to localhost while running in Docker, you break service discovery. The README also says the service starts and you can browse without an LLM key, but question answering and other AI features stay unavailable until you configure one.

Where RAGNotebook breaks or is the wrong tool

The heaviest constraint is that Neo4j is not optional. Knowledge graph storage, chunk vectors and full-text search all live in Neo4j, so there is no configuration that lets you run the master branch against Postgres with pgvector or against a standalone vector database. If you already operate a vector store and wanted to add a note-taking front end on top of it, this project does not meet you there.

The second constraint is the LLM dependency. The README says the service starts without a key but that AI features are unavailable. Since graph extraction, tagging, completion and chat all call the model, an unconfigured deployment is close to a Markdown editor with a login page. The reranker is optional and the README states that failures fall back to the original order, which is sensible, but it also means retrieval quality varies with whether the rerank endpoint is reachable.

The third is operational. The README documents docker compose down for stopping with data retained and docker compose down -v as the way to fully reset everything, warning that it clears the databases. There is no documented rollback procedure for a schema or graph change, and no migration tool is named in the README. If you have months of notes in the graph, an upgrade is a leap of faith that the new code reads the old data. For a single user experimenting, that is acceptable. For anything you would call a system of record, it is not.

Finally, the repository's own history is a caveat. The README describes a deliberate pivot from a basic RAG chat service to a knowledge management platform, and states that the base RAG code is permanently kept on the base-rag branch. A pivot means the master branch is younger than the project name suggests, and the two branches are not the same product.

RAGNotebook versus the base-rag branch, and versus a plain vector store

The most useful comparison is internal, because the repository ships both approaches. The base-rag branch is a document upload, vector retrieval and question answering service that the README calls ready to use out of the box and explicitly not GraphRAG. It targets developers who want to integrate retrieval quickly or learn how a RAG pipeline is assembled. The master branch adds note management, knowledge graph construction, spaced repetition, AI writing assistance and cross-source recommendations, and the README says it is the version with long-term maintenance support.

The practical difference is the number of moving parts. base-rag needs an LLM endpoint and a vector store. master needs an LLM endpoint, an embedding model (or fallback to the LLM credentials), Neo4j, MySQL and Redis, plus optionally a reranker and a web search provider. If your goal is to answer questions over a folder of PDFs, base-rag gives you that with fewer failure points, and switching branches is cheaper than removing a graph layer later.

Against a generic vector store plus a chat UI, the difference is what gets indexed. A vector store indexes chunks. RAGNotebook indexes chunks and entities, and the agent can traverse relationships during retrieval. Whether that pays off depends on your notes containing relationships worth extracting. A journal of daily entries with few named entities will not gain much from a graph, and you will be paying Neo4j's operational cost for nothing. Notes about projects, people, papers and their connections are the case where the design earns its complexity.

Licence terms, upgrade cost and the maintenance picture

RAGNotebook is MIT licensed, and the LICENSE file sits at the repository root. MIT permits commercial use, modification and redistribution provided the copyright notice and permission notice are kept. It gives no patent grant and no warranty, so if you build a product on it, the compliance work is yours. That is a description of the licence text, not legal advice; talk to a lawyer about your own situation.

The repository is not archived, and the last push was on 2026-09-09, with releases v2.3.0 on 2026-08-30, v2.3.1 on 2026-09-07 and v2.3.2 on 2026-09-09. The cadence is recent and the version numbers are moving, which is consistent with the README's claim of long-term maintenance for the master branch, but the README does not describe a support policy, a deprecation window or a compatibility guarantee between minor versions.

Upgrade cost is dominated by the data, not the code. Pulling a new image and running docker compose up -d --build is trivial. What is not documented is what happens to an existing Neo4j graph and MySQL schema when the application code changes, or how to back up and restore the graph before an upgrade. The compose file keeps data in named volumes (mysql_data, redis_data, neo4j_data) and the README says uploaded files and logs live under backend/media, backend/logs and backend/data. Those are the directories and volumes to copy before you upgrade, and the README does not say so.

Editorial conclusion

Adopt RAGNotebook if you are an individual who wants notes, a knowledge graph and a RAG chat in one self-hosted stack, or a developer who wants to read the base-rag branch as a working LangChain example. Do not adopt it if you need a multi-tenant service with a documented upgrade path, or if you cannot run Neo4j, MySQL and Redis alongside it. Before you commit data, verify three things: that your OpenAI-compatible endpoint accepts the model name in OPENAI_MODEL_NAME, that the Neo4j container's initial password matches NEO4J_PASSWORD on first boot, and that you are comfortable with the fact that the only documented full reset is docker compose down -v, which clears the databases.

Frequently asked questions

What does RAG stand for in RAGNotebook?

RAG stands for retrieval-augmented generation, the pattern where a model answers using documents fetched from a store rather than from its parameters alone. RAGNotebook implements it as the core engine of the whole system, with the README stating that RAG remains the central engine across both branches.

Does RAGNotebook use GraphRAG?

Yes, on the master branch. The README describes the master branch as Agentic RAG plus KnowledgeGraph, where an LLM extracts entities and relationships into Neo4j and retrieval expands along graph edges. The base-rag branch is explicitly described as not GraphRAG.

How is RAG different from generative AI in this project?

In RAGNotebook the generative model is only one component. The README describes an agent built with LangChain's create_agent and tools that retrieves from Neo4j and MySQL before answering, and the answer carries citations back to graph nodes and notes instead of being produced from the model alone.

Is RAG still relevant for a project like RAGNotebook?

The README treats it as the foundation rather than a transitional technique, stating that RAG remains the core engine of the system while the surrounding product moved from a plain chat service to a knowledge management platform. The base-rag branch is kept permanently for anyone who only wants the retrieval service.

Official sources

  1. Issues
  2. License: MIT
  3. README
  4. Releases
  5. RMA-MUN/RAGNotebook on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/rma-mun-ragnotebook.svg)](https://hysenlabs.com/projects/rma-mun-ragnotebook)