easy-vecdb: a Chinese-language vector database course that builds one from scratch
📚 从零开始的向量数据库原理与实践教程,在线阅读地址:https://easy-vecdb.datawhale.cc/
At a glance
- What is it?
- easy-vecdb is a Datawhale tutorial repository, not a database. It teaches vector search theory, then drills through Annoy, Faiss and Milvus, and ends with four RAG and agent projects. Here is what the repository actually contains and where it stops.
- Who is it for?
- Adopt easy-vecdb if you want a structured path from brute-force similarity search to a running Milvus RAG example and you read Chinese, or you are willing to follow the Chinese docs alongside the English README. Do not adopt it if you need a maintained Python library, a published package, or English-only teaching material; the repository is a course, and the README lists no releases.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 45 days ago.
- What is it written in?
- Mainly Jupyter Notebook, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What easy-vecdb is, and the reader it is written for
The repository is a course, not a library. Its own description calls it a vector database principles and practice tutorial built from zero, and the navigation splits into a theory half and a practice half. The theory half runs from why a vector database is needed at all, through embedding algorithms such as Word2Vec and Transformer embeddings, through brute-force search and similarity metrics, into a chapter on ANN algorithms covering IVF, PQ, HNSW, LSH and Annoy, and closes with a chapter on implementing your own minimal vector database.
The practice half is organised by tool rather than by concept: three chapters on Annoy, five on Faiss (index types, composite indexes and GPU, tuning and recall evaluation, then engineering and service deployment), and six on Milvus, from architecture and the Collection/Partition/Index data model through PyMilvus APIs to a BM25 hybrid search example, an image retrieval example, and an elective chapter on Milvus internals, rerankers, Milvus Lite and MinerU. Four projects sit at the end: Annoy plus DSSM for recommendation recall, a Faiss RAG, a Milvus agent, and a RAG that pairs Milvus with ArangoDB.
That shape tells you the audience. It is for someone who can read Python and wants to understand why an index returns approximate neighbours before picking one, and for a student or researcher assembling a first RAG pipeline. It is not aimed at a team that wants a drop-in dependency, because there is nothing here to install as a database.
How the material is organised: VitePress, docs, data, src
The top level holds docs, data, src, tmp, plus a package.json whose name is easyvecdb-docs. That package.json is the only build system in the repository, and it is a documentation site, not a runtime. Its scripts are docs:dev, docs:build and docs:preview, all of which invoke vitepress against the docs directory, and its engines field requires Node 18 or newer. Two dev dependencies are declared, markdown-it-mathjax3 and vitepress, and the single runtime dependency is vitepress-plugin-image-viewer.
The mechanism is therefore conventional static-site generation. Markdown chapters live under docs, VitePress compiles them into the online book at the project's GitHub Pages address, and MathJax renders the formulas in the theory chapters. The src directory is described in the README as project-related code and holds the notebooks and scripts the chapters reference; data is a shared sample-data directory and tmp is for temporary files. The primary language of the repository is Jupyter Notebook, which matches the teaching intent: the algorithm chapters are meant to be run cell by cell rather than read as prose.
One practical consequence: because the site is generated from the same Markdown you would read on GitHub, you can clone the repository and run the site locally to read it offline, but you cannot import anything from it into your own application. The code you take away is example code from the chapters, not a packaged API.
Installing easy-vecdb and reading the first chapter locally
There is no installable package. The README points readers at the online site and at the src directory for source code, and the only installation the repository supports is building the documentation site. The package.json engines field requires Node 18 or newer, so check that first.
Clone the repository and install the two dev dependencies plus the image-viewer plugin:
git clone https://github.com/datawhalechina/easy-vecdb.git
cd easy-vecdb
npm installThen start the local VitePress server. The docs:dev script runs vitepress dev against the docs directory, and VitePress prints a local URL you open in a browser:
npm run docs:devFor a static build instead, docs:build compiles the site and docs:preview serves the built output so you can check it before deploying:
npm run docs:build
npm run docs:previewA first real use is not a command but a reading order. Open the navigation table in the README, start at Chapter 2 on why a vector database is needed, work through Chapter 5 on ANN search algorithms, and only then jump into the tool-specific parts. The Annoy, Faiss and Milvus chapters each begin with their own environment setup chapter, so the Python dependencies for those examples are documented inside the chapters rather than at the repository root. The README does not list a requirements.txt or environment.yml, so treat each chapter's setup section as the source of truth for that chapter's packages.
Where the course stops short
The navigation table marks every listed chapter as complete, and a line under it says the project is continuously updated. That line is aspiration, not a schedule; the README does not state a release cadence, and no releases are listed for the repository. The last push was on 2026-08-16.
More concretely, the material is Chinese-first. The repository ships README.md and README_en.md, so the front page has an English version, but the chapters themselves are Markdown files with Chinese filenames such as the Chapter 5 ANN file and the Milvus chapters. An English-only reader gets an English table of contents and a Chinese course. That is the single biggest constraint on adoption outside the Chinese-speaking community, and the README does not promise translated chapters.
There is also a structural gap in the coverage. The theory chapters explain IVF, PQ, HNSW and LSH, and the practice chapters cover Annoy, Faiss and Milvus, but the related-search terms people use around this space include Chroma and general vector database rankings, and neither appears in the navigation. If your stack is Chroma or another engine not on that list, the conceptual chapters still transfer but the hands-on chapters will not match your tooling. Finally, the repository's own directory listing includes a tmp folder and a data folder with no documented schema, so example data is something you inspect rather than rely on.
How easy-vecdb differs from a library such as Faiss or Annoy
The obvious alternative is to skip the course and read the documentation of the engine you plan to use. Faiss and Annoy are libraries with their own APIs and release cycles; easy-vecdb is teaching material that explains what those libraries do internally. If you already know why an HNSW graph gives you approximate recall and you just need the API, the course is overhead. If you have ever tuned an index parameter without knowing what it trades away, the Chapter 5 explanation of IVF, PQ, HNSW and LSH is the part a library README will not give you.
The second alternative is a general RAG tutorial, and the related searches for this project include RAG tutorial and Python RAG variants, which is the right comparison. A typical RAG walkthrough starts at the embedding model and the prompt, and treats the vector store as a black box behind a client call. easy-vecdb inverts that: it spends the first six chapters on embeddings and search algorithms, then a chapter on writing your own minimal vector database, and only afterwards moves to Milvus and the four application projects. The trade-off is real. You reach a working RAG later, but you arrive knowing why the retriever returned what it returned, which matters when recall is bad and you have to decide whether to change the index, the chunking or the embedding model.
Licence and maintenance cost
The licence situation needs care, because two sources disagree. The README's licence section states that the work is licensed under the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International licence and links to the CC BY-NC-SA 4.0 deed, and the badge in the same section repeats that. The repository metadata recorded for this project lists Apache-2.0, and there is a LICENSE file at the top level. The README does not explain the discrepancy.
That matters more than usual for a course, because the artefact you would reuse is the text and the notebooks, not compiled code. CC BY-NC-SA 4.0 carries a non-commercial restriction and a share-alike condition, which is a very different proposition from Apache-2.0 for anyone planning to fold these chapters into paid training material or an internal onboarding course at a company. Open the LICENSE file and the licence section of the README before you reuse anything, and if the two still disagree, ask the maintainers through an issue rather than assuming the permissive reading. This is not legal advice; it is a reason to check the file.
Maintenance cost is low by design. The only thing you can break by upgrading is the documentation build: VitePress is pinned to a caret range on 1.6.4, markdown-it-mathjax3 to 4.3.2, and Node must be 18 or newer. There is no server, no database, and no runtime dependency to track. The cost you carry instead is content drift: if Faiss or Milvus changes an API, the chapters are corrected by hand, and the README gives no version pinning for the Python packages the chapters use.
How to contribute and what the project asks of you
The README describes a standard Datawhale workflow: open an issue for problems, open a pull request for contributions, and if nobody responds, contact the community's support team through the link the README provides. It also points to the Datawhale open source project guide for anyone wanting to start a new project. The contributor list names a project lead and four contributors, and the project thanks one external supporter.
For a reader deciding whether to depend on this material, the contribution model is the relevant signal. Community-maintained course repositories move when someone volunteers, and the navigation table is the only status indicator the repository offers: it shows a check mark per chapter rather than a version number or a changelog. If a chapter you need is marked complete but the code fails against a newer library version, the fix path is an issue, not a release upgrade. That is normal for this kind of project, and it is worth knowing before you build a syllabus around it.
Editorial conclusion
Adopt easy-vecdb if you want a structured path from brute-force similarity search to a running Milvus RAG example and you read Chinese, or you are willing to follow the Chinese docs alongside the English README. Do not adopt it if you need a maintained Python library, a published package, or English-only teaching material; the repository is a course, and the README lists no releases. Before committing time, open the online site, check that the chapter you need is marked complete in the navigation table, and read the licence page, because the README badge says CC BY-NC-SA 4.0 while the repository metadata says Apache-2.0, and that difference decides whether you can reuse the text commercially.
Frequently asked questions
What is easy-vecdb and what does it do?
easy-vecdb is a vector database principles and practice tutorial from the Datawhale community, published as an online book and a repository of Jupyter notebooks. It covers embedding and search fundamentals, ANN algorithms such as IVF, PQ, HNSW and LSH, hands-on chapters for Annoy, Faiss and Milvus, and four application projects. It is teaching material, not a database you install.
How do I build a vector database with easy-vecdb?
The course includes a chapter titled implementing your own vector database, which the README lists as the final part of the theory section. It walks through a minimal implementation rather than a production system, and the later chapters move to Annoy, Faiss and Milvus for real workloads.
Which vector databases does the easy-vecdb tutorial cover?
The navigation lists Annoy, Faiss and Milvus, each with its own multi-chapter section, plus a chapter comparing ANN algorithms including IVF, PQ, HNSW and LSH. Chroma is not covered, and the README does not compare engines by ranking.
Is SQL a vector database in the context of easy-vecdb?
The course material is about dedicated vector search engines and libraries, not SQL databases. The README's chapters and projects reference Annoy, Faiss, Milvus and ArangoDB, and the theory chapters explain similarity search and ANN indexing rather than relational query planning.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/datawhalechina-easy-vecdb)