NucliaDB: a hybrid search database for RAG, and what its AGPL licence means for you
NucliaDB, The AI Search database for RAG
At a glance
- What is it?
- NucliaDB stores unstructured data and searches it with vector, full text and graph indexes. It is written in Rust and Python, ships as a Docker image, and is licensed AGPLv3, which shapes who can adopt it.
- Who is it for?
- Adopt NucliaDB if you need one store for text, vectors, labels and original files, and you can live with AGPLv3 or buy the hosted service instead. Skip it if you need permissive licensing in a closed product, or if you want a single-language codebase you can read end to end.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What problem NucliaDB solves, and who it is for
Most search stacks force a choice. A keyword engine handles exact terms well and semantic similarity badly. A vector store handles embeddings well and gives you nothing for filters, labels or the original file. NucliaDB is built to hold both sides in one database. The README describes it as an out of the box hybrid search database using vector, full text and graph indexes, and the feature list includes storing text, files, vectors, labels and annotations.
The intended user is a team building retrieval augmented generation or semantic search over messy content: documents, links, conversations, HTML, Markdown. The README lists field types of text, file, link and conversation, which tells you the data model expects more than plain strings. If your corpus is already clean embeddings in a table, NucliaDB is more machinery than you need. If your corpus is PDFs, web pages and support threads that you want to search by meaning and by keyword at the same time, the hybrid index is the point.
The project also targets multi tenant deployments. The README states it was designed to index large datasets and provide multi tenant support, and it lists role based security with upstream proxy authentication validation. That is a hosting concern, not a library concern, and it explains why the deployment surface is several services rather than one process.
How the indexes, storage and services fit together
The repository layout shows a split between Python and Rust. Under nucliadb/ sits the Python service code. Under nidx/ sits a separate Rust indexer, with nidx_binding and nidx_protos as workspace members in pyproject.toml. The two halves are packaged separately: the default Dockerfile explicitly builds without the binding and notes that Dockerfile.withbinding produces a version with a compiled binding for standalone use. That is a real architectural fork, and it matters when you deploy.
Persistence is split by data shape. The README lists a storage layer on PostgreSQL and blob support through an S3 compatible API, GCS and Azure Blob Storage. So metadata and resource structure live in Postgres, original files and blobs live in object storage, and the searchable indexes live in the Rust indexer. The README also mentions replication of index storage and distributed search, which are the features you would need to run more than one index node.
On the write path, resources carry multiple fields and metadata, and the README says fields, paragraphs and semantic sentences are indexed separately. That is the detail that makes hybrid retrieval work: a paragraph level index lets a long document return a relevant passage rather than a whole file. On the read path, you can issue keyword queries or vector queries. The README frames the vector side as returning closest matches, and notes that with NLP this finds similar sentences without exact keyword constraints.
One honest caveat about scope. The README markets cloud data and insight extraction through the Nuclia Understanding API and model training through the Nuclia Learning API. Those are hosted services from the vendor, not parts of this repository. Self hosting NucliaDB gives you the database and the indexes. The enrichment that turns raw documents into well structured fields is a separate purchase.
Installing NucliaDB and running a first query
The README points to the quickstart at docs.nuclia.dev rather than giving install steps inline, so the commands below come from the repository files, not from a documented tutorial. The Makefile defines the local build path. The target build-nucliadb-local runs docker build with the tag nuclia/nucliadb:latest against Dockerfile.withbinding.
make build-nucliadb-localThat produces a local image tagged nuclia/nucliadb:latest. If you want the version without the compiled Rust binding, the default Dockerfile builds that variant, and its comment says so directly. The same Dockerfile declares the ports the container listens on, so you know what to publish. It exposes 8080/tcp for HTTP, 8060/tcp for gRPC ingest and 8040/tcp for gRPC train, and its CMD is the nucliadb executable. Those three ports are the whole public surface of the container, and the image puts /app/bin on PATH.
For Python work against the database, the workspace defines client packages. The pyproject.toml dependency group named sdk pulls in nucliadb_sdk and nucliadb_dataset, and the workspace members list both. The Makefile also defines a plain install target that runs uv sync against the whole workspace. The README states you can export data in a format compatible with most NLP pipelines, naming HuggingFace datasets and pytorch, which is what nucliadb_dataset is for.
make installAfter that, the packages are available in your environment. The README does not walk through a first API call, so treat the API reference at docs.nuclia.dev/docs/api as the place to find the actual request shapes rather than guessing endpoints.
Where NucliaDB is the wrong tool
The licence is the first constraint, and it is not a footnote. The README states NucliaDB is open source under the GNU Affero General Public License Version 3, and it spells out the consequence in plain terms: you may use it as long as you do not modify it, and if you do modify it, you have to make the modifications public. For a team embedding the database inside a proprietary product, that is a meaningful restriction, and it is the single most common reason to look elsewhere. The repository also carries LICENSE_Apache_v2.0.txt alongside LICENSE_AGPLv3.0.txt, so the licence situation is per component rather than uniform. The README does not break down which directory falls under which file, and I would not assume.
Operationally, this is not a single binary you drop into an existing app. Postgres and an object store are part of the design, and the README lists S3 compatible storage, GCS and Azure Blob Storage as the blob options. If your environment has neither Postgres nor object storage, you are adopting three pieces of infrastructure, not one.
The README does not document rollback, backup and restore procedures, or an upgrade path between major versions. There are three releases in the recent list, v6.15.1 in July 2026 and v7.0.0 and v7.1.0 in August 2026, and the jump from 6.x to 7.x is exactly the kind of boundary where you want migration notes. The README is silent on them. Check the release notes before planning an upgrade.
Finally, if you want a small codebase you can read in an afternoon, this is not it. Python services plus a Rust indexer plus generated protobuf packages plus a separate SDK is a lot of surface area for a team without someone who can debug across two languages.
NucliaDB compared with Elasticsearch and pgvector
The README answers the Elasticsearch comparison itself. It says the core difference is an architecture built from the ground up for unstructured data, with vector, keyword, graph and fuzzy search behind one API, plus the extracted information from the Nuclia Understanding API. Read that carefully: the claimed advantage is partly the index design and partly the enrichment pipeline that feeds it. Elasticsearch with a dense vector field gives you hybrid retrieval too, but it does not ship a document understanding service, and it does not assume a multi tenant knowledge box model.
The more interesting comparison for a small team is pgvector inside the Postgres you already run. That approach keeps one datastore and one query language, and you write the hybrid ranking yourself. NucliaDB instead separates storage from indexing, runs a dedicated Rust indexer, and gives you the hybrid behavior as a product feature. The trade is operational weight for less glue code. If your retrieval logic is simple and your corpus is small, the glue code is cheaper than the extra services.
Against a plain vector database, the difference is field and paragraph granularity. A vector store typically returns a chunk you inserted. NucliaDB indexes fields, paragraphs and semantic sentences separately, so the unit of retrieval is chosen by the database rather than by your chunking script. That is a genuine design difference, and it is also a constraint: you work within its resource and field model instead of storing arbitrary payloads.
Maintenance, releases and licence obligations
The repository is not archived, and the last push was on 2026-09-10, which is recent. Releases are frequent: v6.15.1 on 2026-07-06, v7.0.0 on 2026-08-03 and v7.1.0 on 2026-08-11. A cadence like that cuts both ways. You get fixes quickly, and you also get major version boundaries that you have to plan for. The README does not describe a supported upgrade path between 6.x and 7.x, so pinning a version in your deployment and reading the release notes before moving is the practical approach.
Contributing has its own paperwork. The repository root contains NucliaDB_individual_CLA.md and a CONTRIBUTING.md, and the README asks contributors to read the Contributor Covenant Code of Conduct and submit a pull request from a fork. The Makefile provides license-check and license-fix targets that run the Apache SkyWalking Eyes license-eye container against the working tree, which is how the project enforces its header policy. If you fork and modify the code, those targets tell you what the project expects.
On licensing, the README states the terms and I am repeating them rather than interpreting them: AGPLv3, free to use as long as you do not modify NucliaDB, and modifications must be made public. The repository also carries an Apache 2.0 licence file, and the README does not explain the split. If your use case turns on which file applies to which component, that is a question for your own counsel, not for this article. The vendor's stated business model is the normalization API plus NucliaDB as a hosted service, which is the usual shape for an AGPL project: the code is open, and the convenient path is paid.
Editorial conclusion
Adopt NucliaDB if you need one store for text, vectors, labels and original files, and you can live with AGPLv3 or buy the hosted service instead. Skip it if you need permissive licensing in a closed product, or if you want a single-language codebase you can read end to end. Before committing, read LICENSE_AGPLv3.0.txt and LICENSE_Apache_v2.0.txt in the repository to see which parts fall under which licence, and confirm whether you need the compiled nidx binding or the Dockerfile without it.
Frequently asked questions
What is the best database for AI?
There is no single answer, and NucliaDB's own positioning is narrower than the question. It is a hybrid search database for unstructured data, using vector, full text and graph indexes, and the README frames its advantage as an architecture built from the ground up for that data rather than as a general purpose AI store.
What is a RAG database?
In this project's terms, it is a store that holds your source content and returns relevant passages for a language model to use. NucliaDB stores text, files, vectors, labels and annotations, indexes fields, paragraphs and semantic sentences separately, and supports both keyword and vector queries over the same resources.
Is there an AI database?
NucliaDB is one example. The README describes it as an out of the box hybrid search database for unstructured data, written in Rust and Python, with a PostgreSQL storage layer and blob support through S3 compatible storage, GCS or Azure Blob Storage.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/nuclia-nucliadb)