Open-source project
chroma-core/chroma avatar
chroma-core/chroma

Chroma: the open-source data infrastructure for AI, and what its four-function API actually gives you

Search infrastructure for AI. Create a DB and try it out in under 30 seconds with $5 of free credits.

29,377 stars2,534 forksRustApache-2.0

At a glance

What is it?
Chroma is an Apache-2.0 vector store with Python and JavaScript clients and a Rust server. It gets you from pip install to a query in minutes, but the repository's own README leaves persistence, scaling and rollback largely to the docs.
Who is it for?
Adopt Chroma if you want a local or self-hosted vector store with a four-function API and you are comfortable reading docs.trychroma.com for anything beyond the README. Do not adopt it if you need a documented rollback path, a published capacity limit, or a guarantee about the current release line: the repository's last push was on 2025-04-01 while tagged 1.5.9 carries a 2026-05-05 timestamp, and nothing in the repository explains that gap.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 4 days ago.
What is it written in?
Mainly Rust, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem Chroma solves, and who it is aimed at

Retrieval over embeddings usually starts as a script: embed some text, keep the vectors in a list, compute cosine similarity, sort. That works until you need persistence, metadata filters, or a second process that also wants to query the same data. Chroma exists to remove that step. The README describes it as "the open-source data infrastructure for AI", and the pyproject.toml names the package chromadb, published under Apache-2.0 with a requires-python of >=3.9.

The audience is narrow but real. You are building a prototype, a RAG demo, or an internal search tool, and you want a store that accepts raw documents and handles tokenization, embedding and indexing on your behalf. The README states exactly that: "we handle tokenization, embedding, and indexing automatically". It also says you can skip that and supply your own embeddings, which is the path most production teams take once they have a model they trust.

What Chroma is not, based on the repository layout, is a general-purpose relational database. The Python dependencies include pypika, a SQL query builder, and the Rust workspace lists a rust/sqlite member, so there is a SQL layer underneath. But the public surface is four functions, and the README presents it that way. If your queries are joins and aggregations over structured rows, this is the wrong abstraction.

How the client, the server and the Rust crates fit together

The repository is a monorepo with two halves. On the Python side, chromadb/ holds the client, and pyproject.toml exposes a console script named chroma pointing at chromadb.cli.cli:app, built with typer. On the Rust side, Cargo.toml declares a workspace with members including rust/frontend, rust/worker, rust/log-service, rust/sysdb, rust/index, rust/segment, rust/blockstore, rust/storage, rust/garbage_collector and rust/python_bindings. The Python client is not a thin HTTP wrapper only: the mypy configuration sets mypy_path to rust/python_bindings, which indicates the client can call into compiled Rust code directly.

That dual nature explains the two deployment shapes in the README. `client = chromadb.Client()` runs in-memory for prototyping, and the README says persistence "can be added easily", without showing how in the excerpt. The other shape is client-server: `chroma run --path /chroma_db_path` starts a server against a directory. The docker-compose.yml confirms the server listens on 8000 and that the default config uses a persist_directory of /data, mounted as a named volume called chroma-data.

The named Rust crates tell you what the server does at scale. There is a log service, a garbage collector, a memberlist crate for cluster membership, and an s3heap for object storage. None of that is documented in the README. If you are evaluating Chroma for a distributed deployment, the crates are the evidence that the design exists, and docs.trychroma.com is where the README sends you for the rest.

Installing chromadb and running a first query

The README gives the install line directly. For Python, it is pip install chromadb. The same block notes that JavaScript users run npm install chromadb, and that client-server mode is started with chroma run --path. There is no build step or system dependency listed in the README.

bash
pip install chromadb # python client
# for javascript, npm install chromadb!
# for client-server mode, chroma run --path /chroma_db_path

After installation, the README's API example is the shortest path to something that works. It creates an in-memory client, creates a collection, and adds documents with metadata. The README calls the core API four functions and links a Google Colab notebook for the full run.

python
import chromadb
client = chromadb.Client()

collection = client.create_collection("all-my-documents")

collection.add(
    documents=["This is document1", "This is document2"],
    metadatas=[{"source": "notion"},

The snippet is truncated in the README at the metadata list, so the second metadata entry and the query call are not shown. Expect to finish it from the docs rather than from the repository front page. The README also mentions that get_collection, get_or_create_collection and delete_collection exist alongside create_collection, and that documents can be updated and deleted, with a "Row-based API coming soon".

For a server deployment, the docker-compose.yml builds from rust/Dockerfile with target cli, maps port 8000, and defines a healthcheck that curls http://localhost:8000/api/v2/heartbeat every 30 seconds with 3 retries. That heartbeat path is the concrete way to confirm a container is serving. The compose file also passes CHROMA_OPEN_TELEMETRY__ENDPOINT, CHROMA_OPEN_TELEMETRY__SERVICE_NAME and OTEL_EXPORTER_OTLP_HEADERS through from the environment, so OpenTelemetry export is configured at the compose level, not in application code.

Where Chroma stops being the right tool

The README is a landing page, not a manual, and that is the first limitation. It does not document rollback, backup, or migration between versions. It does not state index size limits, memory requirements, or how many vectors a single collection is expected to hold. The truncated Python example means a reader cannot copy a complete working program from the repository front page. Anyone who needs those answers is directed to docs.trychroma.com, and the quality of that documentation is not something the repository files let me judge.

The second limitation is release hygiene. The repository's last push is dated 2025-04-01, while the release list shows 1.5.9 and cli-1.4.4 tagged 2026-05-05, and a "latest" release carrying the 2025-04-01 timestamp. Those dates do not line up, and nothing in the README or pyproject.toml explains the discrepancy. The README does state a cadence: tagged pypi and npm releases on Mondays, with hotfixes at any time. That is a policy statement, not evidence about which artifact you get when you run pip install chromadb today.

The third is operational. The compose file sets restart: unless-stopped and a healthcheck, which is sensible, but persistence depends on the chroma-data volume and the default /data path. Lose that volume and you lose the collection; the README describes no export or snapshot command. If your data is reproducible from source documents, that is fine. If it is not, you are trusting a volume you have no documented way to back up.

Chroma against a plain embedded index such as FAISS

FAISS is the obvious alternative for the in-memory case, and the difference is not speed. It is scope. FAISS is a similarity search library: you give it vectors, it gives you nearest neighbours. Chroma wraps that concern in a database abstraction with collections, metadata, and a server mode. The README's four-function API and the automatic tokenization, embedding and indexing are exactly the layer FAISS does not provide.

The trade-off runs the other way too. FAISS has no server, no HTTP API, no collection lifecycle, and no persistence format you would call a database. Chroma adds all of those, and with them the operational questions this article has raised: which release, which volume, which config keys. If you already have an embedding pipeline and only need nearest-neighbour search inside one process, Chroma's extra surface is overhead.

There is a middle option visible in the repository itself. The README says you can add your own embeddings instead of letting Chroma compute them, which keeps the storage and metadata layer while removing the bundled ONNX model from the critical path. The pyproject.toml lists onnxruntime >= 1.14.1 and tokenizers >= 0.13.2 as core dependencies, so the default install carries a local embedding stack whether or not you use it.

Licence, upgrade cost and what the release cadence implies

Chroma is Apache-2.0, and both the README and the pyproject.toml classifiers confirm it. That is a permissive licence with an explicit patent grant, and it does not impose copyleft obligations on your application. It also does not grant trademark rights, and nothing in the repository suggests otherwise. This is not legal advice; read the LICENSE file in the repository root if your situation has specific constraints.

The upgrade cost is harder to estimate. The Python package declares requires-python >= 3.9 and depends on pydantic >= 2.0, pydantic-settings >= 2.0 and grpcio >= 1.58.0, all of which have their own release cycles. The dev extra pins chroma-hnswlib==0.7.6 exactly, which suggests the HNSW index binding is version-sensitive; the runtime dependency list does not pin it, so the index implementation you get may differ from the one the maintainers test against.

The cadence statement in the README, Monday releases plus ad-hoc hotfixes, implies you should expect frequent version churn and should pin whatever you deploy. The repository provides a RELEASE_PROCESS.md at the root, which is where the actual procedure lives; the README only states the schedule. Combined with the unexplained date gap between the last push and the 1.5.9 tag, pinning is not optional for a production deployment.

Editorial conclusion

Adopt Chroma if you want a local or self-hosted vector store with a four-function API and you are comfortable reading docs.trychroma.com for anything beyond the README. Do not adopt it if you need a documented rollback path, a published capacity limit, or a guarantee about the current release line: the repository's last push was on 2025-04-01 while tagged 1.5.9 carries a 2026-05-05 timestamp, and nothing in the repository explains that gap. Before you commit, verify which tag you are actually installing and whether your deployment model is the in-memory client or chroma run --path.

Frequently asked questions

How do I install chromadb?

The README gives one line: pip install chromadb for the Python client. JavaScript users run npm install chromadb instead. The same block notes that client-server mode is started with chroma run --path /chroma_db_path.

How can I use Chroma?

The README's example creates an in-memory client with chromadb.Client(), creates a collection with create_collection, and adds documents and metadata with collection.add. It says tokenization, embedding and indexing are handled automatically, and that you can supply your own embeddings instead.

What is Chroma in AI?

The README describes it as the open-source data infrastructure for AI, distributed as the chromadb package under Apache-2.0. It stores documents and embeddings in collections and answers similarity queries over them.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
  4. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/chroma-core-chroma.svg)](https://hysenlabs.com/projects/chroma-core-chroma)