CocoIndex: incremental indexing for AI agent context, reviewed
Incremental engine for long horizon agents 🌟 Star if you like it!
At a glance
- What is it?
- CocoIndex is an Apache-2.0 Python framework that declares a transformation over source data and keeps a derived index in a target store up to date, reprocessing only the delta. It fits teams whose retrieval corpus changes continuously; it is the wrong tool for one-off batch embedding jobs.
- Who is it for?
- Adopt CocoIndex if your retrieval corpus changes continuously and you already run a vector database or graph store you intend to keep as the serving layer; the declarative Python model is a reasonable fit for teams that would otherwise hand-write change detection. Do not adopt it for a one-off embedding job over a frozen dataset, and do not adopt it expecting it to host retrieval for you.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 2 days ago.
- What is it written in?
- Mainly Rust, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem CocoIndex solves: stale context in long-running agents
An agent that answers questions over a codebase, a Slack export, or a folder of meeting notes is only as good as the freshness of the index behind it. The usual approach is a batch job: crawl the source, chunk it, embed it, write to a vector store, run again tomorrow. Between runs the index drifts. When the job does run, it typically re-embeds everything, which costs money on any hosted embedding model and takes time proportional to corpus size rather than to what changed.
CocoIndex targets that gap. The project describes itself as an incremental engine for long horizon agents, and the README frames the pitch as turning codebases, meeting notes, inboxes, Slack, PDFs, and videos into live context with minimal incremental processing. The stated design goal is that only the delta is reprocessed on every change.
The audience is narrower than the tagline suggests. This is for engineers building retrieval or agent memory over sources that mutate, who are willing to describe their pipeline as code and to keep running the process. It is not a hosted service and not a drop-in index. The pyproject.toml description is explicit about the model: users declare the transformation, CocoIndex creates and maintains an index, and keeps the derived index up to date based on source updates.
How the declarative dataflow engine tracks what changed
The architecture is a Rust core exposed to Python through PyO3. The Cargo workspace lists rust/core, rust/py, rust/ops_text, rust/code_ast, rust/code_match, and two SDK crates, while pyproject.toml builds with maturin and declares a single console entry point, cocoindex = "cocoindex.cli:cli". So the scheduling, change detection, and state live in Rust, and the pipeline definition lives in Python.
The programming model is a dataflow graph you build in Python. You describe sources, transformations, and a target store, and the engine diffs the current state of the sources against what it recorded on the previous run. Components that did not change are skipped. The README compresses this into three claims: incremental, only the delta; any scale, parallel by default; declarative, Python, five minutes.
Because the graph is declared rather than scripted, the engine can reason about which downstream nodes depend on which upstream inputs. That is what makes partial reprocessing possible at all. It also means the framework owns the state database, and that state is what you must keep intact across runs. The repository ships a benchmarks directory and a rust/code_ast crate, so AST-level handling of source code is a first-class concern rather than an afterthought.
What the README does not document is rollback. There is no described mechanism for reverting the derived index to an earlier state if a bad transformation lands. Treat the index as rebuildable from source rather than as something you can rewind.
Installing CocoIndex and running a first pipeline
The package is on PyPI as cocoindex. The pyproject.toml requires Python 3.11 or newer and classifies support for 3.11 through 3.14, with free-threaded builds marked beta. Install it with pip in whatever environment you use for the pipeline:
pip install cocoindexThat pulls in the runtime dependencies declared in pyproject.toml, including click, rich, python-dotenv, watchdog, numpy, psutil, and msgspec. The watchdog dependency is consistent with local file watching as a change source. After install, the cocoindex console script is available.
The fastest way to see the model is to copy an example rather than start from a blank file. The repository has a large examples directory, and the names describe the shape of each pipeline: examples/code_embedding, examples/multi_format_indexing, examples/kafka_to_lancedb, examples/meeting_notes_graph_neo4j, examples/docs_to_knowledge_graph, examples/gdrive_text_embedding. Pick the one closest to your source and target combination and read it end to end before editing.
A pipeline follows the same three-part shape in every example. You declare a source, chain transformations over the records, and declare a target store to export into. The exact API surface is documented at cocoindex.io/docs, which the README links as the reference for quickstart, connectors, ops, transformations, and target stores. Run the example as-is first, confirm rows appear in your target store, then edit one source file and run it again. The second run should touch only that file. That behaviour, not the first run, is what you are evaluating.
Where CocoIndex stops being the right tool
The clearest limitation is that CocoIndex does not serve retrieval. It maintains an index in a target store. If you want a search API, you build it on top of the vector database or graph database that CocoIndex writes into. The examples confirm this: targets include LanceDB, Kafka, BigQuery, Neo4j, FalkorDB, and SurrealDB, each of which is a separate system you operate.
A second constraint is state. Incremental behaviour depends on the engine's record of previous runs. Any workflow that treats the pipeline as stateless, for instance rebuilding a container from scratch each time and pointing it at a fresh volume, throws that away and degrades to a full reprocess. The README does not document a supported way to migrate or export that state between environments.
Third, the cost model shifts rather than disappears. Skipping unchanged inputs saves embedding calls and processing time, but you still pay for the source scan that determines what changed. For a corpus that changes in almost every file on every run, incremental processing has little to skip and the framework's advantage narrows to its parallelism.
Finally, this is a young project with a specific shape. The declared classifiers say Production/Stable, and the version line is at 1.0.x, but the API is Python-first and the examples are the practical documentation. If your team wants a managed ingestion service with a support contract, this is not that.
CocoIndex compared with hand-rolled ETL and with hosted ingestion
The honest alternative for many teams is a script. A Python file that walks a directory, hashes each file, skips hashes it has seen, chunks and embeds the rest, and upserts into a vector store is perhaps a hundred lines. It has no Rust core, no state database you did not write, and no framework upgrade to track. What it lacks is the dependency graph: when you add a summarization step or a second derived collection, you reimplement the invalidation logic yourself, and that logic is where these scripts rot.
Against hosted ingestion services, the difference is placement. A hosted service typically runs the crawl and embedding on the vendor's infrastructure and hands you a query endpoint. CocoIndex runs in your process, writes into a store you chose, and leaves serving to you. That is more operational work and more control. It also means your data does not have to leave your network to be indexed, which matters when the source is an internal codebase or a mailbox.
Against a plain vector database client, the difference is that CocoIndex treats the index as derived state with a maintenance contract, rather than as a collection you write to and forget. If your corpus is a frozen snapshot, that contract buys you nothing.
Licence, releases, and the cost of upgrading
CocoIndex is Apache-2.0, declared in both pyproject.toml and the Cargo workspace package metadata. That is a permissive licence with an explicit patent grant and no copyleft obligation on your own code. The distribution includes a THIRD_PARTY_NOTICES.html file, which is standard for a project with this many Rust and Python dependencies; if you redistribute the package, that file is the one to carry along. This is a description of what the repository states, not legal advice.
On cadence: the latest release is v1.0.21, dated 2026-09-05, following v1.0.20 on 2026-08-12 and v1.0.19 on 2026-08-04. The last push to the default branch was on 2026-09-09. The repository is not archived. Patch releases at that spacing suggest active maintenance of the 1.0 line, and the version numbering implies API stability within the major version, though the project has not published a compatibility policy that the README describes.
Upgrade cost concentrates in two places. The Python API surface is what your pipelines import, so a minor release can change how a pipeline is written. The state database format is what your existing index depends on, and a format change on upgrade is the scenario worth testing before you roll it out. The repository does not document a rollback path for either.
Editorial conclusion
Adopt CocoIndex if your retrieval corpus changes continuously and you already run a vector database or graph store you intend to keep as the serving layer; the declarative Python model is a reasonable fit for teams that would otherwise hand-write change detection. Do not adopt it for a one-off embedding job over a frozen dataset, and do not adopt it expecting it to host retrieval for you. Before committing, verify three things against your own data: that the connectors you need exist in the examples directory, that your target store appears among the supported targets in the docs, and that a second run after editing one source file reprocesses only that file rather than the corpus. The last check is the whole product.
Frequently asked questions
What is CocoIndex?
CocoIndex is an Apache-2.0 framework for building and maintaining indexes over source data such as codebases, documents, and meeting notes, with a Rust core and a Python API. Its stated purpose is to keep a derived index continuously up to date while reprocessing only the data that changed.
How do I use CocoIndex?
Install the cocoindex package with pip, then declare a pipeline in Python that names a source, a chain of transformations, and a target store. The repository's examples directory contains complete pipelines for combinations such as code embedding and Kafka to LanceDB, and the docs at cocoindex.io/docs cover connectors, ops, and target stores.
What are the alternatives to CocoIndex?
The main alternative is a hand-written incremental script that hashes files, skips unchanged ones, and upserts into a vector store, which avoids a framework dependency but requires you to reimplement invalidation as the pipeline grows. Hosted ingestion services are the other option; they run the crawl and embedding for you but hand back a query endpoint rather than writing into a store you operate.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/cocoindex-io-cocoindex)