Model or dataset
cocoindex-io/cocoindex-code avatar
cocoindex-io/cocoindex-code

cocoindex-code: an AST-aware code search CLI for coding agents

A super light-weight embedded code search engine CLI (AST based) that just works - improves speed and efficiency for coding agent 🌟 Star if you like it!

2,727 stars225 forksPythonApache-2.0

At a glance

What is it?
cocoindex-code wraps a Rust-based indexing engine in a Python CLI that exposes semantic code search to Claude Code, Grok and other agents. It installs in one command, but the README is thinner than the pitch.
Who is it for?
Adopt cocoindex-code if you already drive Claude Code or Grok and want semantic retrieval without running a separate vector database; the plugin marketplace and the ccc skill make that path short. Skip it if you need a documented, stable search API for production tooling, if your repository's language coverage is unclear, or if you cannot accept a local embedding model download on first run.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 8 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem cocoindex-code targets: agent context windows filled by grep

A coding agent that answers a question about a large repository usually starts by grepping. Grep returns lines, not structure, so the agent reads whole files to reconstruct meaning, and the token budget disappears before the actual reasoning starts. cocoindex-code attacks that step. It builds a semantic index of the codebase and lets the agent retrieve the relevant region instead of the whole file. The README frames the payoff as "Instant token saving by 70%", which is a claim from the project, not a measured result here.

The intended user is narrow and specific. This is for someone running an agent that supports skills or MCP, working in a repository large enough that full-file reads are expensive. If your codebase fits in a single context window, or you are not using an agent at all, the index is overhead you will pay for and rarely query.

How the index is built: tree-sitter ASTs, sqlite-vec and a Rust engine

The repository topics list tree-sitter, ast, rust and cocoindex, and the README describes the tool as "Built on CocoIndex, a Rust-based ultra performant data transformation engine." That gives the shape of the data flow. Source files are parsed into syntax trees rather than split on whitespace, so the unit being embedded is a syntactic construct, not an arbitrary line window. The chunks are embedded and stored in a local index. The pyproject.toml dependencies confirm the storage layer: sqlite-vec sits alongside numpy, msgspec, pydantic and typer, and the package requires Python 3.11 or newer.

Two embedding paths exist. The slim install depends on LiteLLM and therefore a cloud embedding provider with an API key. The full install adds sentence-transformers through the embeddings-local extra, so embeddings run locally. The README states the ccc init interactive prompt defaults to Snowflake/snowflake-arctic-embed-xs. Index state lives in .cocoindex_code/, which the Grok hook section references when deciding whether an incremental index is needed.

That AST-first choice is the real differentiator against line-based indexing, and it is also where the risk sits. Parsing depends on grammar availability per language. The README does not publish a language support table, so the honest position is that you should test your primary language before assuming coverage.

Installing cocoindex-code and running a first index

The README gives two package managers. pipx is the first, and the extra in brackets matters: without it you get the slim build and need a cloud embedding key.

bash
pipx install 'cocoindex-code[full]'
pipx upgrade cocoindex-code

uv is the alternative, and the README shows the same extra with --upgrade folded into the install command.

bash
uv tool install --upgrade 'cocoindex-code[full]'

The README notes the full extra pulls in sentence-transformers, roughly 1 GB of torch and transformers, so the first install is not instant even though setup is described as one minute. After install, the agent path avoids manual steps entirely. The skill teaches the agent to run initialization, indexing and search itself, so the documented command is just:

bash
npx skills add cocoindex-io/cocoindex-code

The README states that no ccc init or ccc index is needed on that path. For direct control, the CLI is the other route: ccc init runs the interactive prompt that selects the embedding model, ccc index builds or refreshes the index, and ccc search queries it. The README also mentions typing /ccc inside the agent to invoke the skill explicitly, and asking in plain language, for example "find how user sessions are managed".

Claude Code users can install the same repository as a plugin marketplace instead, which adds version pinning and an update command:

text
/plugin marketplace add cocoindex-io/cocoindex-code
/plugin install cocoindex-code@cocoindex-code

The MCP server and the refresh_index default

The .mcp.json entry in the repository root exposes ccc mcp as a stdio server. The README describes the single tool it offers as search, with refresh_index=true by default. That default is a deliberate trade. Every search call can trigger a refresh, which keeps results current while the agent edits files, at the cost of doing indexing work inside the search path. On a large repository, that is the setting most likely to make a search feel slow.

The Grok plugin bundles the same MCP server alongside a skill and a hook. The hook fires on SessionStart and on PostToolUse for Edit and Write, running an incremental ccc index when .cocoindex_code/ exists. So Grok users get three overlapping mechanisms, and the README explicitly documents how to disable the optional ones if you want the skill alone. Note also that Grok does not import Claude's enabledPlugins or plugin cache, so a team running both agents installs twice.

Where cocoindex-code is the wrong tool

The packaging is candid about maturity: pyproject.toml carries the classifier Development Status :: 3 - Alpha. Treat that as the governing constraint. An alpha tool that rewrites its index on agent edits is a reasonable bet for a solo developer experimenting with agent workflows, and a poor one for a CI pipeline that depends on stable search output.

There are concrete failure modes visible in the design. The hook only runs incremental indexing when .cocoindex_code/ already exists, so a fresh clone with no index directory gets no automatic indexing until something initializes it. The slim install fails without an embedding provider key, which is a silent trap if you copy the bare package name from a blog post rather than the README. The full install downloads a local model on first use, which is a problem on metered or offline machines. And because the retrieval unit is an AST chunk, generated code, heavily templated files, or languages without a grammar will not chunk the way you expect. The README does not document rollback or index corruption recovery, so plan on deleting .cocoindex_code/ and reindexing if state goes bad.

Alternatives and how their approach differs

The closest comparison in the same search space is a plain text index such as ripgrep or the grep tooling agents already ship with. The difference is the unit of retrieval: grep matches strings and returns lines with no notion of function boundaries, while cocoindex-code parses first and embeds syntax-aware chunks. Grep needs no index, no model download and no daemon; cocoindex-code needs all three but returns context the agent can use without reading the surrounding file.

A second alternative is standing up a general vector database and embedding your code yourself. That gives you control over the embedding model, the chunking strategy and the query API, and it is the right answer if you need search across more than one repository or want to tune recall. What you give up is the agent integration: cocoindex-code ships the skill, the .mcp.json server and the Grok hook, so the retrieval is reachable from the agent without you writing a tool wrapper. If your agent is not on the supported list, that advantage disappears and you are left with a Python CLI around an alpha-indexed store.

A third option, which the README itself gestures at, is using the underlying CocoIndex engine directly. cocoindex-code is a thin layer over it, so a team with unusual chunking requirements is better served building on the engine than configuring the CLI.

Licence and the cost of keeping the index current

Both LICENSE and the pyproject.toml license field state Apache-2.0, and the package.json for the OMP extension host repeats it. Apache-2.0 is permissive and includes a patent grant, which matters if you embed the tool in a commercial product. The dependencies are the part to check yourself: cocoindex[litellm], sentence-transformers, torch and transformers each carry their own terms, and the local embedding model you select at ccc init has its own licence on Hugging Face. Nothing here is legal advice; read the licence files of the model you actually pick.

Maintenance cost is dominated by index freshness rather than the tool itself. The last push to the repository was on 2026-09-08, and the most recent release listed is v0.2.41 from 2026-08-07, so the project is being changed frequently. Frequent releases on an alpha package mean upgrades are worth pinning: the README's own upgrade commands (pipx upgrade cocoindex-code, or uv tool install --upgrade) will move you to whatever is current, and the plugin marketplace path offers /plugin marketplace update instead. Budget for reindexing after an upgrade, since the embedding model or chunking can change between versions.

Editorial conclusion

Adopt cocoindex-code if you already drive Claude Code or Grok and want semantic retrieval without running a separate vector database; the plugin marketplace and the ccc skill make that path short. Skip it if you need a documented, stable search API for production tooling, if your repository's language coverage is unclear, or if you cannot accept a local embedding model download on first run. Before committing, verify three things: that your language is parsed by the AST layer, that .cocoindex_code/ is excluded from version control, and that the archived-or-not status of the upstream CocoIndex engine still matches what your agent integration assumes.

Frequently asked questions

How does code indexing work in cocoindex-code?

The tool parses source files into ASTs using tree-sitter, embeds the resulting chunks, and stores them in a local index under .cocoindex_code/ backed by sqlite-vec. ccc index builds or refreshes that index, and ccc search queries it.

What is an index for coding?

In this project it is a local store of embedded, syntax-aware code chunks that an agent can query semantically instead of reading whole files. The README presents this as a way to save tokens, claiming 70% savings.

How to index code for an LLM with cocoindex-code?

Install the full extra so local embeddings work, run ccc init to choose the model, then ccc index. Alternatively, install the ccc skill and the agent handles initialization, indexing and searching itself, with no ccc init or ccc index required.

Official sources

  1. cocoindex-io/cocoindex-code on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/cocoindex-io-cocoindex-code.svg)](https://hysenlabs.com/projects/cocoindex-io-cocoindex-code)