# Claude Context MCP: Semantic Code Search for AI Coding Agents

> Claude Context is an MCP plugin from Zilliz that gives Claude Code and other AI coding agents semantic search over entire codebases by storing indexed code in a Zilliz Cloud vector database and retrieving relevant snippets on demand. It requires a Zilliz Cloud account for the vector store and an OpenAI API key for embeddings, and installs through a single command in Claude Code.

**zilliztech/claude-context** — Code search MCP for Claude Code. Make entire codebase the context for any coding agent.

- Repository: https://github.com/zilliztech/claude-context
- Website: https://github.com/zilliztech/claude-context/tree/master/docs
- Stars: 12,535 · Forks: 923
- Language: TypeScript
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/zilliztech-claude-context

## The Context Window Problem This Plugin Addresses

AI coding agents have finite context windows. Loading entire project directories into every prompt is one way to give the agent codebase awareness, but for large projects it is expensive in token cost and often still insufficient. A million-token context window sounds large, but real production codebases easily exceed it, and loading unnecessary files raises the cost of every request.

Claude Context addresses this with retrieval augmented generation. It works as a Model Context Protocol (MCP) server that integrates with Claude Code and other compatible AI coding agents. When Claude Code needs to understand a part of the codebase not currently in context, it calls the Claude Context MCP tool with a natural language query. The server finds semantically relevant code from the indexed repository and injects those snippets into the current context. The README describes this as letting Claude retrieve results without multi-round discovery.

The trade-off is external infrastructure. The plugin requires a Zilliz Cloud account for vector storage and an external embedding model API to convert code into vectors. Teams that need a fully local or air-gapped setup cannot use the default configuration without additional work.

## How Semantic Indexing Works Under the Hood

Claude Context indexes a codebase by splitting source files into chunks, generating a vector embedding for each chunk using the configured embedding model, and storing those vectors in a Milvus collection on Zilliz Cloud. When an agent issues a code search query, the plugin embeds the query using the same model and retrieves the most semantically similar code chunks from the vector database. The retrieved chunks are returned to the agent as context.

The `.env.example` file in the repository documents the configurable parameters. The embedding provider defaults to OpenAI (`EMBEDDING_PROVIDER=OpenAI`) using the `text-embedding-3-small` model. Batch size during indexing defaults to 100 chunks per API call (`EMBEDDING_BATCH_SIZE=100`); the README notes that increasing this value reduces indexing time for providers that support higher throughput.

The provider list in `.env.example` includes OpenAI, VoyageAI, Gemini, and Ollama. Selecting Ollama removes the OpenAI dependency and runs the embedding model locally, though this requires a running Ollama instance and the performance characteristics of local embedding differ from cloud-hosted models. The remaining three options all call external APIs.

## Prerequisites: Zilliz Cloud Account and OpenAI API Key

Before running the install command, two external accounts are required. First, a Zilliz Cloud account to provision the vector database. The README points to the Zilliz Cloud sign-up page and describes where to find the public endpoint URL and API key in the cloud console. Second, an OpenAI API key for the default embedding model. The README notes that OpenAI keys always start with `sk-`.

Both are paid services with usage-based pricing. The README positions the external vector database as a cost-reduction mechanism: loading entire directories into context on every request costs tokens at scale, and targeted retrieval keeps per-request context smaller. However, this exchanges per-request token costs for a standing infrastructure cost covering the Zilliz Cloud subscription and OpenAI embedding API calls during the indexing phase.

The Node.js runtime must be version 20.0.0 or newer. No local build step is required since the server runs via `npx`. No other local dependencies are needed for the default configuration.

A VS Code extension (`zilliz.semanticcodesearch`) is separately available in the VS Code Marketplace for users who prefer a GUI-based workflow over the command line.

## Registering the MCP Server in Claude Code

The primary install path for Claude Code is a single command that registers the MCP server and injects credentials as environment variables:

```bash
claude mcp add claude-context \
  -e OPENAI_API_KEY=sk-your-openai-api-key \
  -e MILVUS_ADDRESS=your-zilliz-cloud-public-endpoint \
  -e MILVUS_TOKEN=your-zilliz-cloud-api-key \
  -- npx @zilliz/claude-context-mcp@latest
```

This registers a new MCP server named `claude-context` in the Claude Code configuration, sets the three required environment variables that the server reads at startup, and specifies `npx` as the launch mechanism. Once registered, Claude Code starts the MCP server at the beginning of each new session.

The three required environment variables are `OPENAI_API_KEY` (or the key for the chosen embedding provider), `MILVUS_ADDRESS` (the Zilliz Cloud public endpoint), and `MILVUS_TOKEN` (the API key). The README links to the Claude Code MCP documentation for details on listing and removing registered servers.

## Embedding Provider Configuration via the Environment File

Beyond the inline environment variables in the `claude mcp add` command, a persistent configuration can be written to `~/.context/.env`. The `.env.example` file documents all available variables. The `.env.example` explicitly warns against placing the configuration file in the codebase directory, noting it may conflict with the project's own environment variables.

The main configurable settings are the embedding provider and model (`EMBEDDING_PROVIDER` and `EMBEDDING_MODEL`), the batch size for indexing (`EMBEDDING_BATCH_SIZE`), and the provider-specific API keys. For VoyageAI the key variable is `VOYAGEAI_API_KEY`; for Gemini it is `GEMINI_API_KEY`; for Ollama no API key is needed but a base URL may be set via `OPENAI_BASE_URL` (Ollama exposes an OpenAI-compatible API).

For Codex CLI, the equivalent configuration goes in `~/.codex/config.toml` using a TOML format:

```toml
[mcp_servers.claude-context]
command = "npx"
args = ["@zilliz/claude-context-mcp@latest"]
env = { "OPENAI_API_KEY" = "your-openai-api-key", "MILVUS_TOKEN" = "your-zilliz-cloud-api-key" }
startup_timeout_ms = 20000
```

The README notes that the Codex CLI config file uses `mcp_servers` as the top-level key rather than `mcpServers`, which is a difference from most other clients.

## Using Claude Context with Cursor, Windsurf, Gemini CLI, and Other Clients

The MCP server works with any client that supports the Model Context Protocol. The README documents configurations for seven clients beyond Claude Code. Gemini CLI uses a JSON settings file at `~/.gemini/settings.json` with an `mcpServers` key. Qwen Code uses the same JSON structure at `~/.qwen/settings.json`. Cursor accepts a global `~/.cursor/mcp.json` or a project-level `.cursor/mcp.json`. Windsurf uses a JSON MCP settings file accessible through its settings UI. Claude Desktop on macOS uses a configuration file at the standard application support path.

All these clients use the same npx command (`npx @zilliz/claude-context-mcp@latest`) and the same three environment variables. The structural difference across clients is only in the configuration file path and JSON key naming conventions. The `-y` flag appears in some client configurations to skip npx confirmation prompts.

The monorepo in the repository contains three packages: `@zilliz/claude-context-core` (the indexing and retrieval logic), `@zilliz/claude-context-mcp` (the MCP server), and a VS Code extension (`semanticcodesearch`). All three build from the same TypeScript codebase using pnpm workspaces.

## The Zilliz Cloud Dependency and the Absence of a Local Vector Store Quick Start

The quick start guide requires a Zilliz Cloud account. There is no documented path in the README for using a local or self-hosted vector database through the quick start. Milvus, the open-source project underlying Zilliz Cloud, can be self-hosted, but the README provides no configuration instructions for a local Milvus instance. The `MILVUS_ADDRESS` variable in `.env.example` has a placeholder noting it should be set to the Zilliz Cloud public endpoint, suggesting the design assumes cloud hosting.

Teams with data-residency requirements that prohibit sending source code to external services need to configure Ollama for local embeddings and connect to a self-hosted Milvus instance. Neither setup is documented step by step in the README. The README does not document a `MILVUS_ADDRESS` value format for local instances, though local Milvus typically listens on a different port and protocol than the cloud service.

The repository is not archived. The last recorded push was on 2026-07-14.

## Conclusion

Teams working on large codebases with Claude Code should evaluate Claude Context if they are spending significant context tokens on directory loading each session. The plugin reduces that cost by replacing broad file inclusion with targeted vector retrieval. Before committing to it, confirm that the project's data-sensitivity policy allows source code to be sent to Zilliz Cloud and OpenAI for indexing, since both services receive codebase content. Switching the embedding provider to Ollama and connecting to a self-hosted Milvus instance removes the cloud dependency but requires configuration work the README does not document step by step. The npm package is available as `@zilliz/claude-context-mcp`, currently at version 0.1.15.

## FAQ

### How do I set up Claude Context with Claude Code?

Run the `claude mcp add` command with your Zilliz Cloud endpoint and API key as `MILVUS_ADDRESS` and `MILVUS_TOKEN`, and your OpenAI key as `OPENAI_API_KEY`, pointing to the `@zilliz/claude-context-mcp@latest` package via npx. The README shows the exact command syntax. Node.js 20.0.0 or newer is required.

### How do I add files from my codebase to Claude Context's index?

Claude Context indexes the codebase automatically when the MCP server runs for a session in Claude Code. The indexing process reads source files, generates embeddings using the configured provider, and stores them in the Zilliz Cloud vector database. The batch size and embedding model are configurable through the `~/.context/.env` file.

### How can I give Claude more context about a large codebase?

The Claude Context MCP plugin addresses this by indexing the entire codebase into a Zilliz Cloud vector database and retrieving only semantically relevant code snippets per query. This avoids loading entire directories into context for every request. The README positions this as reducing per-request token cost for large projects.

## Sources

- [Issues](https://github.com/zilliztech/claude-context/issues)
- [License: MIT](https://github.com/zilliztech/claude-context/blob/master/LICENSE)
- [Project website](https://github.com/zilliztech/claude-context/tree/master/docs)
- [README](https://github.com/zilliztech/claude-context/blob/master/README.md)
- [zilliztech/claude-context on GitHub](https://github.com/zilliztech/claude-context)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/zilliztech-claude-context
