# mcp-server-qdrant: a semantic memory layer for MCP clients

> An official Qdrant MCP server that exposes two tools, qdrant-store and qdrant-find, so an LLM client can write to and search a Qdrant collection. It is a small, focused server, and the documentation is thin in a few places worth knowing before you wire it into an IDE.

**qdrant/mcp-server-qdrant** — An official Qdrant Model Context Protocol (MCP) server implementation

- Repository: https://github.com/qdrant/mcp-server-qdrant
- Website: https://qdrant.tech
- Stars: 1,540 · Forks: 307
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/qdrant-mcp-server-qdrant

## What mcp-server-qdrant actually solves

LLM clients forget. A chat session ends, the context window fills, and the notes you accumulated are gone. mcp-server-qdrant exists to give those clients a place to put information and a way to get it back. The README describes it as an official Model Context Protocol server for "keeping and retrieving memories in the Qdrant vector search engine", acting as a semantic memory layer on top of the Qdrant database.

The audience is narrower than the repository name suggests. It is for people who already have Qdrant running, or are willing to run it, and who use an MCP-capable client such as Claude, Cursor or Windsurf. The server does not manage Qdrant for you, does not create collections through a UI, and does not offer a general query language. It exposes exactly two tools: qdrant-store writes a string plus optional JSON metadata into a collection, and qdrant-find runs a semantic search and returns the stored information as separate messages. Everything else is Qdrant's job.

That narrowness is the point. If you want a memory layer and nothing more, two tools are easier to reason about than a full client. If you want to filter by payload, scroll through a collection, or run hybrid search, this is the wrong layer and you should talk to qdrant-client directly.

## The two tools and how a request flows

The mechanism is straightforward. Your MCP client calls qdrant-find with a query string. The server embeds that query using the configured embedding model, sends a vector search to Qdrant with a limit controlled by QDRANT_SEARCH_LIMIT (default 10), and returns the matching stored information as separate messages. When the client calls qdrant-store, the server embeds the information, attaches the optional metadata, and writes a point into the collection.

Embedding happens inside the server process. The EMBEDDING_PROVIDER variable currently accepts only fastembed, and the default model is sentence-transformers/all-MiniLM-L6-v2. That default matters more than it looks: the model that embedded your stored points must match the model that embeds your queries, or the search returns noise. The README does not warn about this explicitly, but the configuration makes it your responsibility.

Collection selection is conditional. If COLLECTION_NAME is set, the collection_name argument on both tools is disabled and every call goes to that default. If it is not set, the client must pass collection_name on each call. This is a sensible design, though it means a misconfigured server can silently write into the wrong collection if the client always supplies a name.

One constraint is stated plainly: you cannot provide both QDRANT_URL and QDRANT_LOCAL_PATH at the same time. Pick a remote server or a local path, not both.

## Installing mcp-server-qdrant with uvx and running it

The README's preferred path needs no installation step. With uvx, the package is fetched and run in one command. The example sets QDRANT_URL, COLLECTION_NAME and EMBEDDING_MODEL, then starts the server:

```bash
QDRANT_URL="http://localhost:6333" \
COLLECTION_NAME="my-collection" \
EMBEDDING_MODEL="sentence-transformers/all-MiniLM-L6-v2" \
uvx mcp-server-qdrant
```

With no --transport flag the server uses stdio, which is what local MCP clients expect. If the command starts and then appears to hang, that is normal: stdio servers wait for a client to speak. Your MCP client config should point at the same command with the same environment variables.

For a remote client, switch to SSE or streamable-http:

```bash
QDRANT_URL="http://localhost:6333" \
COLLECTION_NAME="my-collection" \
uvx mcp-server-qdrant --transport sse
```

The server listens on port 8000 by default, and FASTMCP_SERVER_PORT changes it. The README gives this override:

```bash
QDRANT_URL="http://localhost:6333" \
COLLECTION_NAME="my-collection" \
FASTMCP_SERVER_PORT=1234 \
uvx mcp-server-qdrant --transport sse
```

If you would rather run it in a container, the repository ships a Dockerfile based on python:3.11-slim. It installs the package with uv, exposes port 8000, sets COLLECTION_NAME to default-collection, and starts the server with SSE:

```dockerfile
FROM python:3.11-slim
WORKDIR /app
RUN pip install --no-cache-dir uv
RUN uv pip install --system --no-cache-dir mcp-server-qdrant
EXPOSE 8000
ENV COLLECTION_NAME="default-collection"
CMD uvx mcp-server-qdrant --transport sse
```

The image sets QDRANT_URL and QDRANT_API_KEY to empty strings as defaults, so you must override them at runtime or the server has nothing to connect to. Note also that the container always runs SSE, so a stdio-only client cannot use this image as-is.

## Read-only mode and the limits you will hit

QDRANT_READ_ONLY defaults to false. Setting it to true disables the qdrant-store tool, leaving only retrieval. For a shared collection where several people point their assistants at the same data, that is the safer default, and the README treats it as a first-class setting rather than an afterthought.

The bigger limitation is scope. There is no tool for deleting a memory, no tool for listing what is stored, and no way to filter by metadata on retrieval even though metadata can be written. If you store a memory with the wrong content, you have to remove it through Qdrant itself. The README does not document a delete path, so plan for that gap before you let a client write freely.

The embedding model is another boundary. Only fastembed is supported as a provider, and the model name is a single string. If your collection was built with a different model or a different vector size, this server will not adapt; it will embed with whatever EMBEDDING_MODEL says and search. There is no dimension check described in the README.

Finally, the tool descriptions are configurable through TOOL_STORE_DESCRIPTION and TOOL_FIND_DESCRIPTION, with defaults in settings.py. That is useful when a client keeps calling the wrong tool, but it also means the behavior an LLM sees is partly your prose, not the server's.

## How it compares with a plain qdrant-client script

The obvious alternative is not another MCP server; it is writing a small script with qdrant-client, which this project already depends on. The difference is the interface. A qdrant-client script gives you full control: filters, payload indexes, scroll, delete, batch upsert, and any embedding model you like. It also gives you nothing for free in an MCP client, because the client has no way to call it.

mcp-server-qdrant trades that control for discoverability. The client sees two named tools with descriptions, and the LLM decides when to call them. That is the whole value proposition, and it is a real one for chat and IDE workflows where you cannot inject arbitrary Python. The cost is that every capability not exposed as a tool is invisible to the model. If your use case needs a metadata filter, you are back to writing code, and at that point the MCP layer is overhead.

A second comparison is with the many community MCP servers for Qdrant. The README does not discuss them, so there is no feature matrix here. What this repository offers is the official label and a small surface area. If you need retrieval plus something else, check whether a community server exposes the extra tool before assuming this one can be extended.

## Maintenance, licence and upgrade cost

The repository is not archived, and the last push was on 2026-09-04, which is recent enough that the project is being touched. The latest release listed is v0.8.1 from 2025-12-10, with v0.8.0 in 2025-06-27 and v0.7.1 in 2025-03-11. The gap between the last push and the last tagged release is worth noting if you depend on tagged versions rather than the default branch.

The dependency pins are the real upgrade cost. pyproject.toml pins fastmcp to exactly 2.7.0, requires pydantic between 2.10.6 and 2.12.0, and requires Python 3.10 or newer. A hard pin on fastmcp means the FastMCP environment variables documented in the README are tied to that version, and the README itself notes that server-specific FASTMCP_SERVER_ settings may change in future versions. Upgrading this package can therefore change your configuration surface, not just your bug list.

The licence is Apache-2.0, declared both in the repository metadata and in pyproject.toml. That is a permissive licence with an explicit patent grant, and it is compatible with commercial use. It is not legal advice, and if you redistribute the server inside a product you should read the licence text in the repository rather than this summary. There is no separate commercial edition mentioned, and no pricing page in the repository, so there is nothing to upgrade to beyond the open source package.

## Before you point a client at it

The failure modes are mostly configuration. Setting both QDRANT_URL and QDRANT_LOCAL_PATH is explicitly disallowed. Leaving QDRANT_URL empty in the Docker image connects you to nothing. Choosing an embedding model that does not match the one used to build the collection produces plausible-looking but wrong results, because vector search has no way to tell you the embeddings are from different spaces. And if COLLECTION_NAME is unset, every client call must pass collection_name, which some clients will not do.

There is also a transport mismatch to watch. The provided Dockerfile runs SSE only. A client that speaks stdio needs the uvx path instead, or a modified image. The README documents stdio, sse and streamable-http, but the image hardcodes one of them.

Given the small tool surface, a reasonable first step is to run the server read-only against a throwaway collection, confirm that qdrant-find returns what you stored through the Qdrant API directly, and only then enable qdrant-store. That sequence catches an embedding mismatch before it becomes a memory layer full of points that never come back.

## Conclusion

Adopt mcp-server-qdrant if you already run Qdrant and want a Claude, Cursor or Windsurf client to keep notes and retrieve them semantically with almost no code. Do not adopt it if you need a general-purpose Qdrant client, hybrid search, or a documented upgrade path, because the README does not cover migrations and the server exposes only two tools. Before wiring it in, verify that your client can pass environment variables, that QDRANT_URL and QDRANT_LOCAL_PATH are not both set, and that the embedding model you choose matches the one used to create the collection.

## FAQ

### What exactly does mcp-server-qdrant do?

It is an official Model Context Protocol server that acts as a semantic memory layer on top of Qdrant. It exposes two tools, qdrant-store for writing information and qdrant-find for retrieving it by semantic search.

### Do I need mcp-server-qdrant?

You need it if you want an MCP-capable client such as Claude, Cursor or Windsurf to store and retrieve notes in Qdrant without writing code. If you only need to query Qdrant from a script, the qdrant-client library this project depends on is enough.

### How is mcp-server-qdrant different from an API?

The README describes MCP as an open protocol for connecting LLM applications with external data sources and tools, and this server implements it with two named tools that the model can call. A plain API would require the client to know the endpoints and call them itself.

### Can mcp-server-qdrant run as a hosted server?

Yes, in the sense that it supports the sse and streamable-http transports, which the README describes as suited to remote clients. The Dockerfile in the repository runs the server with SSE and exposes port 8000.

## Sources

- [License: Apache-2.0](https://github.com/qdrant/mcp-server-qdrant/blob/master/LICENSE)
- [Project website](https://qdrant.tech)
- [qdrant/mcp-server-qdrant on GitHub](https://github.com/qdrant/mcp-server-qdrant)
- [README](https://github.com/qdrant/mcp-server-qdrant/blob/master/README.md)
- [Releases](https://github.com/qdrant/mcp-server-qdrant/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/qdrant-mcp-server-qdrant
