# PromptCache: a self-hosted semantic cache proxy for OpenAI, Mistral and Claude traffic

> PromptCache is a Go proxy that sits between your app and an LLM provider, embeds incoming prompts, and can serve a previously cached answer when a new prompt looks close enough. Here is what the repository documents, and where the semantic matching model breaks down.

**messkan/prompt-cache** — Cut LLM costs by up to 80% and unlock sub-millisecond responses with intelligent semantic caching.A drop-in, provider-agnostic LLM proxy written in Go with sub-millisecond response

- Repository: https://github.com/messkan/prompt-cache
- Website: https://messkan.github.io/prompt-cache
- Stars: 409 · Forks: 43
- Language: Go
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/messkan-prompt-cache

## The repeated-prompt problem PromptCache targets

Production LLM traffic is rarely as varied as it looks. The README names three patterns it was built around: RAG applications that keep answering recurring internal questions, agents that re-run similar reasoning or tool-use steps, and support bots fielding near-identical customer questions. In each case the upstream provider is paid to regenerate an answer the system has already produced.

PromptCache is aimed at the team that owns that traffic and wants the saving without changing application code. It presents an OpenAI-compatible surface, so an existing SDK only needs a new base_url. It is not a general vector database, not an orchestration framework, and not a gateway that adds routing or fallback logic. Its scope is one thing: intercept a chat completion request, decide whether a stored response is close enough to reuse, and either return it or forward the request.

## Two thresholds, a gray zone, and a verifier model

The matching design is the part worth understanding before deployment. PromptCache embeds the incoming prompt and compares it against stored embeddings, then sorts the result into three bands using two environment variables.

At or above CACHE_HIGH_THRESHOLD (default 0.70) the request is a direct hit and the cached response is returned. Below CACHE_LOW_THRESHOLD (default 0.30) it is a clear miss and the request goes upstream. Between the two lies the gray zone, where ENABLE_GRAY_ZONE_VERIFIER (default true) sends the pair to a smaller model to decide whether the prompts are actually interchangeable. The README is explicit that the higher threshold must stay above the lower one.

The verifier is the honest part of the design. Semantic similarity is a distance in embedding space, not a statement about meaning, and the project says so: matching is probabilistic and does not guarantee two prompts are interchangeable. Turning the verifier off cuts provider calls and latency, and the README notes it may also reduce matching accuracy. That is a real trade-off, not a footnote.

Storage is BadgerDB, an embedded key-value store, which is why the deployment is a single container with a volume rather than a cluster. The Dockerfile runs the binary as a non-root user and grants group ownership to the root group so the image also runs under an arbitrary UID, a detail aimed at OpenShift-style security contexts.

## Installing PromptCache with docker-compose and making a first cached call

The README gives two install paths. The Docker route clones the repository, exports the provider and key variables, and brings up the compose file, which maps port 8080 and mounts ./badger_data into the container.

```bash
git clone https://github.com/messkan/prompt-cache.git
cd prompt-cache

export EMBEDDING_PROVIDER=openai
export OPENAI_API_KEY=your_key_here

# Strongly recommended for every non-local deployment
export API_AUTH_TOKEN=your-secret-token

docker-compose up -d
```

The compose file defaults EMBEDDING_PROVIDER to openai and passes through MISTRAL_API_KEY, ANTHROPIC_API_KEY and VOYAGE_API_KEY as well, so the same file works for the other providers once the matching key is present. If you prefer to build locally, the repository also offers a script and a Make target:

```bash
./scripts/run.sh
# or
make run
```

With the server listening on 8080, point an OpenAI-compatible client at it. The only change from a normal call is base_url:

```python
from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8080/v1",
    api_key="your-openai-api-key",
)

client.chat.completions.create(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "Explain quantum physics"}],
)
```

The first request is forwarded to the configured provider. A later request judged sufficiently similar may be served from cache, and the README describes this as the expected behaviour rather than a guarantee. To confirm the cache is doing anything, query the stats endpoint with the token you set:

```bash
curl http://localhost:8080/v1/stats \
  -H "Authorization: Bearer your-secret-token"
```

Provider selection is a single variable. For Mistral, set EMBEDDING_PROVIDER=mistral and MISTRAL_API_KEY; the defaults are mistral-embed for embeddings and mistral-small-latest for verification. For Claude, set EMBEDDING_PROVIDER=claude with ANTHROPIC_API_KEY and VOYAGE_API_KEY, since embeddings come from voyage-3 through Voyage AI while verification uses claude-3-haiku-20240307. The README states that provider names identify compatibility only and that the project is not affiliated with any of these vendors.

## Streaming, cache warming and runtime configuration in v0.4.0

The v0.4.0 release notes list four additions: Bearer-token authentication for management endpoints, SSE streaming support, runtime threshold configuration, and cache warming. Each has a caveat worth reading.

Streaming works in both directions. On a miss, PromptCache forwards the provider stream and buffers the assembled response so it can be cached. On a hit, it synthesizes OpenAI-compatible SSE chunks from the stored response. The buffering step means a streamed miss is held in memory and written to BadgerDB after assembly, which is a different resource profile from a pass-through proxy.

Runtime configuration means thresholds can be changed without a restart, which matters because a threshold that suits a support bot will not suit an agent. The README does not document rollback for a bad threshold change, so treat a runtime edit as a live experiment on production traffic.

Authentication covers the management surface: /metrics, /v1/stats, /v1/config, /v1/config/provider, /v1/cache and /v1/cache/warm. If API_AUTH_TOKEN is unset, management authentication is disabled and PromptCache logs a warning. The inference endpoint /v1/chat/completions is deliberately outside that scheme, and the README says to put PromptCache behind the authentication and authorization controls your application requires.

## Where semantic caching is the wrong tool

The most important limitation is stated by the project itself: semantic similarity is not an authorization mechanism, and cache matching must not be used as a security boundary between users, tenants or authorization scopes. If two users can produce similar prompts, a shared cache can hand one user a response generated for another. The RESPONSIBLE_USE.md file is referenced for sensitive and multi-user data, and that reference is the right place to start, not an afterthought.

Retention is the second constraint. Prompts, responses, embeddings and metadata can be persisted in BadgerDB, with a TTL that defaults to 24 hours. The README notes that a TTL does not replace a data-retention or privacy policy, and that the BadgerDB directory and its backups need the same protection as any other store of application data. A 24 hour default is convenient and also means yesterday's prompts are still on disk unless you change it.

Third, the matching model assumes prompts repeat. A workload of unique, long, context-heavy requests will produce misses, pay the embedding call, and still forward upstream, so the cache adds a hop without saving a generation. The README's own performance table is qualified: results depend on provider, model, prompt distribution, hit rate, hardware and configuration, and the benchmark figures printed there are labelled examples rather than guarantees. If your traffic has no repetition, this is the wrong layer to optimize.

## PromptCache compared with provider-side prompt caching

The closest alternative is the prompt caching built into the providers themselves, and the difference is architectural rather than a matter of degree. Provider-side caching, as documented by OpenAI, Anthropic and others, keeps a prefix of your prompt warm inside the vendor's infrastructure and discounts the input tokens that match. The cache lives on their side, keyed by their rules, and you cannot inspect or move it.

PromptCache inverts that. The cache lives in your BadgerDB directory, in your container, keyed by embedding similarity rather than exact prefix match. It can serve a response for a prompt that was never sent verbatim, which provider prefix caching does not attempt, and it works the same way across OpenAI, Mistral and Claude behind one endpoint. The cost is operational: you run the store, you own the retention policy, and you accept the probability that a near-match is wrong. There is also a hybrid worth naming: the two are not mutually exclusive, since a PromptCache miss still goes upstream where provider-side caching may apply.

The other neighbours are local inference caches such as the KV cache in llama.cpp, which reuse attention state within a running model rather than storing finished responses. That is a different layer entirely and does not transfer to a hosted API.

## Maintenance, licence and what a PromptCache upgrade costs

The repository is not archived, and the last push was on 2026-08-18. Releases are spaced months apart: v0.2.0 on 2025-12-28, v0.3.0 on 2026-01-18, and v0.4.0 on 2026-04-24. The v0.4.0 notes describe additive features (auth, streaming, runtime thresholds, cache warming) rather than a rewrite, which suggests upgrades are not disruptive, but the repository does not document a migration path between versions and the README does not document rollback. One detail to watch: the Makefile still declares VERSION=0.3.0 while v0.4.0 is the current release, so anything reading that variable will report the older number.

The licence is MIT, which permits commercial use, modification and redistribution provided the copyright notice and permission notice are included. That is a permissive licence with no copyleft obligation on your own code. It says nothing about the data you put in the cache, and the provider keys you configure remain governed by your agreements with those providers. Nothing here is legal advice; the LICENSE file is the authoritative text.

Ongoing cost is mostly operational: a container, a persistent volume for BadgerDB, and embedding calls for every request that reaches the matching stage, including the ones that miss. The verifier adds a small-model call for gray-zone prompts. Those embedding and verification calls are the price of admission and should be counted against whatever the cache saves.

## Conclusion

Adopt PromptCache when your workload has genuinely repeated prompts (internal RAG questions, support bots, agent loops) and you can put it behind your own authentication layer. Do not adopt it as a tenant isolation boundary: the README states plainly that semantic similarity is not an authorization mechanism, and the inference endpoint is not an application-level authorization system. Before you route real traffic, verify three things yourself: that CACHE_HIGH_THRESHOLD and CACHE_LOW_THRESHOLD match your prompt distribution, that API_AUTH_TOKEN is set so the management endpoints are not open, and that your retention policy covers the prompts and embeddings BadgerDB will persist on disk.

## FAQ

### What is a prompt cache?

In PromptCache it is a store of previous LLM responses that sits between your application and the provider. The proxy embeds an incoming prompt and can return a stored response instead of making another model call.

### How do I enable prompt cache in PromptCache?

Start the server with a provider selected, for example EMBEDDING_PROVIDER=openai plus OPENAI_API_KEY, then point an OpenAI-compatible client at http://localhost:8080/v1. Caching is on by default; CACHE_HIGH_THRESHOLD, CACHE_LOW_THRESHOLD and ENABLE_GRAY_ZONE_VERIFIER tune the matching.

### Is prompt caching good?

The README frames it as a trade-off rather than a win by default: a hit avoids an upstream generation and lowers latency for that request, but semantic matching is probabilistic and does not guarantee two prompts are interchangeable. It also warns that results depend on the provider, model, prompt distribution, hit rate, hardware and configuration.

### What is prompt caching in Claude?

In PromptCache, Claude support means setting EMBEDDING_PROVIDER=claude with ANTHROPIC_API_KEY and VOYAGE_API_KEY, because embeddings come from voyage-3 through Voyage AI while verification uses claude-3-haiku-20240307. The README states the provider name identifies compatibility only and that the project is not affiliated with Anthropic.

### What is prompt cache hit rate in PromptCache?

A hit is a request whose similarity to a stored prompt reaches CACHE_HIGH_THRESHOLD, default 0.70, at which point the cached response is returned and no upstream generation is made. Prompts scoring between CACHE_LOW_THRESHOLD and the high threshold go to the gray-zone verifier when ENABLE_GRAY_ZONE_VERIFIER is true.

## Sources

- [License: MIT](https://github.com/messkan/prompt-cache/blob/main/LICENSE)
- [messkan/prompt-cache on GitHub](https://github.com/messkan/prompt-cache)
- [Project website](https://messkan.github.io/prompt-cache)
- [README](https://github.com/messkan/prompt-cache/blob/main/README.md)
- [Releases](https://github.com/messkan/prompt-cache/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/messkan-prompt-cache
