Model or dataset
zilliztech/GPTCache avatar
zilliztech/GPTCache

GPTCache: A Semantic Cache Between Your App and the LLM API

Semantic cache for LLMs. Fully integrated with LangChain and llama_index.

8,204 stars595 forksPythonMIT

At a glance

What is it?
GPTCache sits in front of an LLM client, embeds each prompt, and answers repeats and near-repeats from a local vector store instead of paying for another API call. It is a Python library with a Docker server, an MIT licence, and a warning in its own README that the API may change at any time.
Who is it for?
GPTCache fits teams running high volumes of repetitive or near-duplicate prompts through a Python service, who can accept a cache that may return a semantically similar answer rather than the exact one. It does not fit anyone who needs a frozen API surface, or who cannot tolerate a wrong hit, such as per-user personalised answers or anything where a stale response is worse than a slow one.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 8 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The cost problem GPTCache was built to absorb

Every call to a hosted LLM API costs money and takes time. In a chat product, a large share of that traffic is not new: users retype the same question, rephrase it slightly, or hit a suggestion button that sends a prompt someone already sent. GPTCache exists to intercept those calls before they leave your process. The README frames the goal as storing LLM responses so that repeated and similar requests are served locally, and the setup.py description calls it "a memcache for AIGC applications, similar to how Redis works for traditional applications." That analogy is the clearest statement of intent: it is a cache layer, not a model, not a router, not an observability tool. The audience is Python developers building chat or agent applications on top of OpenAI-style APIs, especially those already using LangChain or llama_index, both of which the project says it is fully integrated with. If your traffic is low or your prompts are all unique, there is nothing here for you to gain.

How the cache decides a prompt has been seen before

The mechanism has four moving parts, and the README's similar-search example wires all of them together explicitly. First, an embedding function turns the incoming prompt into a vector; the example uses the Onnx embedder and reads its dimension to size the vector store. Second, a data manager combines two stores: a CacheBase for the cached answers and a VectorBase for the embeddings, with the example pairing sqlite with faiss. Third, a similarity evaluation decides whether a stored vector is close enough to the new one to count as a hit; the example uses SearchDistanceEvaluation. Fourth, cache.init() ties the three together and cache.set_openai_key() supplies the API key.

The data flow on a request is therefore: prompt in, embed, search the vector store, score the nearest candidates, and either return the stored answer or fall through to the real API and write the new pair back. The exact-match path is the degenerate case of the same pipeline, and the README's first example shows it needs only cache.init() and cache.set_openai_key() with no embedder or data manager configured. The repository layout mirrors this design: examples/ has separate directories for embedding, data_manager, similarity_evaluation, eviction, processor and session, which tells you the project treats each of those as a pluggable slot rather than a fixed choice.

Installing GPTCache and running your first cached call

The README gives the install as a single pip command, and setup.py declares python_requires of 3.8.1 or higher, so check your interpreter before anything else.

bash
python --version
pip install gptcache

The README warns that if a dependency fails to install because your pip is old, you should upgrade pip first with python -m pip install --upgrade pip. The default install is deliberately thin: requirements.txt lists only numpy, cachetools and requests. Features that need more, such as the Onnx embedder or faiss, pull their libraries in automatically when you reference them. That is convenient, and it is also the reason a container built from this library can grow unexpectedly; pin your environment if size matters.

For a first real use, the README's exact-match example is the smallest thing that works. It imports the cache and the OpenAI adapter, initialises, and then sends the same question twice in a loop.

python
from gptcache import cache
from gptcache.adapter import openai

cache.init()
cache.set_openai_key()

question = "what's github"
for _ in range(2):
    response = openai.ChatCompletion.create(
      model='gpt-3.5-turbo',
      messages=[{'role': 'user', 'content': question}],
    )

The README states that on the second iteration the answer comes from the cache without requesting ChatGPT again, so the wall-clock time printed by the surrounding loop should drop sharply on the second pass. Note that this example requires OPENAI_API_KEY to be set in your environment; the README suggests echo $OPENAI_API_KEY to confirm and export OPENAI_API_KEY=YOUR_API_KEY to set it temporarily. To move from exact matching to similar matching, add the Onnx embedder, a sqlite plus faiss data manager, and SearchDistanceEvaluation to cache.init(), as the README's second example does.

The API stability warning is the real adoption risk

The README says plainly that the project "is undergoing swift development, and as such, the API may be subject to change at any time." That sentence should drive your decision more than any feature list. A cache sits in the request path of your application, which means an API change here is not a background annoyance; it is a change to code that every user request touches. The release history reinforces the point: the most recent release listed is 0.1.44 from 2024-08-01, and before it 0.1.43 from 2023-11-28 and 0.1.42 from 2023-09-28. A version still in the 0.1.x line after that span signals the maintainers are not treating the surface as frozen.

There is a second, quieter constraint. The README states that because the number of large models keeps growing and their APIs keep changing, the project no longer adds support for new APIs or models, and it directs users to the get and set API with a demo at examples/adapter/api.py. In practice that means the bundled openai adapter is a starting point, not a guarantee of coverage for whatever provider you use today. If you are on a newer or less common API shape, budget for writing your own adapter against get and set rather than expecting an integration to appear.

When a semantic hit is worse than a cache miss

The failure mode that matters is a false positive: two prompts that are close in embedding space but should not share an answer. "What is the refund window for the Pro plan" and "What is the refund window for the Enterprise plan" are near neighbours by any distance metric and have different correct answers. GPTCache gives you the similarity evaluation slot precisely because this threshold is workload-specific, and the repository ships an examples/similarity_evaluation/ directory to experiment with. The README does not document a rollback path or a way to invalidate a single bad entry after the fact, so if you cache a wrong answer, correcting it is your problem.

That makes GPTCache the wrong tool in several concrete situations. Any response that depends on per-user context, session state or a timestamp should not be served from a shared semantic cache, because the cache key is the prompt text and its embedding, not the user. Anything regulated or audited, where you must be able to say exactly which response a user received and why, is a poor fit for a layer that silently substitutes a stored answer. And if your prompts are genuinely unique, the embedding and vector search you pay for on every request is pure overhead on top of the API call you were going to make anyway.

GPTCache compared with a plain exact-match cache

The obvious alternative is not another semantic cache library; it is the caching you already have. Python's functools.lru_cache, or a Redis keyed on a hash of the prompt string, gives you exact-match caching with no embedding model, no vector store and no threshold to tune. The difference in approach is the whole point of GPTCache: a hash cache can only ever answer the identical string, while GPTCache embeds the prompt and searches for neighbours, so "what's github" and "what is github" can share one API call. That is a real capability, and it is also where the risk lives, because the second query is not identical to the first and the returned answer was generated for a different wording.

If your traffic is dominated by literal repeats, a hash cache is simpler, has no false-positive mode, and adds no model to your dependency tree. GPTCache earns its complexity only when the near-duplicate share is high enough that the extra hits pay for the embedding step, the vector store and the tuning work. The project itself points at this spectrum: the README describes both an exact match cache and a similar search cache as separate examples, which is an admission that the cheap mode is worth having on its own.

Licence, server mode and the cost of keeping up

GPTCache is MIT licensed, which permits commercial use and modification; the LICENSE file sits at the repository root and setup.py points its license field at the MIT text. Nothing in the README or setup.py suggests a copyleft obligation, but the dependency question is separate and worth checking: because optional libraries install automatically when you use a feature, the effective licence of your deployment depends on which embedders and vector stores you actually pull in. That is an engineering check, not a legal opinion, and for a commercial product you should have someone confirm the set you ship.

The upgrade cost is the more practical concern. The README's own note that the API may change at any time means you should pin the version and read the release note before moving, rather than tracking the main branch. The project also ships a server mode: the README mentions a GPTCache server Docker image that lets any language use the cache, and setup.py registers a console script named gptcache_server pointing at gptcache_server.server:main. That is the path to take if your application is not Python, and it moves the version-pinning problem from your app into a container you control. What the README does not describe is how the server handles the same similarity configuration, so read docs/usage.md before assuming the server exposes every knob the library does.

Editorial conclusion

GPTCache fits teams running high volumes of repetitive or near-duplicate prompts through a Python service, who can accept a cache that may return a semantically similar answer rather than the exact one. It does not fit anyone who needs a frozen API surface, or who cannot tolerate a wrong hit, such as per-user personalised answers or anything where a stale response is worse than a slow one. Before adopting, verify the similarity threshold your workload needs by running the examples under examples/similarity_evaluation/ against your own prompt set, and confirm that the packages GPTCache pulls in automatically are ones you are willing to ship.

Frequently asked questions

What is GPTCache and what does it cache?

GPTCache is a Python library that stores LLM responses and returns them for repeated or similar prompts instead of calling the API again. The README describes it as a semantic cache for LLM queries, built to cut API cost and response time.

How does semantic caching work in GPTCache?

An embedding function turns the prompt into a vector, a data manager stores that vector alongside the answer, and a similarity evaluation scores the nearest stored vectors to decide whether the request is a hit. The README's similar-search example configures all three explicitly with Onnx, sqlite plus faiss, and SearchDistanceEvaluation.

When should you not cache with GPTCache?

The README does not give a list of cases to avoid, but the design implies one: the cache key is the prompt text and its embedding, not the user or the session. Prompts whose correct answer depends on per-user state, or workloads where a semantically close but wrong answer is worse than a slow one, are poor fits.

Official sources

  1. License: MIT
  2. Project website
  3. README
  4. Releases
  5. zilliztech/GPTCache on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/zilliztech-gptcache.svg)](https://hysenlabs.com/projects/zilliztech-gptcache)