# Cognee: self-hosted memory for AI agents, and where it stops

> Cognee is an Apache-2.0 Python library that turns ingested documents into a knowledge graph plus vector index so agents can recall context across sessions. It installs with pip, needs an LLM key, and its own README warns that its defaults trade latency for memory quality.

**topoteretes/cognee** — Cognee is the open-source AI memory platform for agents. Give your AI agents persistent long-term memory across sessions with a self-hosted knowledge graph engine.

- Repository: https://github.com/topoteretes/cognee
- Website: https://www.cognee.ai
- Stars: 31,221 · Forks: 3,133
- Language: Python
- License: Apache-2.0
- Published: 2026-08-04 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/topoteretes-cognee

## What Cognee is for, and who actually needs it

A chat agent forgets. Every new session starts from an empty context window, so the user re-explains their project, their preferences and the documents they already uploaded. Cognee's answer is to move that context out of the prompt and into a store the agent queries: a knowledge graph built from your documents, paired with vector embeddings, both hosted on your own infrastructure.

The README frames the audience in three groups. Teams that want a company brain, meaning internal documents unified in one searchable place. Teams running agents that are supposed to learn from feedback and share knowledge across agents. And teams with trust requirements, since the README lists agentic user and tenant isolation, traceability, an OTEL collector and audit traits. That last group is the interesting one. If you are building a single-user demo, a vector store and a long system prompt will do, and Cognee adds an LLM call you did not need. The library earns its place when the same agent must serve several tenants and you cannot ship their documents to a managed memory service.

## The four operations and what happens behind them

The public surface is deliberately small. The README describes four operations: remember, recall, forget and improve. remember runs add, cognify and improve together, which is the whole ingestion path in one call. recall queries the store and, per the README, auto-routes to pick the best search strategy rather than making you choose between graph traversal and vector similarity. forget deletes. improve is the self-tuning step.

Underneath, the library keeps a relational store for metadata and a graph store for relationships. The docker-compose file shows the default relational provider is embedded SQLite, with a postgres profile available by setting DB_PROVIDER=postgres and DB_HOST=postgres. The Dockerfile installs optional dependency groups by name, and the list it syncs is a fair map of the supported graph and vector backends: neo4j, postgres, llama-index, fastembed, ollama, mistral, groq, anthropic and aws. So the architecture is not one engine but a pluggable pipeline: an LLM extracts entities and relationships from your text, those land in a graph, embeddings land in a vector index, and recall combines the two.

Two flags change that pipeline's cost profile. AUTO_FEEDBACK defaults to on and, per the README, makes one LLM call after each answered query so memory self-tunes. CACHING, also on by default, backs session memory; turning it off disables remember(session_id=...) entirely. The README is blunt that anyone benchmarking Cognee should leave CACHING on, because switching it off measures the system with its memory layer removed.

## Installing Cognee and running a first remember and recall

Cognee needs Python 3.10 to 3.14 and an LLM API key. The README shows installation through uv, though pip, poetry and other Python package managers work the same way.

```bash
uv pip install cognee
```

Configuration is environment based. The .env.example file is marked deprecated in favour of .env.template, and the README gives the direct route for a quick check:

```python
import os
os.environ["LLM_API_KEY"] = "YOUR OPENAI_API_KEY"
```

The environment template also names LLM_MODEL and LLM_PROVIDER, so you can point at a non-OpenAI provider without touching code. With the key set, the smallest useful program stores a fact and reads it back. Note that the operations are asynchronous:

```python
import cognee
import asyncio

async def main():
    await cognee.remember("Cognee turns documents into AI memory.")
    results = await cognee.recall("What does Cognee do?")
    for result in results:
        print(result)

asyncio.run(main())
```

The first remember call is the slow one, because it runs the full extraction pipeline against your LLM. If it returns without error, the fact is in the graph; recall should print it back. There is a CLI for the same three operations, and a local UI:

```bash
cognee-cli remember "Cognee turns documents into AI memory."
cognee-cli recall "What does Cognee do?"
cognee-cli -ui
```

The README adds a caveat to that last command: the MCP server launched by cognee-cli -ui runs inside a Docker container, so you need Docker Desktop, Colima or another OCI-compatible runtime with a working docker CLI. A reader without Docker should expect the UI path to fail even though the Python library itself works fine.

## The defaults are tuned for memory quality, not speed

This is the part most introductions skip. Cognee's README states plainly that its defaults favour memory quality over raw latency, and it names the two knobs that matter. AUTO_FEEDBACK=false removes the post-query LLM call, making reads faster and cheaper while session memory keeps working. CACHING=false disables session memory altogether.

That is an honest disclosure, and it is also a warning about the shape of the system. Every remember call spends tokens on entity and relationship extraction. Every answered query spends tokens again unless you disable feedback. On a high-traffic agent, that second cost is the one people miss, because it scales with questions rather than with documents. The README's instruction to leave CACHING on when benchmarking is worth reading twice: it implies that a plausible-looking configuration produces numbers that describe something other than Cognee.

A third flag appears in the truncated tuning section, DATASET_QUEUE_ENABLED=false, which the README says removes the per-process concurrency guard on datasets and risks file-lock leaks. The naming of a directory in the repository root, working_dir_error_replication, suggests file locking around the working directory has been a real source of trouble. I would treat that flag as off-limits in production.

## Where Cognee is the wrong tool

Cognee is beta software. pyproject.toml classifies it as Development Status 4 - Beta, and the version string in the same file is 1.5.4 while the most recent tagged release listed is v1.5.3.dev1 from 2026-08-26. The last push to main was on 2026-08-26, so the project is moving, but moving projects rename things. If you are building against Cognee, pin a version.

The dependency block is a second constraint. litellm is capped below 1.97.0 with a comment explaining that 1.97.0 fails to build on Python 3.10 and 1.98.0+ needs typing features from 3.11. pydantic-settings excludes a range of versions for a security advisory. These are the marks of a library sitting on top of a fast-moving stack, and they mean your resolver may disagree with Cognee's pins.

Finally, the operational surface is not small. A graph store, a vector index, a relational database and an LLM provider all have to be reachable and configured. If your problem is retrieving passages from a document set, a plain vector database with an embedding model is fewer moving parts and no extraction bill. Cognee is for memory that has to persist and connect, not for search.

## Cognee against a graph-native alternative

The closest comparison in the search data is Cognee versus Graphiti, and the difference is architectural rather than cosmetic. Graphiti is a temporal knowledge graph framework: it models facts with validity intervals, so the graph records when something became true and when it stopped being true. Cognee's README describes ontology generation grounded in cognitive science and a graph that evolves as your knowledge does, but it does not present temporal validity as its organising idea. If your questions are of the form what did we believe in March, a temporal graph is the right primitive.

The other axis is scope. Cognee bundles ingestion, graph construction, vector search, session memory, a CLI, an MCP server, a web UI and a Docker deployment. That is a platform, and it is why the README can offer four verbs as the entire API. A narrower library asks you to assemble those pieces yourself, which is more work and less magic when something breaks. Cognee's own repository layout reflects the breadth: separate directories for the frontend, the MCP server, distributed workers and deployment.

## Licence, upgrades and the cost of staying current

Cognee is Apache-2.0, declared both in the README badge and in pyproject.toml. That permits commercial use and modification, and it includes a patent grant. It does not cover the models you connect: your OpenAI or Anthropic bill is separate and, as noted above, scales with queries when AUTO_FEEDBACK is on. The repository also carries a NOTICE.md and a licenses/ directory, which is where third-party attributions live; check NOTICE.md before redistributing a bundled build rather than assuming the top-level licence tells the whole story. This is a description of the files, not legal advice.

Upgrade cost is driven by the pinning. Because litellm, pydantic-settings and instructor all carry constraints with written justifications, a Cognee upgrade is also a dependency resolution event. The release notes for v1.5.3 mention search relevance and ingestion reliability, and v1.5.2 mentions stability and search improvements, which is the pattern you would expect from a project still finding its footing on retrieval quality. Read the notes for the version you are moving to, and run the pipeline against a copy of your own data before promoting it.

## Conclusion

Adopt Cognee if you want agent memory you host yourself and you are willing to pay for an LLM call on every ingest and, by default, after every answered query. Skip it if you need a stable API today: pyproject.toml still carries Development Status 4 - Beta, and the README's own tuning notes show how easily a benchmark can end up measuring Cognee with its memory layer switched off. Before you commit, verify two things in your own environment: that your Python version sits inside the 3.10 to 3.14 window, and that AUTO_FEEDBACK and CACHING are set the way you intend, because CACHING=false stops remember(session_id=...) from working at all.

## FAQ

### What is Cognee?

Cognee is an open-source AI memory platform for agents. It ingests data in any format and builds a self-hosted knowledge graph plus vector index so agents can recall context across sessions.

### How much does Cognee cost per month?

The library itself is Apache-2.0 and free to self-host. The recurring cost is your LLM provider bill: ingestion spends tokens on extraction, and with AUTO_FEEDBACK left on, Cognee makes one LLM call after each answered query.

### How do I install Cognee?

The README shows uv pip install cognee on Python 3.10 to 3.14, then setting LLM_API_KEY in the environment or in a .env file copied from .env.template.

### How do I use Cognee?

The API exposes four operations: remember, recall, forget and improve. A minimal program awaits cognee.remember(...) to store a fact and cognee.recall(...) to query it, or uses cognee-cli remember and cognee-cli recall from the shell.

### What is cognee ai?

It is the same project: an open-source AI memory platform for agents, described in the README as giving agents persistent long-term memory across sessions through a self-hosted knowledge graph.

## Sources

- [Official documentation](https://www.cognee.ai)
- [Official README](https://github.com/topoteretes/cognee#readme)
- [Project repository](https://github.com/topoteretes/cognee)
- [Release notes](https://github.com/topoteretes/cognee/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/topoteretes-cognee
