# Hindsight stores agent memory in PostgreSQL, and reflect is the extra step

> Hindsight is a self-hostable memory server for agents, reached over HTTP on port 8888 with a UI on 9999. Its three calls, its provider list, and its sampling defaults are unusually well specified. Its authentication story is not in the quick start.

**vectorize-io/hindsight** — Hindsight: Agent Memory That Learns

- Repository: https://github.com/vectorize-io/hindsight
- Website: https://hindsight.vectorize.io/
- Stars: 42,564 · Forks: 5,735
- Language: Python
- License: MIT
- Published: 2026-09-15 · Updated: 2026-09-15 · Language: en
- Canonical page: https://hysenlabs.com/projects/vectorize-io-hindsight

## retain, recall, and reflect are three calls against a bank_id

Everything a client does hangs off a `bank_id`, and there are three operations against it. `retain` stores content, `recall` searches what is stored, and `reflect` returns a disposition-aware response rather than raw matches.

```python
from hindsight_client import Hindsight

client = Hindsight(base_url="http://localhost:8888")

# Retain: Store information
client.retain(bank_id="my-bank", content="Alice works at Google as a software engineer")

# Recall: Search memories
client.recall(bank_id="my-bank", query="What does Alice do?")

# Reflect: Generate disposition-aware response
client.reflect(bank_id="my-bank", query="Tell me about Alice")
```

That third call is where Hindsight separates itself from a plain retrieval setup. In a RAG system you embed chunks yourself and paste the retrieved text into a prompt. Hindsight adds a generation step that turns stored material into an answer, and the configuration exposes internal operations named `verification`, `retain`, `reflect`, and `consolidation`, so there is work happening behind the three public calls that you never invoke directly. The Node client exposes the same retain and recall shape, and the contents index also lists memory types, observations, and mental models with knowledge pages, which are storage concepts the quick start does not define.

## Four install paths, and the pip one never names a database

The recommended path is a container, and the documented command publishes two ports and mounts a named volume at a dot-directory inside the image.

```bash
export OPENAI_API_KEY=sk-xxx

docker run -it --pull always --name hindsight --restart unless-stopped -p 8888:8888 -p 9999:9999 \
  -e HINDSIGHT_API_LLM_API_KEY=$OPENAI_API_KEY \
  -v hindsight-data:/home/hindsight/.pg0 \
  ghcr.io/vectorize-io/hindsight:latest
```

The API answers on `http://localhost:8888` and the UI on `http://localhost:9999`. The volume path `/home/hindsight/.pg0` is where the embedded database keeps its files, which is worth knowing before you try a bind mount to a path of your own choosing.

The compose path is the one that asks for a database password, `HINDSIGHT_DB_PASSWORD`, before `docker compose up` in `docker/docker-compose`. The Helm path sets `postgresql.enabled=true` alongside the provider and key. The pip path sets only an LLM credential and starts the server:

```bash
pip install hindsight-api
export HINDSIGHT_API_LLM_API_KEY=sk-xxx

hindsight-api
```

No storage backend is named on that route, and `HINDSIGHT_API_LLM_PROVIDER` is not set either, although `.env.example` lists it as required with `HINDSIGHT_API_LLM_MODEL=gpt-4o-mini`. The README sends readers to the installation guide for Windows and air-gapped setups, so the bare-metal storage question is answered there rather than here.

## --pull always against the :latest tag moves your server without you

The container command combines `--pull always` with the image reference `ghcr.io/vectorize-io/hindsight:latest`. Both halves matter. A moving tag means the image named `latest` is not the same artifact tomorrow, and `--pull always` means Docker fetches it on every start rather than reusing what is cached.

The practical effect is that a memory server holding long-lived facts about people can change version while nobody deploys anything, and the two published ports come back on a new build. Nothing in the documented command pins a digest or a version tag. The compose file is the escape hatch, because a compose file can hold a fixed tag, and the Helm chart takes an explicit version at install time.

The release history is worth reading next to that. v0.10.0 shipped on 2026-09-14, v0.10.1 on 2026-09-21, and v0.10.2 on 2026-09-29, with the last push on 2026-09-29. A project cutting a release most weeks and sitting on a 0.x number is one that expects its own interface to move, which is a reasonable thing to expect and an awkward thing to hold under a floating tag.

## Two ports published, one LLM key, and no credential for the API

`-p 8888:8888 -p 9999:9999` binds both the API and the web UI to the host, and Docker publishes to every host interface unless told otherwise. The only key in the command is `HINDSIGHT_API_LLM_API_KEY`, which is a credential for your model provider, not for Hindsight.

So on a machine with a routable interface, anything that can reach port 8888 can call `retain` and `recall` against a bank whose contents are facts about real people. The quick start does not document a token, an API key, or an environment variable that authenticates a client to a self-hosted Hindsight. The one place a client credential appears is the managed path, where the instruction is to point any client at `https://api.hindsight.vectorize.io` with your API key.

`--restart unless-stopped` extends the exposure window, since the container returns after a reboot without anyone deciding to run it again. If you run this on a shared or cloud host, the first decision is not which LLM to pick. It is putting a proxy with real authentication in front of 8888, and remembering that `--restart unless-stopped` will bring the container back without consulting you.

## The vision slot is chunk-scoped, so a screenshot does not tax text retains

`.env.example` carries a second model configuration that has nothing to do with the first. Alongside `HINDSIGHT_API_LLM_PROVIDER`, `HINDSIGHT_API_LLM_API_KEY`, and `HINDSIGHT_API_LLM_MODEL=gpt-4o-mini`, there are commented-out `HINDSIGHT_API_VLM_PROVIDER`, `HINDSIGHT_API_VLM_API_KEY`, `HINDSIGHT_API_VLM_MODEL`, and `HINDSIGHT_API_VLM_BASE_URL` entries.

The comment above them explains the scoping. Only retain chunks that carry an inline attachment use the vision model. Every text-only chunk stays on the retain LLM, so a bank that ingests the occasional screenshot does not need to run or pay for a vision model across all of its text. Leave the block unset and attachments fall back to the retain LLM as before.

That is a cost decision encoded in configuration rather than left to the caller, and it changes how you size the deployment. A workload with mixed content does not need one capable model for everything, and turning the slot on is a per-chunk routing change with no code edit. The provider list for the vision slot is not spelled out in the file beyond the `gemini` example, so check the supported models page before assuming the same 25 providers apply.

## verification=0.0 next to reflect=0.9 names the four internal calls

One environment variable documents more of the architecture than any diagram. `HINDSIGHT_API_LLM_TEMPERATURE` sets a value in `[0.0, 2.0]`, or the string `none` to omit the parameter entirely, which the comment says is required for models that reject explicit temperatures and names Azure gpt-5.5 as the case in point.

The valuable part is the override table that follows. The global value applies to every operation, but per-operation values take precedence, and the defaults are `verification=0.0`, `retain=0.1`, `reflect=0.9`, `consolidation=0.0`. Read as a system description, that is four distinct internal operations, two of them pinned near-deterministic and one of them, reflect, deliberately run hot.

`HINDSIGHT_API_LLM_REASONING_EFFORT` follows the same shape. Accepts `none`, `low`, `medium`, `high`, and `xhigh`, and the comment is explicit that the value is sent as given whatever the model is called, with `none` there to stop a self-hosted reasoning model (vLLM, Ollama, llama.cpp, TGI) from emitting thinking blocks. Unset it and no reasoning parameter goes out at all, leaving each model at its own default. Both variables are compatibility switches for a provider matrix you did not choose, and both are worth setting explicitly rather than inheriting.

## unsafe-best-match lets uv resolve a name from any configured index

The root `pyproject.toml` is a workspace definition, not a package definition. Eight members are listed: `hindsight-all`, `hindsight-api`, `hindsight-api-slim`, `hindsight-all-slim`, `hindsight-dev`, `hindsight-mcp-server`, `hindsight-clients/python`, and `hindsight-embed`. Alongside it, `index-strategy = "unsafe-best-match"` is set with a comment about preventing resolution failures when using a pytorch index together with PyPI.

The flag name is the warning. Best-match resolution across all configured indexes means a package name that exists on more than one index can be satisfied by whichever one uv judges best, which is exactly what you want during a dependency conflict and exactly what you do not want if an untrusted index carries a colliding name. The setting is deliberate and commented, so it is a known trade rather than an accident, but it is a supply-chain decision your environment inherits from this file.

The same root holds three lockfiles, `uv.lock`, `package-lock.json`, and `deno.lock`, next to `ruff.toml` and a `.python-version`. Clients ship separately too: PyPI carries `hindsight-api` and `hindsight-client`, npm carries `@vectorize-io/hindsight-client`, Go installs from `github.com/vectorize-io/hindsight/hindsight-clients/go`, and the CLI comes from `curl -fsSL https://hindsight.vectorize.io/get-cli | bash`. The npm workspaces list `hindsight-clients/typescript` and the uv members list `hindsight-clients/python`, and neither root manifest tracks the Go directory, so that client is versioned outside both.

## LongMemEval is the project's own yardstick, and the data is dated

The performance section is the strongest claim in the repository and it deserves reading carefully. The README calls Hindsight the most accurate agent memory system ever tested, says it reaches top results on the LongMemEval benchmark, and dates its comparison table to January 2026. It adds that Hindsight's numbers were independently reproduced by researchers at the Virginia Tech Sanghani Center for Artificial Intelligence and Data Analytics and by The Washington Post, while other vendors' scores are self-reported.

Two things follow. The benchmark is the project's own selection, published on a site the project controls. And a 0.10.2 server in active development is a moving target underneath a dated table. The alternative the README names is worth keeping in view, because the difference is architectural rather than a score. RAG makes you own the embedding, the chunking, and the prompt assembly, and returns text you paste into a model call.

Hindsight's `reflect` operation performs a generation step over stored material, and `consolidation` and `verification` exist alongside it. That is a different amount of machinery between your code and the model's context, and it is a decision to make on the shape of your application.

## Conclusion

Adopt Hindsight if your agent needs durable facts about a person across sessions and you are willing to operate a server with a database, two published ports, and a 0.x version that can upgrade itself. Do not adopt it if that memory store must be reachable only by authenticated callers, because the self-hosted quick start documents no credential for the Hindsight API itself. Check three things before deploying. Which PostgreSQL the pip path expects, since the bare-metal instructions set only an LLM key. Which model your internal operations will call at the documented default temperatures. And whether a private index can shadow a package name, given index-strategy = "unsafe-best-match" in pyproject.toml.

## FAQ

### How do I install Hindsight?

The recommended route is a container that publishes the API on port 8888 and the UI on 9999, with a named volume at `/home/hindsight/.pg0` for the embedded database. Alternatives are `docker compose up` in `docker/docker-compose`, a Helm chart from `oci://ghcr.io/vectorize-io/charts/hindsight`, `pip install hindsight-api` followed by `hindsight-api`, or the hosted Cloud endpoint.

### Which LLM providers can Hindsight use?

The README states 25 or more, selected through `HINDSIGHT_API_LLM_PROVIDER`. Hosted names include `openai`, `anthropic`, `gemini`, `groq`, `bedrock`, and `vertexai`; fully local options include `ollama`, `lmstudio`, and `llamacpp`; gateways include `litellm` and `litellmrouter`. `openai-codex`, `claude-code`, `cursor`, and `github-copilot` need no API key.

### What does Hindsight store, and how do I read it back?

You pass a `bank_id` to each call. `retain` stores content into that bank, `recall` searches it, and `reflect` returns a disposition-aware response. The documentation also indexes memory types, observations, mental models with knowledge pages, and memory banks as core concepts.

## Sources

- [License: MIT](https://github.com/vectorize-io/hindsight/blob/main/LICENSE)
- [Project website](https://hindsight.vectorize.io/)
- [README](https://github.com/vectorize-io/hindsight/blob/main/README.md)
- [Releases](https://github.com/vectorize-io/hindsight/releases)
- [vectorize-io/hindsight on GitHub](https://github.com/vectorize-io/hindsight)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/vectorize-io-hindsight
