# seekdb: an embedded, MySQL-compatible search database for agent state

> seekdb from OceanBase unifies vector, full-text and scalar filtering behind one SQL plan, with copy-on-write database forks for agent sandboxes. The trade-off is that you inherit the OceanBase SQL engine to get it.

**oceanbase/seekdb** — The AI-Native Search Database. Best for agent storage, it unifies vector, text, structured, and semi-structured data into a single engine. This all-in-one database makes agents smarter, easier to run, and more stable.

- Repository: https://github.com/oceanbase/seekdb
- Website: https://seekdb.ai
- Stars: 3,009 · Forks: 351
- Language: C++
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/oceanbase-seekdb

## The problem seekdb targets: agent state that is written and searched at the same time

Most retrieval stacks are assembled from parts. A vector index holds embeddings, a relational store holds metadata, and a full-text engine holds documents. An agent that writes a memory and then immediately searches for it has to cross all three, and the merge usually happens in application code. seekdb's pitch is that this split is the reason agent storage is hard to operate, and that one engine can hold vector, text, structured and semi-structured data together.

The intended user is not a general application developer. The README frames the project as "The State Store for AI Agents" and lists ai-agents, rag, langchain and llamaindex among its topics. That is a narrow audience on purpose: teams building agents that persist memory, tool output or conversation state, then query it back within the same session. If your workload is a nightly batch job that rebuilds an index, the problem seekdb solves is not your problem.

## Change Stream and two-level HNSW: how writes become searchable without blocking

The mechanism the README describes is an async index pipeline called Change Stream. A write commits and returns without waiting for index construction. Separately, the pipeline consumes the redo log and updates a delta HNSW index. Queries then read both the delta index and a snapshot index under fine-grained read locks.

The claim that follows is that P99 latency stays flat under concurrency, because the write path never blocks on the index build. The README reports 1,523 QPS with 21.7 ms concurrent P99 and describes P99 jitter of 1.1x as concurrency rises, against roughly 10x for Elasticsearch and Milvus on the same workload. Those numbers come from the project's own benchmark, and the README points to https://github.com/oceanbase/vdb-streambench for reproduction. Treat them as the project's claim, not an independent result. The architectural detail worth taking seriously is the two-level index split: it is what makes a just-written vector visible to a query without a full index rebuild, and it is the part you would have to verify against your own write rate.

## Installing seekdb and running a first hybrid query

The README's 30-Second Try section gives a single install command. pyseekdb is the Python SDK for seekdb, and the README states that embedded mode runs in-process with no servers, no schemas and no embedding setup.

```bash
pip install -U pyseekdb   # pyseekdb is the Python SDK for seekdb
```

After that, the README says you can switch to server or OceanBase mode with one line, and points to images/demo.py for the demo source. The README does not show that one-line switch inline, so check demo.py or the documentation at docs.seekdb.ai before assuming the exact call.

The query shape below is taken from the README's hybrid search example. Vector distance, full-text match and a scalar filter all sit in one statement, and the ordering uses APPROXIMATE.

```sql
SELECT id, title, l2_distance(emb, '[0.12,0.34,...]') AS dist
FROM docs
WHERE MATCH(content) AGAINST('quarterly report')
  AND author_id = 42
  AND created_at > '2026-01-01'
ORDER BY dist APPROXIMATE LIMIT 10;
```

The point of this example is that the three conditions are pushed into one execution plan rather than merged client-side. What you should see is a single result set ordered by vector distance, already filtered by the text match and the scalar predicate. If you are coming from a stack where you query a vector store and then filter in Python, that is the difference to look for.

## FORK DATABASE sandboxes and what MERGE TABLE actually commits

The second distinctive feature is copy-on-write branching. FORK DATABASE takes a snapshot of an entire database in seconds, and the README states there is no data copy because the mechanism is kernel-level COW rather than application-layer save and restore. An agent can then write to the fork, and the parent either accepts the work with MERGE TABLE or discards it with DROP DATABASE.

```sql
FORK DATABASE agent_state TO agent_sandbox_42;

USE agent_sandbox_42;
INSERT INTO memory (session_id, embedding, content) VALUES (...);

MERGE TABLE agent_sandbox_42.memory INTO agent_state.memory STRATEGY THEIRS;
DROP DATABASE agent_sandbox_42;
```

The merge takes a conflict strategy, and the README names three: FAIL, THEIRS and OURS. That is the interesting design decision. A merge across a fork is not a fast-forward, so the project forces you to declare what happens when the parent and the fork both changed a row. The README does not document rollback semantics for a completed merge, and the repository's fork tests live under tools/deploy/mysql_test/test_suite/fork_table/, so read those before relying on merge behaviour in production.

## Where seekdb is the wrong tool

The clearest limitation is the dependency. seekdb is built on the OceanBase SQL engine, and the README presents it in three shapes: embedded library, single-node server, and the OceanBase distributed cluster. Choosing seekdb means choosing that engine, its storage layout and its operational model. If you wanted a small index you can drop into an existing Postgres or SQLite application, this is a larger commitment than the install command suggests.

The second limitation is scope. The README does not document rollback for MERGE TABLE, and it does not describe what happens to open sessions on a fork when the parent is dropped. If your agent workflow needs to undo a merge rather than discard a fork, that path is not established.

Third, the benchmark figures are self-reported and the README links to the project's own reproduction repository. That is a reasonable practice, but it means the performance case rests on a workload the project chose. If your queries are dominated by large analytical scans rather than the streaming write plus search pattern described, the Change Stream design is not aimed at you.

## seekdb compared with a dedicated vector database

The obvious alternative is a purpose-built vector store such as Milvus, or an embedded index library. The difference in approach is where filtering happens. A dedicated vector store generally performs approximate nearest-neighbour search first and applies metadata filters afterwards, which means either over-fetching candidates or accepting recall loss when a filter is selective. seekdb pushes the scalar predicate and the full-text match into the same plan as the vector ordering, so the filter constrains the search rather than trimming its output.

The second difference is the branching model. Vector stores typically give you collections and snapshots, not a writable fork of the whole database with a defined merge strategy. If your agent needs to try a change and roll it back atomically, that is the feature to compare against, not raw query latency.

The cost of the seekdb approach is that you are running a SQL engine. A dedicated vector store will usually be a smaller process with a narrower API. The README's own comparison is against Milvus and Elasticsearch on QPS and P99 jitter, which is a throughput argument, not a simplicity argument.

## Licence, releases and what upgrading costs you

seekdb is licensed under Apache-2.0, and the repository ships a LICENSE and NOTICE file. That is a permissive licence, but the NOTICE file matters in practice: if you redistribute seekdb, Apache-2.0 requires you to carry the notices forward. This is not legal advice; read the LICENSE and NOTICE yourself if you plan to ship it inside a product.

The release cadence visible in the repository is roughly one minor version every one to three months: v1.2.0 on 2026-04-15, v1.3.0 on 2026-05-25, and v1.4.0 on 2026-08-27. The last push to the default branch was on 2026-09-10. Upgrade cost is not documented in the README, so the thing to check before pinning a version is whether FORK DATABASE and MERGE TABLE semantics stayed stable across v1.3.0 and v1.4.0, since those are the features most likely to change shape while the project is young.

## Conclusion

Adopt seekdb if you are building agents that need vector, full-text and scalar filters in one query and you want embedded or single-node deployment without running a separate vector store. Skip it if your search needs are pure nearest-neighbour over a static corpus, where a dedicated index library is simpler, or if you cannot accept the OceanBase SQL engine as your dependency. Verify first that the FORK DATABASE and MERGE TABLE statements behave as documented on your target version (v1.4.0 is the release listed), and confirm the embedded mode covers the concurrency you expect, since the README describes embedded, single-node server and distributed cluster as three separate deployment shapes.

## FAQ

### Which database is best for searching?

There is no single answer, but seekdb's position is that vector, full-text and scalar search belong in one engine rather than three. The README shows a single SQL statement combining l2_distance, MATCH ... AGAINST and a scalar predicate in one execution plan.

### What is a query database?

In seekdb's case the query path reads two indexes at once: a delta HNSW index holding recently written vectors and a snapshot index holding the rest. The README describes queries hitting both under fine-grained read locks.

### What is a search database?

seekdb is described in the README as an AI-native search database that unifies vector, text, structured and semi-structured data in a single engine. It is MySQL-compatible and can run embedded, as a single-node server, or in the OceanBase distributed cluster.

## Sources

- [License: Apache-2.0](https://github.com/oceanbase/seekdb/blob/master/LICENSE)
- [oceanbase/seekdb on GitHub](https://github.com/oceanbase/seekdb)
- [Project website](https://seekdb.ai)
- [README](https://github.com/oceanbase/seekdb/blob/master/README.md)
- [Releases](https://github.com/oceanbase/seekdb/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/oceanbase-seekdb
