# RushDB: A Graph and Vector Memory Layer for AI Agents

> RushDB accepts arbitrary JSON, infers labels, types and relationships on write, and supports graph traversal combined with server-side vector search on top of Neo4j. Here is how it installs, what it does well, and where it stops.

**rush-db/rushdb** — RushDB is a graph + vector database and memory layer for AI agents. Push any JSON, get typed, searchable, relationship-aware records back — no schema, no migrations. Built on Neo4j.

- Repository: https://github.com/rush-db/rushdb
- Website: https://rushdb.com
- Stars: 324 · Forks: 26
- Language: TypeScript
- License: not declared
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/rush-db-rushdb

## The three-database problem RushDB is aimed at

An agent that remembers things usually ends up with three storage systems: a key-value store for session state, a vector store for semantic recall, and a graph database for relationships. Keeping them consistent is the actual work. A write has to land in all three, and a delete has to be propagated or the vector index starts returning records whose graph edges are gone.

RushDB's pitch is that one write is enough. The README states the project replaces all three and that you query with graph traversal, semantic search, or both in one call. The audience is narrow and specific: developers building agent memory, RAG pipelines, or an app backend who do not want to define a schema up front. If you already have a settled relational model and a DBA, this is not aimed at you.

## How RushDB turns arbitrary JSON into a graph

The mechanism is inference at write time. According to the README, nested objects and arrays of objects become linked records, and labels, types and relationships are inferred on write. There is no migration step, because there is no declared schema to migrate.

The README gives an example where a COMPANY record contains a DEPARTMENT array, which in turn contains an EMPLOYEE array. Each nested key becomes a label, each object becomes a record, and containment becomes a relationship. That is the whole data model: your JSON key names double as labels.

Queries follow the same shape. Related labels go inside where, not in labels. The README is explicit about this, and it is the detail most likely to trip up a first attempt, because putting DEPARTMENT in the labels array looks natural and produces the wrong query. Underneath, the project is built on Neo4j, so traversal depth is a property of the graph engine rather than something RushDB reimplements.

Vector search is separate from the write path. You register a property for embedding once, and the server embeds it on every subsequent write. The README describes this as managed embeddings, server-side, which means no embedder in your process and no vectors array in your payload.

## Installing the SDK and storing your first agent memory

There are two paths. The cloud path uses a managed instance and an API key from app.rushdb.com. The self-host path is a Docker image plus your own Neo4j instance; the README points to a self-hosting section rather than listing compose files inline.

Start with the JavaScript SDK from npm:

```bash
npm install @rushdb/javascript-sdk
```

Or the Python package, published on PyPI as rushdb:

```bash
pip install rushdb
```

Construct a client with your key, then register the property you want embedded. This call is a one-time setup step, not something you repeat per write:

```typescript
import RushDB from '@rushdb/javascript-sdk'

const db = new RushDB('RUSHDB_API_KEY')

await db.ai.indexes.create({ label: 'MEMORY', propertyName: 'output' })
```

After that, writes are plain JSON. No embedder, no vectors array:

```typescript
await db.records.create({
  label: 'MEMORY',
  data: {
    agent_id: 'agent-42',
    session_id: 'sess-001',
    action: 'summarized',
    topic: 'Q4 results',
    output: summaryText
  }
})
```

Recall combines both query types in one request. The vector search targets the embedded property, while where filters on the graph side:

```typescript
const memories = await db.records.vectorSearch({
  labels: ['MEMORY'],
  propertyName: 'output',
  query: 'what did we decide about Q4?',
  where: { agent_id: 'agent-42' },
  limit: 10
})
```

What you should see back is a ranked set of MEMORY records, already narrowed to agent-42 before the semantic ranking matters. If the index was never registered, the embedded property will not exist and the search has nothing to rank against, so register it before the first write rather than after.

## CSV import and the type-suggestion trade-off

Bulk loading goes through importCsv, which takes a label, the raw CSV text, options and a parseConfig. Two options matter. suggestTypes infers numbers, booleans and dates from strings. skipEmptyValues treats blank cells as unset instead of storing empty values, and the README notes that 0 and false are kept, which is the correct behaviour and worth stating because naive blank handling usually destroys falsy values.

The trade-off is that type inference on strings is a guess. A column of zip codes or product IDs that happen to be numeric will be inferred as numbers, and leading zeros will not survive. If your identifiers are numeric-looking strings, the inference is working against you and you should not enable it for those columns.

## Aggregations, and the limit trap the README warns about

select shapes output with $sum, $avg, $count, $min and $max, while groupBy controls the dimensions. The README gives a worked example over PROJECT records returning a single row with totalBudget, avgBudget and projectCount, and a second example grouping by $record.status to get one row per status.

The warning attached to this is the most useful line in the section: do not add limit to an aggregation, because it would scan only the first N records and skew the totals. That is a real failure mode rather than a style note. An aggregation with a limit returns a plausible number that is simply wrong, and nothing in the response tells you so. Ordering is described as late-ordering, applied after the full dataset is aggregated.

Aggregations compose with traversal. The README's example computes headcount and payroll per department by aliasing the related EMPLOYEE records. That combination, aggregate over a traversed relationship, is the part a plain document store cannot do without application-side joins.

## Where RushDB is the wrong tool

Schema inference is the feature and also the constraint. If you need a contract enforced at write time, rejecting a record whose salary is a string, RushDB will not give you that. It infers a type and stores the record. Applications that depend on strict validation at the storage boundary should keep that validation in their own layer, because the database will not do it.

Operating Neo4j is a second boundary. Self-hosting means Docker plus your own Neo4j instance, so the operational surface is a graph database, not an embedded file. Teams without Neo4j experience are taking on a real dependency.

Third, this is not an analytics warehouse. The aggregation support is real, but the README's examples operate on records the same system serves to agents. A workload that is mostly large scans over historical data belongs in a columnar store, not here.

The README also does not document rollback behaviour for imports, and it does not describe how the server-side embedding model is chosen or changed. If either matters to you, treat it as unresolved rather than assume a default.

## How this differs from Neo4j and from a vector store

Direct Neo4j is the obvious alternative, and RushDB is built on it. The difference is who writes the graph. With Neo4j you write Cypher and decide the node labels and relationship types yourself, which gives you exact control and a query language with a long history. With RushDB you push JSON and the labels and edges are derived from your key names. The trade is control for speed of iteration, and it is a real trade: inferred labels are convenient until you want to rename one, at which point you are migrating data rather than editing a schema file.

A dedicated vector store is the other alternative. It handles embeddings well and knows nothing about relationships. RushDB's claim is that a graph filter and a semantic ranking happen in one call, so you do not fetch candidates from one system and re-filter in another. If your queries are pure nearest-neighbour with no structural constraints, a vector store is simpler and you are paying for graph capability you will not use.

## Conclusion

Adopt RushDB if you are building agent memory or an app backend and want graph traversal and semantic search over the same records without designing a schema or running a separate embedding pipeline. Skip it if you need a schema contract enforced at write time, if you cannot operate Neo4j, or if your workload is plain relational reporting. Before committing, verify three things yourself: that the Apache 2.0 licence file referenced by the README badge is present in the packages you depend on, that the aggregation behaviour you need matches the documented warning about limit, and that your deployment target fits the Docker plus Neo4j self-hosting path the README describes.

## FAQ

### What is RushDB used for?

It is a graph and vector database positioned as a memory layer for AI agents and apps. You push JSON, and it infers labels, types and relationships on write, then lets you query with graph traversal, semantic search, or both in one call.

### Does RushDB require a schema or migrations?

No. The README states there is no schema to design and no migration step, because labels, types and relationships are inferred when records are written. Nested keys become labels and containment becomes a relationship automatically.

### How do I install RushDB?

For the cloud path, install the JavaScript SDK with npm install @rushdb/javascript-sdk or the Python package with pip install rushdb, then construct a client with an API key. Self-hosting uses Docker plus your own Neo4j instance, as described in the README's self-hosting section.

### Can I self-host RushDB?

Yes. The README lists a self-host path alongside the managed cloud option, and describes it as Docker plus your own Neo4j instance. The README does not inline the compose files; it points to the self-hosting section.

## Sources

- [Issues](https://github.com/rush-db/rushdb/issues)
- [Project website](https://rushdb.com)
- [README](https://github.com/rush-db/rushdb/blob/main/README.md)
- [Releases](https://github.com/rush-db/rushdb/releases)
- [rush-db/rushdb on GitHub](https://github.com/rush-db/rushdb)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/rush-db-rushdb
