Model or dataset
spiceai/spiceai avatar
spiceai/spiceai

spiceai/spiceai: a real-time analytics node beside PostgreSQL, MySQL and MongoDB

Add a real-time analytics node to your operational database. Spice is a portable, accelerated SQL query, search, and LLM-inference engine in Rust for data-grounded AI apps and agents.

3,097 stars231 forksRustApache-2.0

At a glance

What is it?
Spice is a Rust SQL query, search and LLM-inference engine you run as a sidecar. Its 2.0 line adds CDC replication from operational databases, and the README claims about two-second freshness with no ETL, no Debezium and no Kafka.
Who is it for?
Adopt Spice if you already run PostgreSQL, MySQL or MongoDB and want analytical SQL next to the application without standing up Debezium and Kafka, and if a single Rust binary or container fits your deployment. Do not adopt it if you need a mature BI layer with a large user community, or if your team cannot operate a second query engine.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Rust, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The operational database problem Spice is aimed at

Analytical queries against a production PostgreSQL, MySQL or MongoDB instance compete with the transactions that keep the application running. The usual escape is a pipeline: change data capture, a message bus, a warehouse. Spice's answer is to put the analytics engine next to the application instead. The README describes adding a real-time analytics node to an operational database by pointing Spice at PostgreSQL, MySQL or MongoDB, after which it maintains a sandboxed, analytics-ready replica using CDC replication. The stated targets are sub-second queries, roughly two-second freshness and zero analytical load on production. The audience is teams building data-grounded AI apps and agents who want SQL, search and inference on localhost without writing pipelines. If your workload is a nightly batch report, this is more machinery than the problem needs.

How the CDC replica and query federation actually fit together

The repository layout makes the architecture legible. Cargo.toml lists crates such as data-connector-api, data_components, runtime-cluster, runtime-datafusion, runtime-checkpoint-postgres and cayenne, so connectors, the DataFusion-based execution layer, cluster coordination and accelerator storage are separate crates rather than one monolith. The README says the engine is built on Apache DataFusion, Apache Ballista, Apache Arrow, Apache Iceberg, Vortex, DuckDB and SQLite, and that a distributed deployment uses Ballista with multi-active schedulers coordinated through object storage. Query path: incoming SQL over HTTP, Arrow Flight, Arrow Flight SQL, ODBC, JDBC or ADBC is planned against the accelerated working set, then delegated to the cluster for the long tail. Ingest path: native CDC from WAL, binlog or change streams, plus DynamoDB Streams, feeds the local replica. The trade-off is that freshness is a replication property, not a query property. The README states roughly two-second freshness; the README does not document what happens to that figure when the source is under heavy write load or the network between source and Spice is slow.

Installing Spice and running a first query

The README points to docs.spiceai.org for documentation and offers a quickstart for a local machine. The Makefile shows the two binaries the project builds: `spice` (the CLI) and `spiced` (the runtime), with `make build` producing both. The Dockerfile builds `spiced` from a Rust base image and copies it into a Debian runtime image, so a container is the other supported route. The README does not give a copy-paste install command for every platform, so check the docs quickstart for the current one before assuming a package manager entry exists.

bash
make build

After `make build` finishes, the release binaries for the CLI and the runtime are in the Cargo target directory. The runtime reads a spicepod definition; the repository root contains a `spicepod.yml` you can use as a starting shape. The README does not document the full schema of that file, so treat the checked-in example and the docs as the source of truth rather than guessing at keys.

bash
make build-cli
make build-spiced

Those two targets build the CLI and the runtime separately, which is faster when you are iterating on one side. To start the runtime from a container instead, build the project Dockerfile with the default feature set; the Dockerfile accepts `CARGO_FEATURES`, `CARGO_NO_DEFAULT_FEATURES` and `RUST_PROFILE` as build arguments, so a slim build is possible without editing the file. The examples directory contains runnable scenarios such as `examples/cdc-debezium-ingest/` and `examples/cosmosdb-connector/`, which is where to look for a working configuration before writing your own.

Where Spice is the wrong tool

Spice is a query and inference engine, not a database of record. Nothing in the README describes it as durable primary storage, and the accelerator is described as a sandboxed working set, which is a cache-shaped role. If you need a system that owns your data, this is the wrong layer. The second limitation is operational surface. Adding Spice means adding a process that holds a connection to your production database, a replication slot or equivalent, and a local copy of your data. The README's claim of zero analytical load on production is about query load; it does not mean the CDC reader is free, and the README does not document rollback, slot cleanup or what happens when the replica falls behind. Third, the connector list is broad but the README names 30+ connectors without enumerating all of them, so verify the specific source you need in the repository before designing around it. Finally, the project is Rust with a large workspace and a Dockerfile that installs a long list of system packages; building from source is not a five-minute exercise.

How Spice differs from DuckDB and from a warehouse pipeline

DuckDB is the closest comparison for the accelerator role. The README describes Spice Cayenne, its Vortex-based accelerator, as generally available and claims 1.5x faster than DuckDB with 3x less memory on TPC-H SF100, along with 100x faster random access versus Parquet. The architectural difference matters more than the numbers: DuckDB is an embedded database you open from a process, while Spice is a long-running runtime that federates across connectors and serves multiple protocols. If your only need is to query a few Parquet files inside a Python process, DuckDB is simpler and the README's own example file `examples/runtime_demo.py` shows Spice being driven from Python rather than replacing it. Against a warehouse pipeline, the difference is where the transformation happens. A Debezium-plus-Kafka pipeline moves changes into a separate store and you query that store; Spice keeps the replica local and serves it from the same process that handles search and inference. The repository even ships `examples/cdc-debezium-ingest/`, which is a fair signal that the project sees Debezium as a source it can read from rather than only a thing to replace.

Maintenance, release cadence and the Apache-2.0 licence

The last push to the default branch was on 2026-09-10, and the most recent release is v2.3.0 from 2026-09-09, following v2.2.1 on 2026-09-02 and v2.2.0 on 2026-08-24. That is a release roughly every one to two weeks across the visible window, and the repository is not archived. Fast cadence cuts both ways: you get fixes quickly, and you also carry upgrade work. The workspace is large, with crates split by connector, accelerator, checkpoint backend and runtime concern, so a version bump can touch several of them at once. Pinning a version and reading the release notes before upgrading is the practical posture. On licensing, the project is Apache-2.0, which permits commercial use and modification and includes a patent grant. Apache-2.0 also requires that you preserve notices and state changes; if you redistribute a modified `spiced`, that obligation follows you. This is a description of the licence, not legal advice, and any redistribution plan should go past whoever handles licensing at your organisation.

Editorial conclusion

Adopt Spice if you already run PostgreSQL, MySQL or MongoDB and want analytical SQL next to the application without standing up Debezium and Kafka, and if a single Rust binary or container fits your deployment. Do not adopt it if you need a mature BI layer with a large user community, or if your team cannot operate a second query engine. Verify first that your database's CDC source is reachable and its retention window is long enough for the initial snapshot, then check that the connectors you need appear in the repository's crates/data-connectors directory before you plan around them.

Frequently asked questions

What is Spice AI used for?

The README describes Spice as a portable, accelerated SQL query, search and LLM-inference engine for data-grounded AI apps and agents, run as a sidecar or as a distributed cluster. Its 2.0 line adds a real-time analytics replica of PostgreSQL, MySQL or MongoDB using CDC replication.

Is Spice AI safe?

The README does not make a security claim of that kind. It lists HashiCorp Vault and Azure Key Vault secret stores, mTLS, read-only API keys and OpenTelemetry observability as enterprise features, and the repository carries a SECURITY.md and a CodeQL workflow, which are the concrete things to review.

Who created Spice AI?

The README does not name the founders or the organisation behind the project. The repository is spiceai/spiceai, the homepage is docs.spiceai.org, and the README links to a Slack workspace and a blog at spice.ai.

what is spice ai

Spice is a Rust runtime that serves SQL, search and LLM inference over your existing data sources, either as a sidecar next to your application or as a multi-node cluster. It federates more than 30 connectors and can accelerate a working set locally for millisecond queries.

Official sources

  1. License: Apache-2.0
  2. Project website
  3. README
  4. Releases
  5. spiceai/spiceai on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/spiceai-spiceai.svg)](https://hysenlabs.com/projects/spiceai-spiceai)