Model or dataset
airweave-ai/airweave avatar
airweave-ai/airweave

Airweave: a self-hosted retrieval layer that keeps agent context in sync

Project brief: Open-source context retrieval layer for AI agents. Open-source context retrieval layer for AI agents and RAG systems.

6,560 stars825 forksPythonMIT

At a glance

What is it?
Airweave connects to apps, databases and documents, syncs them continuously, and answers agent queries from one search interface. It is MIT licensed and runs from a single start.sh, but the stack is heavy and the last push was on 2026-06-05.
Who is it for?
Adopt Airweave if you run several agents over the same SaaS and database sources and do not want each one to carry its own connector code, and if you can operate Docker Compose or Kubernetes with PostgreSQL, Vespa, Temporal and Redis. Do not adopt it if one small static corpus is enough, or if you need a project whose last commit is recent: the last push was on 2026-06-05.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
No. The owners have archived the repository on GitHub, so it is read-only and no longer receives changes.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem Airweave targets: context plumbing rebuilt per agent

Every agent that needs internal knowledge ends up with the same private pipeline: OAuth handling for each SaaS tool, a scheduler that re-pulls changed records, chunking and embedding, a vector store, and a query endpoint. The README describes the project as sitting "between your data sources and AI systems as shared retrieval infrastructure" and handling authentication, ingestion, syncing, indexing and retrieval. That is the whole pitch, and it is a real category of work rather than a feature.

The intended user is a team running more than one agent or more than one integration over the same sources. A single developer with one PDF folder is not the audience; the README's own framing is shared infrastructure, which only pays off when the pipeline would otherwise be duplicated. The README claims 50+ integrations and points at docs.airweave.ai/connectors/overview for the list, so the connector inventory is the first thing to check against your own stack.

Connect, sync, query: the four steps Airweave exposes

The README reduces the system to four steps. You connect apps, databases and documents. Airweave syncs, indexes and exposes them through a unified retrieval layer. Agents query that layer through SDKs, a REST API, MCP, or native integrations with agent frameworks. Agents then retrieve grounded context on demand.

The stack behind those steps is named in the README: FastAPI for the backend, React and TypeScript with ShadCN for the frontend, PostgreSQL for metadata, Vespa for vectors, Temporal for orchestration and Redis for pub/sub. That is a five-service minimum before your first query, and it explains why the quickstart shells out to Docker Compose rather than asking you to run a Python process.

The retrieval interface is a collection. The Python SDK example searches a collection by readable_id with a natural-language query, which means the unit of isolation is a named collection you populate from one or more sources. Embedding behaviour is configured rather than fixed: .env.example lists dense embedders (openai_text_embedding_3_small up to 1536d, openai_text_embedding_3_large up to 3072d, mistral_embed at a fixed 1024d, local_minilm at a fixed 384d) and sparse embedders (fastembed_bm25, which needs no API key).

Self-hosting Airweave with start.sh

The README's self-hosted path is three commands, and it requires Docker and docker-compose. Clone the repository, enter it, and run the start script.

bash
git clone https://github.com/airweave-ai/airweave.git
cd airweave
./start.sh

According to the README, the script creates .env from .env.example, generates the required secrets (ENCRYPTION_KEY and STATE_SECRET), starts all services with health checks, and optionally prompts for OpenAI or Mistral API keys. It then tells you to wait for services to become healthy, which the README says may take two to three minutes on the first run, and to verify the app at http://localhost:8080.

Three flags are documented in the README's own troubleshooting prompt: ./start.sh --restart to restart services, ./start.sh --skip-frontend for a backend-only run, and ./start.sh --destroy to clean up everything. If startup fails, the README points at docker logs airweave-backend and docker logs airweave-frontend, and lists port conflicts on 8080, 8001, 5432, 6333, 6379, 7233, 8081 and 8088 as a common cause.

Before starting, check the embedding block in .env.example, because the comment there states that all three variables are required and the app will not start without them.

bash
DENSE_EMBEDDER=openai_text_embedding_3_small
EMBEDDING_DIMENSIONS=1536
SPARSE_EMBEDDER=fastembed_bm25

A first real query uses the Python SDK. The README gives this example, with the API key and collection id replaced by your own values.

python
from airweave import AirweaveSDK

client = AirweaveSDK(api_key="YOUR_API_KEY")
results = client.collections.search.instant(
    readable_id="my-collection",
    query="Find recent failed payments"
)

Install it with pip install airweave-sdk, or npm install @airweave/sdk for TypeScript. There is also a CLI, installed with pip install airweave-cli, that authenticates with airweave auth login and then searches with airweave search "quarterly revenue figures" --collection finance-data. The README notes the CLI prints interactive output in a terminal and clean JSON when piped, which is what makes it usable from a script or another agent.

What Airweave costs you in operations

The honest limitation is the deployment surface. PostgreSQL, Vespa, Temporal, Redis, a backend and a frontend is not a library you drop into an existing process. The README lists Docker Compose for dev and Kubernetes for production, which is a signal that the intended production story is a cluster, not a single VM.

Embedding choice is a second constraint that bites quietly. If you pick local_minilm at 384 dimensions or mistral_embed at 1024, EMBEDDING_DIMENSIONS has to match, and switching embedders later means re-indexing whatever Vespa holds. The .env.example comment that all three embedding variables are required and the app will not start without them means a misconfiguration surfaces as a failed boot rather than a degraded search.

Authentication is off by default: AUTH_ENABLED=false in .env.example, with Auth0 settings (AUTH0_DOMAIN, AUTH0_AUDIENCE, AUTH0_RULE_NAMESPACE) required only if you turn it on. For a local trial that is convenient. For anything reachable from a network it is a decision you have to make deliberately, and the README does not document a built-in alternative to Auth0.

Finally, the last push to the default branch was on 2026-06-05, and the most recent release in the list is v0.9.73 on the same date. The version number is still below 1.0. That does not make the project unusable, but it does mean you should expect interface movement and read the release notes before pinning a version.

Airweave against rolling your own pipeline or a hosted retrieval API

The closest alternative in kind is a hosted retrieval service: you send documents or connect sources through a vendor console, and the vendor runs the index. The difference is where the data and the operational burden sit. Airweave is self-hostable under MIT, so the index lives in your Vespa instance and your PostgreSQL, and the embedding keys stay in your .env. A hosted service removes Temporal, Redis and Vespa from your plate but puts your source credentials and indexed content in someone else's system.

The other alternative is building it yourself: a scheduler, per-source OAuth clients, a chunker, an embedder and a vector store. That is entirely reasonable when you have one or two sources. The point at which it stops being reasonable is when the third agent needs the same Salesforce and Postgres data, because that is when the connector code starts being copied rather than shared. Airweave's bet is that the shared layer is cheaper than the copies.

A narrower comparison is a framework-native retriever. Those usually assume you supply the documents and they handle chunking and search. Airweave's scope is upstream of that: it owns authentication, ingestion and continuous sync, and exposes the result over REST, SDKs and MCP so a framework can query it. If your documents already sit in one store you control and never change, Airweave adds services without adding much.

Licence, upgrade cost and what to check before committing

The repository is MIT licensed, which permits commercial use, modification and redistribution provided the copyright notice and permission notice are retained. That is the most permissive common option and it does not impose a copyleft obligation on your own code. It says nothing about the services Airweave talks to: Vespa, Temporal, PostgreSQL and Redis carry their own licences, and OpenAI or Mistral API usage is billed by those providers. This is a description of the licence text, not legal advice.

Upgrade cost is dominated by re-indexing, not by the code. Releases are frequent and patch-level (v0.9.71, v0.9.72, v0.9.73 within a few weeks), so the practical question is whether a schema or embedding change forces a rebuild of the Vespa index. The README does not document a rollback procedure for a failed upgrade, and it does not describe index migration. Treat a version bump as a change that needs a staging run and a fresh sync.

Before adopting, verify three things in this order. Confirm your sources appear in the connector list at docs.airweave.ai/connectors/overview, since no amount of retrieval quality compensates for a missing connector. Confirm your embedding choice and EMBEDDING_DIMENSIONS agree, and that the matching API key is set. Then confirm your target environment can run Vespa and Temporal alongside your existing services, because that, not the Python code, is what determines whether this fits.

Editorial conclusion

Adopt Airweave if you run several agents over the same SaaS and database sources and do not want each one to carry its own connector code, and if you can operate Docker Compose or Kubernetes with PostgreSQL, Vespa, Temporal and Redis. Do not adopt it if one small static corpus is enough, or if you need a project whose last commit is recent: the last push was on 2026-06-05. Verify first that your sources are in the connector list, that your embedding choice matches EMBEDDING_DIMENSIONS, and that your deployment target can hold the Vespa and Temporal services.

Frequently asked questions

What does Airweave do?

It connects to apps, tools and databases, syncs their data continuously, and exposes it through a unified search interface that AI agents query. The README describes it as a context retrieval layer sitting between your data sources and your AI systems.

What is Airweave AI?

Airweave is the open-source project at airweave-ai/airweave, an MIT-licensed retrieval layer for AI agents and RAG systems. Agents query it through SDKs, a REST API, MCP, or framework integrations to get grounded context from multiple sources in one request.

Is Airweave legit?

The repository is public, MIT licensed, not archived, and its most recent listed release is v0.9.73 on 2026-06-05. The version is still below 1.0, so evaluate it against your own sources rather than treating it as a finished product.

What is Airweave made of?

The README names the stack: FastAPI for the backend, React and TypeScript with ShadCN for the frontend, PostgreSQL for metadata, Vespa for vectors, Temporal for orchestration and Redis for pub/sub, deployed with Docker Compose in dev and Kubernetes in production.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
  4. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/airweave-ai-airweave.svg)](https://hysenlabs.com/projects/airweave-ai-airweave)