Zleap-AI/SAG review: a knowledge base app built on the SAG retrieval architecture
A new SOTA for RAG — an original retrieval architecture and an open-source knowledge base for humans and agents.
At a glance
- What is it?
- Zleap-AI/SAG packages the SAG retrieval architecture from arXiv:2606.15971 into a local-first knowledge base with a REST API, an OpenAI-compatible endpoint and an MCP server. It is easy to start with SQLite and LanceDB, but the README leaves operational details thin.
- Who is it for?
- Adopt Zleap-AI/SAG if you want one self-hosted service that turns documents into searchable, traceable knowledge for both people and agents, and you are comfortable running SQLite plus LanceDB locally. Do not adopt it if you need multi-tenant access control, a documented production deployment guide, or a project whose README answers operational questions without reading the source.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Zleap-AI/SAG is a knowledge base application, not a retrieval library
The repository ships an application, not just an algorithm. The README describes a pipeline that starts with sources and documents and ends with cited agent answers, and lists the integration surface as a self-hosted REST/OpenAPI service, an OpenAI-compatible chat endpoint, an MCP server, and the zleap-sag Python package. That places it in a different category from a library you import and wire into your own service.
The intended audience is stated plainly: the product is "deliberately local-first and single-user." It starts with SQLite and LanceDB and requires no external database, with a stated path to PostgreSQL/pgvector and other production backends. If your problem is a single person or a small team accumulating documents that need to be searchable and traceable, the scope fits. If your problem is a shared corpus with per-user permissions, the single-user framing is the first thing to test against your requirements.
The capabilities table in the README lists ingestion, search in Fast (vector) and Precise (multi) modes, source tracing back to the original chunk, a knowledge graph view of events and entities, multi-turn chat with clickable citations, and the integration surfaces above. Those are the product's claims, not benchmark results. The benchmark claims live in the paper and in the separate SAG-Benchmark repository.
Events, entities and query-time hyperedges: the mechanism behind the retrieval modes
SAG is presented as a third architecture rather than a combination of dense RAG and GraphRAG. The README rejects the framing of fusing the two, and the paper title is explicit: SQL-Retrieval Augmented Generation with Query-Time Dynamic Hyperedges.
The data model is three lines in the README. A chunk becomes one semantically complete event. The same chunk also becomes multiple indexing entities. The event and those entities form one latent hyperedge. The distinction matters: an event is meant to carry the full meaning of a chunk rather than being split into independent triples, and an entity is described as a lightweight index and expansion point, not a replacement for that meaning. The hyperedge is not pre-built. It is created locally at query time when SQL joins events that share entities around the current query, which is why the README says SAG does not globally maintain those structures.
Offline indexing is described as four steps: parse a document into semantically coherent chunks; extract one event and multiple entities from each chunk in parallel; persist chunks, events, entities and their associations to relational storage; then persist chunk, event and entity representations to vector and full-text indexes. Online retrieval is where the README text given here stops, so the exact ranking logic for the Precise (multi) mode is not documented in the excerpt. The original evidence boundary is stated: selected events map back to source chunks for generation and citation. That is what makes source tracing possible, and it is the property to verify first if citation fidelity is why you are evaluating this.
Installing Zleap-AI/SAG with Docker Compose and running a first search
The repository offers two entry points. The Makefile has install targets for local development, and compose.yaml defines a Docker Compose stack with an api service and a web front end. The Compose file comments state that the default targets personal local use with SQLite and LanceDB and no additional database service, and that data is written to a sagdata volume so rebuilding the container does not lose it.
For the Compose route, the Makefile provides the commands directly:
make compose-config
docker compose up -d --build
make compose-pscompose-config validates the default Docker configuration, and the Makefile describes it as checking SQLite plus LanceDB. After the stack is up, compose-ps shows service status. Ports come from .env.example: WEB_PORT=3000 and API_PORT=8000. The API binds to 127.0.0.1 by default, and the file notes that LAN or server deployments should change BIND_ADDRESS to 0.0.0.0.
For the local development route, dependencies install per app:
make install-api
make install-webThe Makefile runs these in two terminals. install-api creates a virtual environment under apps/api and installs the package in editable mode with the dev extra. install-web runs npm ci under apps/web. Then:
make api
make webThe api target starts uvicorn on sag_api.main:app with --reload on 0.0.0.0 port 8000. The web target runs the Next dev server with -H 0.0.0.0. The README states that LLM and embedding settings can be left empty at startup and configured later in the settings page, and .env.example repeats that the SAG_LLM_* values are first-start defaults that can be changed in the settings page and take effect immediately. Changing the LLM configuration requires a service restart, according to the same file.
Configuration and provider assumptions you inherit on first start
The defaults are opinionated. .env.example sets SAG_LLM_PROVIDER=openai with SAG_LLM_BASE_URL=https://api.302ai.cn/v1 and SAG_LLM_MODEL=qwen3.6-flash, and sets the embedding model to bge-large-en-v1.5 through the same gateway. The file explains that Anthropic or Gemini native access means changing SAG_LLM_PROVIDER to the corresponding value and leaving the base URL empty to use the official endpoint. If you already have a provider relationship, you are replacing these values, not extending them.
Two keys deserve attention. SAG_EMBEDDING_DIMENSIONS must match the embedding model, and the file gives Qwen3-Embedding-8B = 4096 as an example, with 1024 as the default when left empty. SAG_LLM_STRUCTURED_OUTPUT_MODE defaults to auto, which the file says prefers json_schema and falls back to json_object only when the gateway clearly does not support it. Both are the kind of setting that produces confusing failures when wrong rather than a clear error at startup.
There is also a locking switch. SAG_LOCK_LLM_CONFIG defaults to false; setting it to true fixes the generation model configuration for the deployment, which the file says is for operators who need to pin it. Document parsing defaults to SAG_DOCUMENT_PARSER=auto, which prefers MinerU for PDFs and falls back to MarkItDown when MinerU is not configured or fails. MinerU has its own provider switch between 302 and official endpoints.
Where Zleap-AI/SAG is the wrong tool
The single-user, local-first framing is the main boundary. The README states it directly, and the authentication defaults reinforce it: SAG_AUTH_MODE=local means name-based identity on the local machine, while password mode is described in .env.example as the option to enable for external access. SAG_ALLOW_REGISTRATION defaults to false, and the file notes that even in password mode the first credential can still be created, after which registration is closed. That is a workable personal setup and a poor fit for an organization that needs per-user accounts and role separation, because the available text describes no role model.
The security defaults are explicitly development-oriented. compose.yaml sets SAG_SECRET_KEY to a value labeled a known weak key for local quick start and states that the application refuses it when switching to prod. .env.example says production deployments must set a strong random value and gives openssl rand -hex 32 as the generation command. Anyone deploying this beyond a laptop is responsible for that step.
The README does not document rollback, backup and restore procedures, or a migration guide for the archived v1 branch beyond the note that it is no longer maintained. OCTX import and export with integrity validation and failure recovery is mentioned in the changelog for cross-instance migration and backup, but the README excerpt here does not describe the procedure. If your adoption decision depends on a documented restore path, that documentation is not in the README.
Finally, the retrieval claims are benchmark claims. The README says experiments on HotpotQA, 2WikiMultiHopQA and MuSiQue show the best retrieval and end-to-end QA performance on every benchmark, and points to a separate benchmark repository. That is a reason to reproduce, not a reason to assume your corpus behaves the same way.
How it compares with GraphRAG and with a plain vector store
The README's own comparison is the useful one. Traditional dense RAG retrieves chunks mainly by semantic similarity. GraphRAG adds offline graph construction but pays for triple extraction, entity merging, relation normalization, global maintenance and difficult incremental updates. SAG's stated position is that it replaces the choice between those two with its own data model and execution path rather than wrapping both.
The concrete difference is where the graph work happens. GraphRAG builds and maintains a graph ahead of time, which is what makes incremental updates hard. SAG builds hyperedges at query time from SQL joins over events that share entities, so there is no global structure to keep consistent. The cost moves to query execution, and the benefit is that new documents do not require rebuilding a graph. That is the trade-off to evaluate against your update pattern.
Against a plain vector store, the difference is the event and entity layer. A vector store gives you similarity search over chunks. SAG keeps chunks as the evidence boundary but indexes them through events and entities so that relational expansion is available in the same system. If your queries are single-fact lookups, the extra indexing work is overhead. If your queries span multiple documents and need the connecting entities, the architecture is aimed at exactly that case.
Maintenance, release cadence and licence
The repository is not archived, and the last push was on 2026-09-09. Releases are frequent: v1.8.6 on 2026-09-03, v1.8.5 on 2026-09-02 and v1.8.4 on 2026-08-30. The changelog shows a steady stream of additions, including OCTX import and export on August 13, 2026, the @zleap-ai/sag-cli command-line client on July 31, 2026, and the DeepSeek Harness connector @zleap-ai/dsh-sag on August 30, 2026. The July 14, 2026 entry records a complete rewrite on the zleap-sag package with a redesigned UI, and states that the previous version is archived in the v1 branch and no longer maintained.
That rewrite is the main upgrade cost to plan for. The Makefile includes release and release-dry-run targets that create and push stable tags for reviewed public/main, with a required VERSION argument, which suggests the project manages its own release process rather than publishing on every commit. For consumers, that means version pinning is meaningful, but the README does not describe a compatibility policy for the Python package or the API across minor versions.
The licence is MIT, which is permissive and imposes no copyleft obligation on your own code. The README does not discuss third-party service terms. Note that the default configuration points at a hosted gateway for LLM, embedding and MinerU parsing, so your data flow depends on those providers even though the knowledge base itself runs locally. That is a deployment consideration, not a licence one.
Editorial conclusion
Adopt Zleap-AI/SAG if you want one self-hosted service that turns documents into searchable, traceable knowledge for both people and agents, and you are comfortable running SQLite plus LanceDB locally. Do not adopt it if you need multi-tenant access control, a documented production deployment guide, or a project whose README answers operational questions without reading the source. Before committing, verify the LLM and embedding provider configuration in .env.example against your own gateway, confirm that SAG_AUTH_MODE=password behaves the way you expect for remote access, and read the paper at arXiv:2606.15971 to judge whether the retrieval claims match your workload.
Frequently asked questions
How do I install Zleap-AI/SAG and start it locally?
The Makefile provides make install-api and make install-web for local dependencies, then make api and make web in two terminals. For containers, the Makefile provides make compose-config and docker compose up -d --build, with the API on port 8000 and the web front end on port 3000 by default.
Does Zleap-AI/SAG need an external database like PostgreSQL?
No. The README states the product starts with SQLite and LanceDB and requires no external database, and compose.yaml describes the default as targeting personal local use with data written to a sagdata volume. The README also mentions a path to PostgreSQL/pgvector and other production backends, and the repository contains a compose.postgres.yaml file.
What is the difference between the Fast and Precise search modes in Zleap-AI/SAG?
The README lists Fast as vector and Precise as multi, without describing the ranking logic for the Precise mode in the available text. The Dify-specific setting SAG_DIFY_SEARCH_STRATEGY defaults to vector to reduce latency, and .env.example describes multi as entity expansion plus LLM reranking with higher latency.
How do I connect Zleap-AI/SAG to Codex or Claude Code?
The changelog for July 31, 2026 states that the official command-line client @zleap-ai/sag-cli mounts the SAG Knowledge MCP with a single command, sag agent connect codex or sag agent connect claude-code, without copying a JWT or editing configuration files by hand.
Can Zleap-AI/SAG be used by multiple people at once?
The README describes the product as deliberately local-first and single-user. Authentication defaults to SAG_AUTH_MODE=local for name-based local identity, with password mode described as the option to enable for external access, and SAG_ALLOW_REGISTRATION defaults to false. The README does not describe per-user roles or permissions.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/zleap-ai-sag)