Open-source project
spiculedata/saiku avatar
spiculedata/saiku

Saiku: One Cube for Excel, Dashboards and AI Agents

Open-source semantic layer: one cube for Excel (MDX/XMLA), dashboards, and AI agents (MCP). Mondrian + Apache Calcite.

1,325 stars656 forksJavaApache-2.0

At a glance

What is it?
Saiku is an open source semantic layer built on a Mondrian fork and Apache Calcite. It gives the same cube to MDX clients, browser dashboards and LLM agents, and the trade-off is a heavyweight Java stack plus a build that needs GitHub Packages credentials.
Who is it for?
Adopt Saiku if your measures already live in a Mondrian-style cube and you need that same model reachable from Excel, a dashboard and an MCP-capable agent without writing a second semantic definition. Do not adopt it if you want a library you can embed in an existing JVM service, or if you cannot hand out GitHub Packages credentials to every machine that builds the project.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 5 days ago.
What is it written in?
Mainly Java, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem Saiku targets: three consumers, one cube definition

Most analytics stacks end up with the same measures defined more than once. A Mondrian or XMLA cube serves Excel and other MDX clients. A BI tool holds a second copy of the joins and aggregations for its dashboards. Then an AI agent arrives and someone writes a third description of the data, usually as prompt text that drifts from the real schema within weeks.

Saiku's answer is to keep one cube and expose it through several surfaces. The README describes the project as an open-source semantic layer that offers drag-and-drop analysis in the browser, SQL through Mondrian and Calcite, and a typed REST surface so agents can query without ever seeing MDX. The intended audience is a team that already has, or is willing to write, a Mondrian-style cube and wants the browser, the spreadsheet and the agent to read from it.

That framing matters for who should look at this. If your data is a pile of Parquet files and you have no cube, Saiku is not a query engine you point at files. It is a layer over a model you have to define. The repository ships a FoodMart demo cube so you can see the shape before writing your own.

How the Mondrian fork, Calcite planner and Arrow cellsets fit together

The engine is a fork of Mondrian, versioned 4.8.1.x in the README, with Apache Calcite providing a SQL planner alongside the older SqlQuery builder. Calcite is the default path; the legacy builder is still reachable with -Dmondrian.backend=legacy. That flag is the clearest sign this is a migration in progress rather than a finished rewrite, and anyone with existing Mondrian tuning should expect the two backends to behave differently on the same cube.

Results travel as Apache Arrow cellsets. The README's reason is that the browser and any programmatic consumer then share one result envelope rather than each parsing its own JSON shape. The server side is Jetty 12 EE10 with Jersey 3.1, Spring 6 and Spring Security 6.5, launched from a single JAR by Picocli. The SvelteKit 5 front end lives in a separate repository and is served from inside that same JAR under /ui/.

The AI surface sits on top of the same engine. Hierarchies, levels, measures and synonyms are discoverable through /ai/cubes and /ai/schema, and a single POST /ai/query turns a JSON description of a question into validated MDX, runs it, and returns typed cells with value, formatted and unit fields. When validation fails, the response carries a status, field and available envelope so the agent can correct itself instead of parsing a stack trace. A SQL-side twin lives under /ai/ossie/* for datasets modelled with the Ossie semantic layer, with the same list, schema and query shape, plus POST /ai/ossie/ask for a plain-language question. Anomaly and forecast endpoints share the envelope.

Running Saiku from the container image

The README leads with a one-line demo. The image is published to GitHub Container Registry, the container listens on 8080, and demo mode bundles a self-contained H2 database with the FoodMart cube.

bash
docker run -d -p 8080:8080 --name saiku -e SAIKU_DEMO=true ghcr.io/spiculedata/saiku

After that, the UI is at http://localhost:8080/ui/ and the credentials are admin / admin. Drag fields onto rows, columns or filters and the SPA writes the MDX for you. That is the fastest way to see whether the cube model matches how your analysts think.

For anything reachable on a network, the README is explicit: drop SAIKU_DEMO=true and set SAIKU_ADMIN_PASSWORD instead. The launcher refuses to start on the default admin/admin pair once it is network-reachable, so one of the two settings is required. The container also declares a volume at /app/saiku-home, which is where skills and agent spaces are read from.

bash
docker run -d -p 8080:8080 \
  -e SAIKU_ADMIN_PASSWORD='a-strong-password' \
  -v saiku-home:/app/saiku-home \
  ghcr.io/spiculedata/saiku:latest

Wiring an MCP host is a config edit rather than a new process. The README's example for Claude Desktop or Cursor execs the bundled wrapper inside the running container:

json
{
  "command": "docker",
  "args":    ["exec", "-i", "saiku", "saiku-mcp"]
}

The Dockerfile comments add a detail worth knowing: since issue #878 the MCP endpoint lives inside saiku-webapp at /rest/saiku/api/mcp, and MCP hosts authenticate with per-user Basic credentials over the same chain as the AI REST API. The stdio wrapper is a convenience, not a separate server.

Building from source requires GitHub Packages credentials before anything compiles

This is the part most likely to cost an afternoon. The build needs JDK 21 and Maven 3.9 or newer, and the README warns that you must set up GitHub Packages authentication first or the build fails before it compiles. Saiku's Mondrian fork, olap4j, saiku-query and Ossie artifacts are published to GitHub Packages, which requires an authenticated token even though the packages are public. Without it you get a bare 401 Unauthorized on pentaho:mondrian that never mentions tokens.

The token must be a classic personal access token with only the read:packages scope. The README states that GitHub's Maven registry does not accept fine-grained tokens, and that the token creation UI defaults to fine-grained, which is an easy way to lose an hour. Because the packages are public, no repo scope or organisation membership is needed. Five server entries then go into ~/.m2/settings.xml: github-mondrian-saiku, github-olap4j, github-olap4j-xmlaserver, github-saiku and one more the truncated README does not show.

The Dockerfile explains why this does not affect image users: the jar job builds the fat JAR, the docker job stages it as build-context/saiku.jar, and the Dockerfile only copies it in. No Maven runs at image build time, so the container never needs package credentials. If you only consume the image, the whole section is irrelevant. If you build from source, it is a prerequisite with a confusing failure mode.

Skills, spaces and the annotation namespace that make agents usable

Exposing a query endpoint to an LLM is not the hard part. Telling the model what the fields mean is. Saiku's mechanism for this is a set of saiku.semantic.* annotations on the cube, documented in docs/schema-annotations.md, which the /ai/cubes and /ai/schema endpoints surface as hierarchies, levels, measures and synonyms. An unannotated cube still answers queries, but the agent has only raw names to work from.

Two admin-level extensions avoid code changes. Skills are markdown files under saiku-home/skills/*.md with YAML frontmatter, discoverable from /ai/ask. A user can invoke one by prefixing an ask with /skill-name, or ask naturally and let the LLM route through the skill catalogue. Spaces are JSON files under saiku-home/agent-spaces/*.json that define a named persona scoped to a system prompt, a cube allowlist and a skill allowlist. POST /ai/spaces/{id}/ask enforces the persona server-side, so a request for a cube outside the allowlist returns a 403 the user cannot override.

That server-side enforcement is the design decision worth noting. It means the allowlist is a real boundary rather than a prompt instruction the model may ignore. The cost is that every new cube a persona needs has to be added to the JSON file, and there is no sign in the README of a UI for editing spaces or skills.

Where Saiku is the wrong tool

Saiku assumes a cube. If your analysts work from flat tables and you have no intention of modelling hierarchies, levels and measures, the semantic layer is overhead you will pay for and not use. A plain SQL engine plus a notebook will get you further.

Operationally, this is a Java application server, not a library. Jetty, Jersey, Spring and Spring Security arrive in one JAR, and the deployment unit is a container with a persistent volume at /app/saiku-home. Teams that wanted to embed an OLAP query layer inside an existing service will find the process boundary and the REST contract awkward.

The build constraint is a second boundary. Requiring a classic personal access token with read:packages for every machine that compiles the project is friction that a self-hosted Maven mirror would avoid, and the README's own warning about the misleading 401 suggests this has bitten people. The README also notes that observability has gaps: docs/observability.md describes Tier 2 custom spans for ThinQueryService and similar as not yet covered. If you need per-query internal tracing rather than Jetty, Jersey and JDBC spans, that is a known hole.

Finally, the release line is at v4.8.0-RC2. Release candidates in production are a choice, not a default, and the README does not document a rollback path for a schema or repository migration.

How Saiku differs from embedding Calcite directly

The obvious alternative is to skip the semantic layer and use Apache Calcite on its own. Calcite is a SQL planner and query federation framework; you write a schema adapter, register tables, and get SQL planning over your sources. Saiku uses Calcite for exactly that planning job, but wraps it in Mondrian's cube model, which is where hierarchies, levels, measures and MDX come from.

The practical difference is what you get for free. With Calcite alone you have SQL and no notion of a dimension. Excel cannot connect to it over XMLA. There is no /ai/schema endpoint that lists measures with synonyms, because there are no measures. You would build the MDX layer, the XMLA endpoint, the REST contract and the agent-facing description yourself.

In the other direction, Calcite alone is far lighter. It is a library you add to a service you already run, with no Jetty, no Spring Security and no container. The README points to examples/lakehouse-demo for the Saiku to Mondrian to Calcite to Trino to Iceberg path, which is the case where the semantic layer earns its weight: the cube definition stays stable while the engine underneath it changes.

Maintenance, licensing and what an upgrade costs

The repository is not archived, and its last push was on 2026-08-16, the same day v4.8.0-RC2 was tagged. The two preceding releases, v4.8.0-RC1 and v4.7.1, landed on 2026-08-16 and 2026-08-02. The default branch is development, so the release candidates are cut from the working branch rather than a stabilised line.

Saiku is licensed under Apache-2.0, with a NOTICE file at the repository root. That is a permissive licence, but the project also depends on a Spicule fork of Mondrian and on olap4j, saiku-query and Ossie artifacts distributed through GitHub Packages. Anyone redistributing a build should check the licences of those dependencies rather than assuming the top-level Apache-2.0 identifier covers everything. This is a note about what to verify, not legal advice.

Upgrade cost is dominated by the Mondrian fork. Because the fork is versioned separately and consumed from GitHub Packages, upgrading Saiku means moving the fork and olap4j together, and the mondrian.backend flag means you may be carrying two planner behaviours at once. The README does not document a rollback procedure, so a team that upgrades in place should keep the previous image tag available. The CHANGELOG.md at the repository root is the place the project records what changed between releases.

Editorial conclusion

Adopt Saiku if your measures already live in a Mondrian-style cube and you need that same model reachable from Excel, a dashboard and an MCP-capable agent without writing a second semantic definition. Do not adopt it if you want a library you can embed in an existing JVM service, or if you cannot hand out GitHub Packages credentials to every machine that builds the project. Before committing, verify two things: that your deployment sets either SAIKU_DEMO=true or SAIKU_ADMIN_PASSWORD, since the launcher refuses to start on the default credentials once it is network-reachable, and that your cube annotations use the saiku.semantic.* namespace, because the AI endpoints describe themselves to agents from those annotations and an unannotated cube gives the agent less to work with.

Frequently asked questions

How do I install Saiku and log in for the first time?

The README gives a single docker run command with -e SAIKU_DEMO=true and port 8080, after which the UI is at http://localhost:8080/ui/ and the credentials are admin / admin. Demo mode bundles a self-contained H2 database and the FoodMart cube. For a real deployment you drop SAIKU_DEMO and set SAIKU_ADMIN_PASSWORD instead.

Why does building Saiku from source fail with a 401 Unauthorized on pentaho:mondrian?

The Mondrian fork, olap4j, saiku-query and Ossie artifacts are published to GitHub Packages, which requires an authenticated token even though the packages are public. The README warns that the failure is a bare 401 that never mentions tokens. You need a classic personal access token with the read:packages scope and five server entries in ~/.m2/settings.xml.

Can Saiku serve AI agents without exposing MDX to them?

Yes. The README describes a typed REST surface under /rest/saiku/api/ai/* where a single POST /ai/query translates a JSON description of a question into validated MDX and returns typed cells. Validation failures carry a status, field and available envelope so the agent can correct itself. The same contract has a SQL-side twin under /ai/ossie/*.

How do I connect Saiku to Claude Desktop or Cursor?

The README's MCP example runs the bundled stdio wrapper inside the running container with docker exec -i saiku saiku-mcp. Per the Dockerfile comments, the MCP endpoint itself lives inside saiku-webapp at /rest/saiku/api/mcp since issue #878, and MCP hosts authenticate with per-user Basic credentials over the same chain as the AI REST API.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
  4. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/spiculedata-saiku.svg)](https://hysenlabs.com/projects/spiculedata-saiku)