Open-source project
ytsaurus/ytsaurus avatar
ytsaurus/ytsaurus

YTsaurus: a multitenant big data platform with MapReduce, SQL and a key-value store in one cluster

YTsaurus is a scalable and fault-tolerant open-source big data platform.

2,216 stars221 forksC++Apache-2.0

At a glance

What is it?
YTsaurus bundles a distributed file system, a MapReduce engine, an SQL query engine and an OLTP key-value store behind one multitenant installation. It is aimed at operators who can run a Kubernetes cluster first and worry about query tuning later.
Who is it for?
Adopt YTsaurus if you need several workload types (batch MapReduce, SQL analytics, OLTP key-value) sharing one multitenant cluster and you have people who can run Kubernetes and read C++ build files. Do not adopt it if you want a single-node install or a managed service; the README points at Kubernetes or an online demo, not a local package.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 4 days ago.
What is it written in?
Mainly C++, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What problem YTsaurus solves, and who it is for

Most data platforms force a choice. A distributed file system plus a batch engine covers analytics but not transactional lookups. A key-value store covers lookups but not large scans. YTsaurus is built on the opposite premise: the README describes it as a distributed storage and processing platform with support for the MapReduce model, a distributed file system and a NoSQL key-value database, all in one system. The stated goal is a multitenant ecosystem where MapReduce, an SQL query engine, a job schedule and a key-value store for OLTP workloads are interrelated subsystems, so an organisation does not have to run several installations.

The audience follows from that. This is infrastructure for platform teams at companies with enough hardware to care about hardware utilisation, and enough users to care about tenant isolation. The README claims support for up to 1 million CPU cores, thousands of GPUs, exabytes of data across HDD, SSD, NVME and RAM, and tens of thousands of nodes. Those are ceiling figures from the project's own description, not measured results, and they say more about the design target than about what a small team will see. A single developer wanting a local analytics sandbox is not the intended reader; the README's own getting-started path sends that person to an online demo instead.

The subsystems behind one YTsaurus installation

The architecture visible from the repository is a set of cooperating services rather than a single binary. The top level carries yt/ for the core platform, yql/ for the query language layer, connectors/ for external integrations, and library/, util/ and contrib/ as shared code. CHYT and SPYT are separate components with their own release trains, which is why the recent releases list carries tags like chyt/2.19.0 and yt/chyt/controller/v0.0.18 alongside docker/query-tracker/0.4.1. The controller tag indicates that CHYT is deployed and supervised as a controller inside the cluster rather than installed per user.

The data path works in layers. Storage sits on a distributed file system that spans different media, and the README lists automated replication between servers as the mechanism behind durability, with no single point of failure as the design claim. Compute runs as jobs scheduled against that storage, with the job schedule named as one of the interrelated subsystems. On top of the same storage, CHYT provides a ClickHouse SQL dialect and SPYT provides Apache Spark, so the SQL surface is not one engine but two, each bringing its own compatibility story. Distributed ACID transactions are listed as a feature, which is what makes the key-value store usable for OLTP rather than only for batch writes. The README also mentions secure isolation for compute resources and storage, which is the multitenancy mechanism: tenants share the cluster but not each other's resources. The README does not describe the replication protocol, the scheduler's placement algorithm, or how ACID transactions are coordinated, so anyone evaluating failure behaviour has to go to the documentation site rather than the repository front page.

Installing YTsaurus: Kubernetes, the online demo, or a source build

The README gives exactly two ways to try the platform without building it: a Kubernetes cluster walkthrough linked from the documentation, and an online demo. There is no pip install, no single binary, and no apt package in the README. The repository does carry gradlew and build.gradle.kts at the top level, but the README's build section points at BUILD.md as the entry point for building from source, and that is the only build instruction it offers.

Start with the Kubernetes path. The README links the try-yt page in the documentation, which is where the actual manifests and steps live:

bash
# The README links this page for the Kubernetes walkthrough:
# https://ytsaurus.tech/docs/en/overview/try-yt#kubernetes

If you would rather not run anything, the online demo at ytsaurus.tech is the zero-setup option the README offers. It is the fastest way to see the UI and the query surface before committing hardware.

For a source build, the README defers entirely to BUILD.md:

bash
# The README states: Build from source code (BUILD.md)
# See BUILD.md in the repository root for the actual steps.

The repository root also contains platform-specific CMake entry points (CMakeLists.linux-x86_64.txt, CMakeLists.linux-aarch64.txt, CMakeLists.darwin-arm64.txt, CMakeLists.darwin-x86_64.txt), which tells you a source build is expected to be platform-aware rather than a one-line configure. The README does not state supported compiler versions, build times, or required third-party dependencies beyond what BUILD.md and conanfile.py imply. Treat a source build as a project in itself.

Once a cluster is up, the client surface is what you actually use day to day. The README lists a variety of SDKs and APIs without enumerating them, so the concrete Python, CLI and YQL entry points have to be confirmed in the documentation.

Where YTsaurus is the wrong tool

The honest limitation is operational weight. Every advantage in the README is framed at cluster scale: multitenancy that eliminates multiple installations, replication between servers, automated up and down-scaling, isolation between tenants. None of that helps a team of three with a few terabytes, and all of it is something you have to run. The getting-started path confirms the expectation, since it assumes Kubernetes or a hosted demo, not a laptop.

A second constraint is the split SQL surface. CHYT gives you a ClickHouse dialect and SPYT gives you Apache Spark. If your organisation has standardised on one of those, you are adopting a platform to get an engine you already know, which is a reasonable trade only if the shared storage and multitenancy are worth it. If your queries depend on features from a different dialect, the platform does not promise them. The README describes CHYT as powered by ClickHouse and SPYT as powered by Apache Spark, and that phrasing is accurate about the dependency: you inherit those engines' behaviour, including their differences from each other.

The README also does not document rollback or downgrade for cluster updates. It states updates happen with no loss of computing progress, which is a claim about running jobs, not about reverting a bad release. Anyone who needs a documented rollback procedure should treat that as an open question and check the documentation and release notes directly rather than assume it exists.

YTsaurus compared with Spark, ClickHouse and YDB

These comparisons are the ones people actually search for, and the differences are structural rather than cosmetic.

Against Apache Spark: Spark is a processing engine, and you bring storage (object storage, HDFS, a warehouse) and a scheduler. SPYT is YTsaurus's Spark integration, so the relationship is not either-or at the engine level. The difference is that YTsaurus supplies the storage, the scheduler and the multitenant isolation underneath, and the README describes SPYT as supporting launch and support for multiple mini SPYT clusters with easy migration for ready-made solutions. If you already run Spark against object storage and have no cluster-level tenancy problem, YTsaurus adds a layer you would have to operate.

Against ClickHouse: ClickHouse is the database, and CHYT is YTsaurus's ClickHouse-powered query layer over YTsaurus storage. The README sells this on a well-known SQL dialect and integration with BI tools through JDBC and ODBC. The trade is that you get ClickHouse's query behaviour without ClickHouse's own storage and replication model, because durability and replication come from YTsaurus instead. Teams that want a standalone ClickHouse cluster with its own operational model should stay with ClickHouse.

Against YDB: both come from the same broad ecosystem of distributed data systems, and both offer a key-value store with transactions. The README positions YTsaurus as a platform that also carries MapReduce, a distributed file system, a job schedule and SQL engines, so the choice is about whether you want the OLTP store as one subsystem of a larger platform or as the centre of the design. The README does not compare the two feature by feature, and anyone deciding between them should read both sets of documentation rather than rely on this page.

Maintenance, releases and the Apache-2.0 licence

The repository is not archived, and the last push was on 2026-08-14. Releases are split by component rather than cut as one platform version: the recent list shows YTsaurus Strawberry Controller 0.0.18 and YTsaurus CHYT 2.19.0 both dated 2026-08-14, and YTsaurus QueryTracker 0.4.1 dated 2026-08-12. That versioning scheme is the maintenance cost in miniature. Upgrading the platform, the CHYT controller and the query tracker are separate decisions on separate schedules, and a deployment has to track which combination it is running. The README does not publish a compatibility matrix between these components, so that mapping is something an operator has to establish from the release notes.

Licensing is Apache-2.0, per the LICENSE file at the repository root. That is a permissive licence, which in practice means the usual obligations around attribution and notice retention apply, and it does not impose copyleft on your own code. This is a description of the licence identifier, not legal advice; if you are embedding the platform in a product, have your own counsel read the LICENSE file.

One further cost signal sits in the repository layout: the presence of conanfile.py, a clang.toolchain file and platform-specific CMake entry points means a source build pulls in a C++ toolchain and dependency management. Building from source is a recurring maintenance task, not a one-off, if you choose that route over packaged Kubernetes deployments.

Editorial conclusion

Adopt YTsaurus if you need several workload types (batch MapReduce, SQL analytics, OLTP key-value) sharing one multitenant cluster and you have people who can run Kubernetes and read C++ build files. Do not adopt it if you want a single-node install or a managed service; the README points at Kubernetes or an online demo, not a local package. Before committing, verify three things: that the Kubernetes deployment path in the documentation covers your storage media (HDD, SSD, NVME, RAM are all named as supported), that CHYT or SPYT gives you the SQL dialect your BI tools already speak through JDBC and ODBC, and that you can build or obtain the binaries, since BUILD.md is the only build entry point the README lists.

Frequently asked questions

How does YTsaurus compare with Spark?

Spark is a processing engine that needs separate storage and scheduling, while YTsaurus supplies the distributed file system, the job schedule and multitenant isolation underneath, and offers SPYT as its Apache Spark integration. The README describes SPYT as supporting multiple mini SPYT clusters and easy migration for ready-made solutions.

How does YTsaurus compare with ClickHouse?

ClickHouse is a standalone database, while CHYT is YTsaurus's ClickHouse-powered query layer running over YTsaurus storage, giving a familiar SQL dialect and BI integration through JDBC and ODBC. In that setup durability and replication come from YTsaurus rather than from ClickHouse's own storage model.

How does YTsaurus compare with YDB?

Both provide a key-value store with transactions, but YTsaurus is presented as a broader platform that also includes MapReduce, a distributed file system, a job schedule and SQL engines. The README does not compare the two feature by feature, so the choice depends on whether you want the OLTP store as one subsystem of a larger platform.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
  4. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/ytsaurus-ytsaurus.svg)](https://hysenlabs.com/projects/ytsaurus-ytsaurus)