# Vitess: MySQL Sharding Without Rewriting Your Application

> Vitess is a CNCF database clustering system that puts a sharding layer in front of MySQL so queries stay unaware of where rows live. It is aimed at teams whose single MySQL instance is no longer enough, and it costs a control plane to run.

**vitessio/vitess** — Vitess is a database clustering system for horizontal scaling of MySQL.

- Repository: https://github.com/vitessio/vitess
- Website: http://vitess.io
- Stars: 21,362 · Forks: 2,410
- Language: Go
- License: Apache-2.0
- Published: 2026-09-21 · Updated: 2026-09-21 · Language: en
- Canonical page: https://hysenlabs.com/projects/vitessio-vitess

## The problem Vitess solves: MySQL that outgrows one primary

A single MySQL server has a ceiling. Once write traffic or data volume passes what one primary can absorb, the usual answers are manual sharding in the application, or moving to a different database entirely. Vitess takes a third path: it keeps MySQL underneath and inserts a routing layer on top.

The README describes the goal directly: Vitess "allows application code and database queries to remain agnostic to the distribution of data onto multiple database servers." That is the whole proposition. Your application keeps issuing SQL against what looks like one MySQL endpoint. Vitess decides which shard answers.

Who this is for: teams with existing MySQL schema, existing MySQL operational knowledge, and a sharding key they can commit to. Vitess was a core component of YouTube's database infrastructure from 2011 and grew to tens of thousands of MySQL nodes, and the README lists Slack, Square (now Block) and JD.com among later adopters. Those are large deployments. The project does not position itself as a small-team tool, and the architecture below explains why.

## How Vitess works: vtgate, vtctld, vttablet and the shard map

Vitess is not a storage engine and not a fork of MySQL. It is a set of Go services that sit between clients and unmodified MySQL instances.

Clients connect to vtgate, the proxy that speaks the MySQL protocol. vtgate parses each query, consults the cluster's sharding metadata, and routes it to the relevant shard or shards, merging results when a query spans more than one. Because vtgate speaks MySQL, existing drivers and ORMs can connect without change, which is what makes the "agnostic" claim in the README practical rather than aspirational.

Behind vtgate sit vttablets, one per MySQL instance. A vttablet manages its MySQL server, handles health checks and serves queries on behalf of the shard it owns. Each shard has a primary and replicas, and Vitess handles failover between them.

Above that sits vtctld, the control plane. It holds the topology, which records which shards exist, which tablets serve them, and where the data lives. The repository also ships vtadmin, a web UI built separately from the Go binaries: the Dockerfile has a dedicated node:26-trixie-slim stage that builds web/vtadmin from web/vtadmin/package.json.

The README makes one concrete performance claim about resharding: you can "split and merge shards as your needs grow, with an atomic cutover step that takes only a few seconds." Note what that does and does not say. The cutover is short; the copy of data that precedes it is not, and the README does not describe how long a reshard takes end to end.

## Installing Vitess and running a first local cluster

The README does not contain install instructions. It points to vitess.io for documentation and to the examples/ directory in the repository for runnable setups. The directory listing shows examples/local/, examples/compose/, examples/demo/, examples/operator/, examples/region_sharding/ and examples/vtexplain/, which is where a first run should start.

If you build from source, the Makefile is the entry point. It exports GOOS and GOARCH from the local Go toolchain and installs binaries under a PREFIX you set:

```bash
make install PREFIX=/vt/install
```

The Dockerfile uses the same target but skips the admin UI build, which is handled by a separate Node stage:

```bash
make install PREFIX=/vt/install NOVTADMINBUILD=1
```

The Dockerfile pins its toolchain: golang:1.27.1-trixie for the Go binaries and node:26-trixie-slim for vtadmin. go.mod declares go 1.27.1. If your environment cannot supply those versions, the container build is the path of least resistance.

For a first real use, the repository's own examples are the honest starting point rather than a hand-written config. examples/local/ is the single-machine setup, and examples/compose/ wires the components together with Docker Compose. The README does not document the flags or ports those examples expose, so read the example directory itself before running anything against data you care about. For query planning specifically, examples/vtexplain/ exists to show how a statement would be executed across shards without executing it, which is the cheapest way to find out whether your schema and sharding key produce sane plans.

## Where Vitess is the wrong tool

Vitess adds a distributed system in front of your database. That is a permanent operational cost, not a one-time setup fee.

If your workload fits on one MySQL primary with read replicas, Vitess is overhead. You gain nothing from shard routing when there is one shard, and you now run vtgate, vtctld and vttablet processes that can fail independently of MySQL.

Sharding key choice is the harder constraint. The README frames Vitess as making queries agnostic to data distribution, but that only holds for queries that carry the sharding key. Cross-shard queries have to be scattered and merged, and the README does not document how vtgate bounds that cost. A schema whose primary access pattern cannot be expressed in terms of a single sharding column is a poor fit, and changing the key later means a reshard.

Vitess is also MySQL-specific. The README describes it as "built around MySQL," and the related searches show people asking about Vitess for Postgres. Nothing in the repository material supports that: the go.mod depends on github.com/go-sql-driver/mysql, and the architecture assumes MySQL semantics. If Postgres is your database, this is not the project.

Finally, the README is silent on several things an operator needs: backup and restore procedures beyond the examples/backups/ directory, upgrade ordering between vttablet and vtgate, and rollback of a failed reshard. Treat those as questions for the documentation site and the Slack workspace, not as gaps the README will close.

## Vitess compared with a proxy and with a NewSQL database

Two alternatives come up constantly in the search data, and they differ from Vitess in kind, not degree.

A connection proxy such as ProxySQL sits in front of MySQL and routes traffic. It can send reads to replicas and writes to a primary, and it can do so with far less machinery than Vitess. What it does not do is own the shard map. There is no resharding operation, no atomic cutover, and no component that knows how a logical table maps onto multiple physical MySQL servers. If your problem is read distribution, a proxy is the smaller answer. If your problem is that the data no longer fits on one primary, a proxy does not address it.

A distributed SQL database such as TiDB or CockroachDB takes the other route: it replaces MySQL rather than wrapping it. Storage and consensus are part of the product, and the database itself decides how to split and move ranges. The trade-off runs the opposite way from Vitess. You get automatic range splitting without a separate control plane to operate, and you give up running stock MySQL under the hood with the operational knowledge and tooling that comes with it. Vitess keeps your MySQL instances; TiDB and CockroachDB do not have any. For a team with deep MySQL experience and a schema that already works, that difference is the deciding factor.

## Maintenance, releases and the Apache-2.0 licence

The repository is not archived, and the last push was on 2026-09-19. Recent releases are v24.0.3 and v23.0.6, both published on 2026-09-03, following v24.0.2 on 2026-06-24. Two maintained release lines means security fixes can land on the older line, but it also means you have to decide which line you are on and track it. The README points to a roadmap on vitess.io and a low-frequency blog, and the changelog/ directory in the repository holds release notes.

The source is distributed under the Apache License 2.0, with a FOSSA licence scan badge in the README and a .fossa.yml at the repository root. Apache-2.0 is permissive: it includes an express patent grant and permits commercial use and modification. It does not, however, settle what you owe anyone. Vitess is a CNCF project with a GOVERNANCE.md and a STEERING.md, and the README directs security reports to a CNCF mailing list rather than a private GitHub channel. A third-party security audit was performed by ADA Logics, and the report is linked from the README as doc/VIT-03-report-security-audit.pdf. If your organisation requires an audit trail before adopting infrastructure, that document is the first thing to read, not the last.

Upgrade cost is the part the README does not cover. With vtgate, vtctld and vttablet all in the path, a version bump is a rolling operation across components, and the README gives no ordering guidance. Budget for reading the changelog per release rather than assuming a drop-in replacement.

## Conclusion

Vitess fits teams already running MySQL at a scale where one primary can no longer take the writes, and who can staff a control plane of vtgate, vtctld and per-shard vttablets. It is the wrong choice if you only need read replicas, or if you have not yet decided on a sharding key, because resharding is possible but the key shapes everything downstream. Before committing, read the vtexplain and region_sharding examples in the repository and confirm which release line, v23.0.6 or v24.0.3, your tooling targets.

## FAQ

### What is Vitess used for?

Vitess is used to horizontally scale MySQL by sharding it. It routes queries through a proxy so application code stays unaware of how data is distributed across database servers, and it supports splitting and merging shards as load grows.

### What are the key differences between MySQL and Vitess?

MySQL is the database that stores the data; Vitess is a clustering layer built around MySQL that distributes it across multiple servers. Vitess adds components such as vtgate for query routing and vtctld for topology, while the underlying MySQL instances remain MySQL.

### Is Vitess open source?

Yes. The source files are distributed under the Apache Version 2.0 license, and the project is listed under the CNCF with public governance documents in the repository.

### How do I use Vitess?

The README points to vitess.io for documentation, and the repository ships runnable setups under examples/, including examples/local/ for a single-machine cluster and examples/compose/ for a Docker Compose deployment. Clients connect to vtgate using the MySQL protocol.

## Sources

- [License: Apache-2.0](https://github.com/vitessio/vitess/blob/main/LICENSE)
- [Project website](http://vitess.io)
- [README](https://github.com/vitessio/vitess/blob/main/README.md)
- [Releases](https://github.com/vitessio/vitess/releases)
- [vitessio/vitess on GitHub](https://github.com/vitessio/vitess)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/vitessio-vitess
