TiDB: A Distributed SQL Database With MySQL Compatibility and an HTAP Split
TiDB is built for agentic workloads that grow unpredictably, with ACID guarantees and native support for transactions, analytics, and vector search. No data silos. No noisy neighbors. No infrastructure ceiling.
At a glance
- What is it?
- TiDB separates SQL computation from storage and adds a columnar replica for analytics. This review covers what the repository actually ships, how to start a local cluster, and where the design costs you.
- Who is it for?
- TiDB fits teams that have outgrown a single MySQL primary and need ACID transactions plus analytics without maintaining a separate warehouse. It does not fit a single-node application with a few gigabytes of data, where a plain MySQL instance is simpler and cheaper to operate.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Go, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
DEEP OPEN-SOURCE ANALYSIS
What TiDB replaces, and for whom
TiDB targets the point where a single MySQL primary stops being enough. The README describes it as an open-source, cloud-native, distributed SQL database designed for high availability, horizontal and vertical scalability, strong consistency, and high performance. The practical problem it addresses is the split that appears when transactional data grows: you keep the write path in MySQL, then bolt on a separate analytical store fed by change data capture, and spend the next two years reconciling the two.
TiDB's answer is to keep one SQL surface and move the storage underneath it. Applications keep speaking the MySQL protocol, so existing drivers and ORMs connect without a rewrite, and the analytical path is served from the same cluster rather than a downstream copy. That is the audience: teams with a MySQL-shaped application, growing data volume, and a reporting or vector search requirement that a read replica cannot absorb.
It is not aimed at small deployments. Every architectural choice in the repository assumes a cluster: Raft groups, a placement driver, separate storage nodes. A project with ten gigabytes of data and one writer gets nothing from that machinery.
The TiKV, TiFlash and PD split explained
The architecture separates computing from storage. TiDB servers are stateless SQL layer processes; they parse queries, build plans, and coordinate execution across the storage engines. Data lives in TiKV, which the README describes as a row-based storage engine, and Raft consensus keeps replicas consistent. A transaction commits only after writing to the majority of replicas, which is where the strong consistency guarantee comes from and also where write latency is decided.
TiFlash is the second engine, columnar, and it changes the shape of the system. According to the README, TiFlash uses the Multi-Raft Learner protocol to replicate data from TiKV in real time, so the columnar copy stays consistent with the row store. The TiDB server then decides which engine serves which part of a query. This is the HTAP claim in concrete form: one cluster, two storage formats, a planner that routes between them.
The third component is PD, the placement driver, which the repository layout implies through the scale-out tooling and the operator documentation. It handles scheduling and replica placement. That is the piece most teams underestimate, because it is a stateful service whose availability matters as much as the storage nodes.
Vector search sits on top of the same SQL layer, per the README's list of key features. It is presented as a TiDB capability rather than a separate service, which is consistent with the no-data-silos framing.
Building and running a local TiDB server
The README's Quick Start does not give shell commands. It points to three deployment routes: a local test cluster via the TiDB quick start guide, Kubernetes via TiDB Operator, or TiDB Cloud, which the README calls the recommended option and describes as having a free plan with no credit card required.
The repository's Makefile declares the server target as its default goal, and go.mod declares Go 1.25.12. Building from source produces the tidb-server binary that the Dockerfile later copies.
make serverThe Makefile's buildsucc target prints a confirmation line when the build finishes, so a successful run ends with that message rather than a silent exit.
The repository Dockerfile builds that same binary, copies it to /tidb-server, and exposes port 4000. Its header comment states plainly that the current dockerfile is only used for development purposes, and points production users to a separate artifact repository.
FROM golang:1.25.12@sha256:9006890ecba0a168034d99516084099ae3114d9f2b7d6572c77f2dde57ebc980 as builder
WORKDIR /tidb
COPY . .
ARG GOPROXY
ENV GOPROXY ${GOPROXY}
RUN make serverThat is the builder stage as it appears in the file. What you get from the finished image is a process listening on 4000 speaking the MySQL protocol. There is no TiKV, no PD and no TiFlash in that image, so nothing about replication or HTAP is exercised by it. Once a cluster is reachable, connect with any MySQL client on port 4000 and issue ordinary DDL and DML. The README points to the SQL statement overview for the grammar and to the developer guide for connecting from an application. If you want the HTAP behaviour, you need a deployment that includes TiFlash; a single server binary will not demonstrate it.
Where TiDB is the wrong tool
The strongest argument against TiDB in a given project is usually operational, not functional. A three-component cluster with Raft replication and a placement driver has more failure modes than a single MySQL instance. If your availability target can be met by a primary with a standby, TiDB adds moving parts without adding capability you will use.
The second constraint is write latency. Committing only after a majority of replicas acknowledge is what makes the consistency guarantee real, and it is also a floor on how fast a write can return. For workloads dominated by small, latency-sensitive writes, that floor is the whole story, and a single-node engine with local disk will beat it.
Compatibility is the third area to check rather than assume. The README says TiDB is compatible with MySQL 8.0 and that applications can move with no code changes "or with minimal modifications." That qualifier is doing work. The README links to a dedicated MySQL compatibility page rather than claiming full parity, which is the honest signal: verify your specific statements, functions and driver behaviour against that page before planning a migration. Features that depend on MySQL internals, storage engine selection, or behaviour outside the documented compatibility set are where migrations stall.
Finally, the repository's Dockerfile is not a deployment artifact, and the file says so. Teams that build a production rollout on it are using a development image for a job it was not written for.
TiDB against a plain MySQL primary with a read replica
The obvious alternative is what most teams already run: a MySQL primary, one or more read replicas, and a separate analytical store if reporting gets heavy. The difference is where the complexity lives. In that setup, replication is asynchronous by default, so replicas lag and reads can be stale; analytics runs on a different system with its own schema and its own consistency story.
TiDB inverts that. Consistency is the default because commits require a Raft majority, and there is no separate analytical system to keep in sync because TiFlash replicates from TiKV through the Multi-Raft Learner protocol. The cost moves to operations: you now run a distributed system with a placement driver instead of a single writer with followers.
A second alternative worth naming is TiDB Cloud, which the README recommends over self-hosting and describes as a fully managed service with a free plan. That is not really a different database, it is the same one with the operational burden transferred. For teams whose objection to TiDB is running PD and TiKV, the managed route removes exactly that objection, at the price of giving up control of the deployment.
The decision between them is not about features. It is about whether your team wants to operate a distributed storage layer or pay someone else to.
Licence, releases and what maintenance costs you
TiDB is Apache-2.0, and the README states that all source code is available on GitHub under that licence, including enterprise-grade features. The repository carries a LICENSES/ directory and a ThirdPartyNotices.txt, which is what you would expect from a Go project with a large dependency graph spanning cloud storage SDKs and Apache Arrow. Apache-2.0 permits commercial use and modification and includes a patent grant; it also requires that you preserve notices. If you redistribute a modified tidb-server, the notice files in the repository are the ones to carry forward. This is a description of the licence text, not legal advice, and any redistribution plan should go past whoever handles licensing at your organisation.
The release cadence visible in the tags is a patch line on the 8.5 series: v8.5.6 in April 2026, v8.5.7 in July 2026, and v8.5.8 on 2026-08-27. The last push to the default branch was on 2026-08-27, so the project is being worked on. Patch releases on a stable minor line mean upgrade cost is mostly in the patch direction, but a distributed database upgrade is not a single binary swap. TiKV, PD and TiFlash versions move together, and the scale-out documentation linked from the README is the place the ordering rules live.
Budget for the version matrix, not just the server. A TiDB server that is one patch ahead of its TiKV nodes is a support question you do not want to open.
Editorial conclusion
TiDB fits teams that have outgrown a single MySQL primary and need ACID transactions plus analytics without maintaining a separate warehouse. It does not fit a single-node application with a few gigabytes of data, where a plain MySQL instance is simpler and cheaper to operate. Before committing, verify three things: that your SQL actually runs under the documented MySQL compatibility rules, that you have a working TiUP or TiDB Operator deployment path rather than the development Dockerfile, and that the operational burden of TiKV, PD and TiFlash is something your team can carry. The repository's own Dockerfile states it is for development purposes only, and that sentence should decide your first deployment decision.
Frequently asked questions
What does TiDB stand for?
The README gives the pronunciation as /'taɪdiːbi:/ and states that "Ti" stands for Titanium.
Is TiDB free?
The source code is available on GitHub under the Apache 2.0 license, and the README states this includes enterprise-grade features. The README also describes TiDB Cloud as a fully managed service with a free plan that requires no credit card.
Who uses TiDB?
The repository does not name specific users. The README describes the intended fit as cloud-native deployments with high availability and horizontal scalability requirements, and points to community platforms including Discord, Slack and Stack Overflow for user discussion.
Is TiDB good?
The README positions it for high availability, strong consistency and HTAP workloads, and states compatibility with MySQL 8.0. Whether that fits depends on your write latency tolerance and your willingness to operate TiKV, PD and TiFlash.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/pingcap-tidb)
Community notes