# MatrixOne: an HTAP database with Git-style version control for data

> MatrixOne is a Go-written HTAP database that adds zero-copy snapshots, branching and time travel to a MySQL-compatible engine, with vector and full-text search in the same system. It is aimed at teams tired of stitching MySQL, ClickHouse, Elasticsearch and a vector store together, but the operational surface is larger than a single-purpose database.

**matrixorigin/matrixone** — AI-native HTAP database with Git-for-Data and built-in vector search, serving as the data and memory backbone for intelligent agents and applications.

- Repository: https://github.com/matrixorigin/matrixone
- Website: https://docs.matrixorigin.cn/en
- Stars: 2,010 · Forks: 327
- Language: Go
- License: Apache-2.0
- Published: 2026-08-08 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/matrixorigin-matrixone

## The problem MatrixOne is aimed at, and who feels it

The README opens with a familiar stack diagram: MySQL for transactions, ClickHouse for analytics, Elasticsearch for search, Pinecone for AI. Four systems, several ETL jobs, and a lag between the moment a row is written and the moment it is queryable in the analytical copy. MatrixOne's pitch is that one engine can serve all four workloads, so nothing has to be copied between them.

The second problem is narrower and more interesting. Data mistakes are expensive, and unlike code they have no version history. The README states that MatrixOne is "the industry's first database to bring Git-style version control to data", and points to an arXiv paper titled Version Control System for Data with MatrixOne. The promised operations are snapshots, time travel, branch and merge, rollback and an audit trail.

That combination defines the audience. It is not a team that wants a smaller database. It is a team that already runs several, has been burned by a bad migration or a bad backfill, and is willing to take on a distributed system in exchange for being able to branch a production dataset, run the migration there, and merge or discard it.

## How the pieces fit: one engine, four workloads

The README describes a hyper-converged HSTAP engine that handles transactional, analytical, full-text and vector workloads in a single system. The repository layout is consistent with that claim: the Go module is a single module (go.mod declares module github.com/matrixorigin/matrixone and requires Go 1.26.4), with pkg/, cmd/, proto/ and clients/ at the top level. There is no separate vector service or search service to deploy.

The storage and compute split shows up in the dependency list. The go.mod file pulls in the AWS SDK v2 S3 client, the Aliyun OSS SDK and the HDFS client, which is what you would expect from a system that can put its data on object storage or a distributed filesystem rather than only on local disk. The Makefile builds a single binary named mo-service, and the default target builds it. GPU compilation is opt-in: the Makefile comments describe compiling with MO_CL_CUDA=1 after installing the CUDA toolkit and the cuVS Go bindings, which implies the vector index code has an accelerated path but does not require one.

The vector and full-text capabilities are exposed through SQL and through a Python SDK rather than as separate APIs. The README's example uses create_vector_column for an embedding column and a fulltext_index.create call for a full-text index, then queries with boolean_match. Everything runs through the same connection.

## Installing MatrixOne and running a first vector query

The README's 60-second path is a single Docker command. It starts the server and publishes port 6001.

```bash
docker run -d -p 6001:6001 --name matrixone matrixorigin/matrixone:latest
```

Once the container is up, the README creates a database over the MySQL client. Note the flags as written: user root, password 111, host 127.0.0.1, port 6001.

```bash
mysql -h127.0.0.1 -P6001 -p111 -uroot -e "create database demo"
```

The Python SDK installs from PyPI under the name matrixone-python-sdk.

```bash
pip install matrixone-python-sdk
```

The README's vector example defines a table with an 8-dimensional float32 embedding column and connects to the demo database. The snippet in the README is truncated partway through table creation, so treat the model definition below as the shape of the API rather than a complete runnable script.

```python
from matrixone import Client
from matrixone.orm import declarative_base
from sqlalchemy import Column, Integer, String, Text
from matrixone.sqlalchemy_ext import create_vector_column

client = Client()
client.connect(database='demo')

Base = declarative_base()

class Article(Base):
    __tablename__ = 'articles'
    id = Column(Integer, primary_key=True, autoincrement=True)
    title = Column(String(200), nullable=False)
    content = Column(Text, nullable=False)
    embedding = create_vector_column(8, "f32")
```

The full-text example creates an index over two columns and runs a boolean query with must and must_not clauses. Results come back as a ResultSet whose rows you iterate.

```python
from matrixone.sqlalchemy_ext import boolean_match

client.fulltext_index.create(
    Article, name='ftidx_content', columns=['title', 'content']
)

results = client.query(
    Article.title,
    Article.content,
    boolean_match('title', 'content')
        .must('machine')
        .must('learning')
        .must_not('basics')
).execute()

for row in results.rows:
    print(f"Title: {row[0]}, Content: {row[1][:50]}...")
```

If you prefer to build from source, the Makefile's default target produces the mo-service binary and requires Go 1.26 or later; the file comments note that arch-specific SIMD kernels are built by default on x86_64.

## Where MatrixOne is the wrong choice

The README leans on the phrase "no compromises", and that is the claim to be sceptical about. A database that serves OLTP, OLAP, full-text and vector search from one engine is making a resource-sharing decision on your behalf. Analytical scans and transactional writes now compete inside one process, and the tuning knobs that would let you isolate them do not exist in a single-purpose deployment because the workloads were never in the same place.

The version control story has the same shape. Snapshots are described as zero-copy and measured in milliseconds, which implies they depend on the storage layer's copy-on-write behaviour rather than on duplicating data. That is a real mechanism, but it also means snapshot cost and branch cost are tied to how the engine manages its own files, not to anything you control at the SQL level. The README does not document how long a snapshot is retained, what happens to branches when the underlying files are compacted, or what the merge conflict semantics are when two branches modify the same row. Those are the questions a Git-for-Data user will ask first, and the README is silent on them.

Finally, the deployment model. The README describes storage-compute separation and Kubernetes-native operation. That is a distributed system with the operational surface of a distributed system. If your workload fits comfortably in a single Postgres instance, adding MatrixOne buys you capabilities you will not use and a class of failure modes you did not have.

## How it differs from the tools it replaces

The honest comparison is not MatrixOne versus one database. It is MatrixOne versus the four-system stack the README draws, and the difference is where the data lives.

In the conventional stack, MySQL holds the current rows, an ETL job copies changes into ClickHouse, another pipeline feeds Elasticsearch, and an embedding job writes vectors into a dedicated vector store. Each hop introduces lag and a place for the pipelines to break. Queries that need a transaction and a vector similarity score have to be assembled in application code across two or three clients.

MatrixOne removes the hops. The vector column and the full-text index are on the same table as the transactional columns, so a single query can filter on a normal predicate and rank by embedding distance. That is the actual architectural difference, and it is the reason the Git-for-Data feature is more than a gimmick: a snapshot covers the vectors and the search index too, because they are not separate systems with separate backup schedules.

The cost is that you now depend on one engine for everything. With the four-system stack, a bad ClickHouse upgrade degrades analytics and leaves transactions untouched. With MatrixOne, the blast radius is the whole database. That trade is the entire decision.

## Maintenance, releases and licence

The repository is not archived, and the last push was on 2026-08-26. The release history over the preceding weeks shows v4.1.4 on 2026-07-22, v4.2.0 on 2026-08-21 and v4.2.1 on 2026-08-26, so minor releases arrive on the order of weeks and patch releases can follow within days. That cadence is good for fixes and bad for anyone who wants to pin a version and forget it: you will be reading release notes regularly.

Building from source has its own cost. The Makefile requires Go 1.26 or later, and the file explicitly disables go.work inheritance (override GOWORK := off) so that official targets cannot pick up a parent workspace that replaces dependencies. That is a sensible decision for reproducible builds and an annoying one if you were planning to vendor MatrixOne into a larger Go workspace.

MatrixOne is licensed under Apache-2.0, and the LICENSE file sits at the repository root alongside a LICENSES/ directory, which usually indicates bundled third-party components with their own terms. If you redistribute MatrixOne or ship it inside a product, read that directory rather than assuming the root licence covers everything. This is not legal advice; the LICENSES/ contents are the thing to check with whoever handles licensing on your side.

## Conclusion

Adopt MatrixOne if you are consolidating a transactional database, an analytical store and a vector index, and you want snapshot and branch semantics over the same tables. Do not adopt it if you need a single-purpose store with a narrow operational surface, or if your team has no capacity to run a distributed system. Before committing, verify three things yourself: that the MySQL wire protocol covers the statements your application actually issues, that snapshot and branch behaviour matches what the README claims on your data volume, and that the release cadence (v4.2.1 on 2026-08-26) fits your upgrade window.

## FAQ

### What is MatrixOne?

It is an HTAP database written in Go that combines transactional, analytical, full-text and vector workloads in one engine, and adds Git-style version control over data through snapshots, time travel, branching and rollback. The README describes it as MySQL-compatible and cloud-native.

### what is matrix one

The same project: a database from matrixorigin that the README presents as the first to bring Git-style version control to data, alongside built-in vector search and MySQL compatibility. The repository is licensed under Apache-2.0.

### How do I install MatrixOne and connect to it?

The README starts it with docker run -d -p 6001:6001 --name matrixone matrixorigin/matrixone:latest, then creates a database over the MySQL client on 127.0.0.1 port 6001 as user root. The Python SDK installs with pip install matrixone-python-sdk.

### Does MatrixOne support vector search and full-text search?

Yes. The README shows create_vector_column for embedding columns and IVF/HNSW as the vector index types, plus a full-text index created through the SDK and queried with boolean_match using must and must_not clauses.

### Can MatrixOne replace MySQL?

The README describes it as a drop-in replacement for MySQL that uses existing tools, ORMs and applications without code changes, and the quick start connects with the standard mysql client. The README does not enumerate which MySQL statements are unsupported, so verifying that against your own query set is the first thing to check.

## Sources

- [Official documentation](https://docs.matrixorigin.cn/en)
- [Official README](https://github.com/matrixorigin/matrixone#readme)
- [Project repository](https://github.com/matrixorigin/matrixone)
- [Release notes](https://github.com/matrixorigin/matrixone/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/matrixorigin-matrixone
