Open-source project
alibaba/AliSQL avatar
alibaba/AliSQL

AliSQL 8.0.44: DuckDB Analytics and HNSW Vector Indexes Inside a MySQL Branch

AliSQL is a MySQL branch originated from Alibaba Group. Fetch document from Release Notes at bottom.

5,985 stars903 forksC++NOASSERTION

At a glance

What is it?
AliSQL is Alibaba's MySQL 8.0.44 fork, adding a DuckDB analytical storage engine, native vector columns with HNSW indexes, and flashback queries. It keeps the MySQL client protocol, so existing drivers still connect.
Who is it for?
Adopt AliSQL if you already run MySQL 8.0 and want columnar analytics or vector search reachable through the ordinary MySQL protocol, without standing up a second database. Do not adopt it if you need a packaged binary, a Windows build, or a stable vector API today: the README states vector features are disabled by default and vector indexes require READ-COMMITTED, and the repository ships build scripts rather than installers.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 74 days ago.
What is it written in?
Mainly C++, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The split MySQL plus analytics plus vector stack that AliSQL collapses

A team running MySQL for transactions usually ends up with a second system for anything analytical and a third for embedding search. AliSQL's pitch is that both extra workloads live in the same server process. The README describes it as "an open-source MySQL branch maintained by Alibaba Cloud Database Team", based on MySQL 8.0.44, adding a DuckDB analytical storage engine, HNSW vector indexes, Native Flashback, and binlog optimizations, while retaining the MySQL client protocol. The audience is the operator who already knows MySQL: someone who wants a columnar table for aggregations and a vector column for similarity search without introducing a separate query language, a separate wire protocol, or a separate backup story. The README's own comparison table puts the two headline capabilities side by side: TPC-H SF100 results in the wiki showing more than 200x speedup over InnoDB for several queries in that test environment, and a VECTOR(N) type with HNSW indexes supporting vectors with up to 16,383 dimensions. Those numbers come from the project's own wiki, not from an independent run, so treat them as a claim about a specific test environment rather than a general expectation.

How the DuckDB engine and VIDX vector indexes actually sit in the server

Two mechanisms are visible in the README. The first is a storage engine you select per table. A table declared with ENGINE=DuckDB is stored and executed columnar, with automatic compression, and analytical aggregations over it run through DuckDB rather than InnoDB. This is a per-table decision, not a global mode switch, which means a single schema can hold InnoDB tables for OLTP and DuckDB tables for reporting. The second mechanism is a native vector type. A column declared VECTOR(N) is stored by the server itself, and an HNSW index built over it answers approximate nearest neighbour queries. The README states the index supports COSINE and EUCLIDEAN distance, and the example uses DISTANCE=COSINE at index creation time. Distance is then computed in the query with VEC_DISTANCE_COSINE, and values are converted from text with VEC_FROMTEXT. The design point worth noting is that vector search is not a plugin bolted onto a client library; it is a column type and an index type inside the same transaction and replication machinery as the rest of the schema. The README's footnote shows the cost of that integration: because DuckDB registers an additional 2PC participant, builds that include DuckDB use normal binlog group commit, and large-transaction optimization for DuckDB is deferred to the next release. Integration through the shared commit path is exactly why that optimization does not apply yet.

Building AliSQL from source and running a first vector query

The README gives two quick-start paths. Option 1 builds from source; Option 2 points to a step-by-step guide for setting up a DuckDB analytical node in the wiki. There is no package repository, installer, or container image documented in the README, so source build is the concrete path. The build script takes a build type, an install directory, and optional sanitizer or coverage flags, with release as the build type and the install prefix defaulting to /usr/local/alisql when writable and otherwise $HOME/alisql. Prerequisites listed are CMake 3.x+, Python 3, and GCC 7+ or Clang 5+.

bash
git clone https://github.com/alibaba/AliSQL.git
cd AliSQL
sh build.sh -t release -d ~/alisql
make install

After building, the README initializes the data directory insecurely for a local test and starts the server with DuckDB enabled. The --duckdb_mode=ON flag in the README's example is what makes the analytical engine usable.

bash
~/alisql/bin/mysqld --initialize-insecure --datadir=~/alisql/data
~/alisql/bin/mysqld --datadir=~/alisql/data --duckdb_mode=ON

Vector support is off until you turn it on, and the README's example sets the session isolation level to READ-COMMITTED before creating anything. The first real use is a table with a VECTOR column, an HNSW index over it, and an ordered distance query.

sql
SET GLOBAL vidx_disabled = OFF;
SET SESSION transaction_isolation = 'READ-COMMITTED';

CREATE TABLE embeddings (
    id INT PRIMARY KEY,
    content TEXT,
    embedding VECTOR(3)
) ENGINE=InnoDB;

CREATE VECTOR INDEX idx_embedding ON embeddings(embedding) DISTANCE=COSINE;

Rows are inserted with VEC_FROMTEXT, and similarity is retrieved by ordering on VEC_DISTANCE_COSINE against a query vector, limited to the number of neighbours you want. If the index creation fails, the first thing to check is whether the session isolation level is READ-COMMITTED, because the README ties vector index creation to that setting.

Vector search is off by default and gated on READ-COMMITTED

The most important limitation is stated plainly in the README's own example: vector features are disabled by default, and vector indexes require RC. That is two separate constraints. The first means a fresh server will not expose the vector path until vidx_disabled is turned off globally, so an application cannot assume the capability exists just because the binary was built from this branch. The second means index creation is tied to the READ-COMMITTED isolation level, which is not the default in MySQL and is a setting some transactional workloads deliberately avoid. A team whose application depends on REPEATABLE READ semantics has a real conflict to resolve before it can use HNSW indexes, and the README does not describe a workaround. The second limitation is the DuckDB commit path. The README's footnote states that Free Flush supports large InnoDB transactions, but that in this release DuckDB registers an additional 2PC participant, so AliSQL builds that include DuckDB use normal binlog group commit, and large-transaction optimization for DuckDB will be added in the next release. If you enable DuckDB on a build, you give up the large-transaction binlog optimization described in the features table for the workload that touches DuckDB. Third, the feature table marks Instant DDL, parallel B+tree construction, non-blocking locks, and accelerated crash recovery as Planned, not Available. None of those are in the release. Finally, the README documents no rollback procedure for a schema that has adopted ENGINE=DuckDB tables, and it does not describe how to migrate a DuckDB table back to InnoDB.

Where AliSQL is the wrong choice, and what a plain MySQL plus DuckDB setup does differently

The obvious alternative for the analytical half is running DuckDB as a separate process and moving data into it, which is how most teams use DuckDB today: an embedded columnar engine reading Parquet or its own database files, driven from an application or a notebook, with no shared transaction log with MySQL. The difference in approach is the boundary. Standalone DuckDB gives you a single-writer analytical engine decoupled from your OLTP server, at the cost of moving data across that boundary yourself and having no unified commit. AliSQL puts DuckDB behind the MySQL storage engine interface so a table is created with ENGINE=DuckDB and queried with ordinary SQL through the MySQL protocol, which buys you one connection, one schema, and one replication stream. It costs you the coupling described above: the added 2PC participant and the resulting fallback to normal binlog group commit. For vector search, the alternative is a dedicated vector database reached over its own API. That approach typically gives a more mature and better documented ANN surface than a feature the README says is disabled by default and requires a specific isolation level, but it adds another service to operate and another consistency boundary between your relational rows and your embeddings. AliSQL is the wrong tool when your analytical volume is large enough that you want an independent engine you can scale and tune on its own, when your application cannot move to READ-COMMITTED, or when you need a supported binary distribution rather than a source build. It is also the wrong tool if you are not already on MySQL 8.0.44 semantics; the branch inherits that version's behaviour, and the README does not claim compatibility with other MySQL versions.

Maintenance cadence, release numbering, and the licence discrepancy

The release history is uneven in a way worth reading carefully. AliSQL-8.0.44-1 is dated 2026-01-20 and AliSQL-8.0.44-2 is dated 2026-07-17, so the current line moves on a roughly twice-a-year cadence for this version. The previous entry in the list, AliSQL-5.6.32-9, is dated 2018-05-01, which shows the 5.6 line was left behind years ago rather than maintained in parallel. The most recent push to the repository is dated 2026-07-18, one day after the 8.0.44-2 release, so commit activity is aligned with releases rather than continuous. For upgrade cost, the practical question is how much of upstream MySQL 8.0.44 the branch tracks; the README does not document a merge policy or a patch backlog, so that has to be checked in the repository history rather than assumed. On licensing, the two signals disagree and you should resolve it before shipping. The README badge says GPL 2.0 and links to the LICENSE file, while the repository metadata reports NOASSERTION, meaning no licence was detected automatically. GPL 2.0 is the same family as MySQL's own licensing, which matters if you distribute a modified server, but this is a description of what the files say, not legal advice. Read LICENSE directly and get your own counsel if you plan to redistribute.

Editorial conclusion

Adopt AliSQL if you already run MySQL 8.0 and want columnar analytics or vector search reachable through the ordinary MySQL protocol, without standing up a second database. Do not adopt it if you need a packaged binary, a Windows build, or a stable vector API today: the README states vector features are disabled by default and vector indexes require READ-COMMITTED, and the repository ships build scripts rather than installers. Before committing, verify which upstream MySQL 8.0.44 patches the branch has absorbed, confirm that the DuckDB engine path in your workload is not blocked by the README's note that DuckDB registers an extra 2PC participant and therefore falls back to normal binlog group commit, and check the LICENSE file directly, since the repository metadata reports NOASSERTION while the README badge says GPL 2.0.

Frequently asked questions

What is AliSQL?

It is an open-source MySQL branch maintained by Alibaba Cloud Database Team, based on MySQL 8.0.44, that adds a DuckDB analytical storage engine, HNSW vector indexes, Native Flashback, and binlog optimizations while retaining the MySQL client protocol.

How do I install AliSQL?

The README documents building from source: clone the repository, run sh build.sh -t release -d ~/alisql, then make install. Prerequisites listed are CMake 3.x+, Python 3, and GCC 7+ or Clang 5+. No packaged binary or container image is documented.

Does AliSQL support vector search?

Yes, through a VECTOR(N) column type and HNSW indexes supporting COSINE and EUCLIDEAN distance, with up to 16,383 dimensions. The README states vector features are disabled by default and must be enabled with SET GLOBAL vidx_disabled = OFF, and that vector indexes require READ-COMMITTED.

Can I use DuckDB tables and InnoDB tables in the same AliSQL schema?

Yes. The DuckDB engine is selected per table with ENGINE=DuckDB, so a schema can hold InnoDB tables for transactional work and DuckDB tables for analytical queries. Note that builds including DuckDB use normal binlog group commit because DuckDB registers an additional 2PC participant.

Official sources

  1. alibaba/AliSQL on GitHub
  2. Issues
  3. README
  4. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/alibaba-alisql.svg)](https://hysenlabs.com/projects/alibaba-alisql)