Open-source project
heavyai/heavydb avatar
heavyai/heavydb

HeavyDB: a GPU-accelerated SQL engine for multi-billion row datasets

HeavyDB (formerly MapD/OmniSciDB)

3,061 stars461 forksC++Apache-2.0

At a glance

What is it?
HeavyDB is an Apache-2.0 columnar SQL engine that targets CPUs and Nvidia GPUs through JIT compilation and multi-tier caching. It is built from source with CMake, not installed from a package manager, and that shapes who can realistically adopt it.
Who is it for?
Adopt HeavyDB if you have Nvidia GPUs, a C++ toolchain, and a workload of large columnar scans where you can accept building the engine yourself. Do not adopt it if you need a one-line install, a managed service, or a small operational team without CUDA experience.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly C++, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What HeavyDB is for, and who ends up building it

HeavyDB, formerly MapD Core or OmniSciDB, is a SQL-based relational columnar database engine. The README states its purpose plainly: querying multi-billion row datasets in milliseconds without indexing, pre-aggregation, or downsampling. That sentence defines the workload it was designed around. If your queries are point lookups on a transactional table, the design works against you. If your queries are wide scans and aggregations over hundreds of millions of rows, the design is the point.

The intended audience is narrow. The repository is a C++ codebase with a CMake build, CUDA support on by default, and a Calcite-based SQL frontend. The README's table of contents leads with Building, Testing, and Dependencies. There is no docker compose up, no pip install, and no apt package in the README. It points readers to the product documentation at docs.nvidia.com/heavyai for usage and to heavyai.github.io/heavydb for internal architecture. A team evaluating HeavyDB is really evaluating whether it can build and operate a database engine, not whether it can install one.

Multi-tier caching and JIT compilation as the core mechanism

Two mechanisms carry the performance story. The first is multi-tiered caching of data between storage, CPU memory, and GPU memory, which the README lists as a feature. The second is a Just-In-Time query compilation framework. In practice this means query plans are compiled rather than only interpreted, which shifts cost from per-row execution to per-query setup. That trade-off favours repeated analytical queries over short interactive ones.

The top-level repository layout makes the architecture legible. Catalog/ holds metadata, DataMgr/ manages data, Fragmenter/ handles chunking, QueryEngine/ and QueryRunner/ execute, Parser/ and SQLFrontend/ parse, CudaMgr/ and L0Mgr/ talk to GPUs, and StringDictionary/ handles string encoding. Distributed/ exists, so multi-node deployment is contemplated. A ThriftHandler/ and heavy.thrift suggest a Thrift-based client interface, and Embedded/ suggests an in-process mode. Each of these is a directory name, not a documented API. The README does not describe how a client connects, what port is used, or how the distributed mode coordinates. For that, the product documentation is the only cited source.

Building HeavyDB from source with CMake

The README gives the build sequence directly. After installing the dependencies listed in the Dependencies section, create a build directory, run CMake, and compile. The example below uses debug mode, which is what the README shows first. A release build is what you would actually deploy, and the README documents -DCMAKE_BUILD_TYPE=release as the option for that.

bash
mkdir build
cd build
cmake -DCMAKE_BUILD_TYPE=debug ..
make -j 4

The build options matter more than usual here. -DENABLE_CUDA defaults to on, so a machine without CUDA support will need -DENABLE_CUDA=off or the build will fail. -DENABLE_ONLY_ONE_ARCH=off compiles GPU code for all architectures by default; the README notes that turning it on restricts compilation to the host machine's architecture and speeds up the build. -DENABLE_TESTS defaults to on. -DENABLE_AWS_S3 defaults to on, which means S3 support is compiled in unless disabled.

Once built, the README's testing path is short. The sanity_tests target runs the most common tests.

bash
make sanity_tests

That is the extent of what the README tells you about running the system. It does not document starting a server, loading data, or connecting a client. The repository does contain initdb.cpp, insert_sample_data, SampleData/, and heavyai.conf.sample, but the README does not explain how they fit together. Treat the product documentation at docs.nvidia.com/heavyai as the required next step, and treat the sample configuration file as the starting point for a real deployment rather than a finished one.

Sanitizer builds and the cost of a fresh build directory

The README is unusually explicit about sanitizer builds, and the constraints are worth reading carefully. Both AddressSanitizer and ThreadSanitizer require CUDA to be disabled and require a fresh build directory. That means you cannot run a sanitizer pass in the same build tree as your GPU build.

bash
mkdir build && cd build
cmake -DENABLE_ASAN=on -DENABLE_CUDA=off ..
make -j 4

Before running tests under ASan, the README sets an environment variable that disables two checks.

bash
export ASAN_OPTIONS=alloc_dealloc_mismatch=0:handle_segv=0
make sanity_tests

For TSan, the README points to a suppressions file at config/tsan.suppressions and requires it to be set through TSAN_OPTIONS. This is a real operational cost. A contributor working on query execution needs at least two build trees and a working knowledge of which warnings are third-party noise. The README does not say how long a full build takes, so treat build time as unknown until you measure it on your own hardware.

Where HeavyDB is the wrong tool

The clearest limitation is packaging. The README describes CPack for generating distribution packages and notes that packages built on CentOS with static linking enabled can be used on most other recent Linux distributions. That is a build recipe, not a supported binary distribution. -DPREFER_STATIC_LIBS only works on CentOS according to the README, which means the portability story depends on a platform that is itself aging.

The second limitation is GPU dependency. Nvidia GPUs are what the README says are currently supported. CPU-only operation is documented for X86, Power, and ARM, but ARM is marked experimental support. If your infrastructure is AMD GPUs or Apple silicon, the README offers nothing.

The third limitation is the documentation split. The README is a developer document. It covers building, testing, and packaging. It does not cover SQL syntax, deployment topology, backup, restore, or upgrade paths. The product documentation is hosted on a different domain and the release notes live at docs.heavy.ai/overview/release-notes. A team that needs a database with a single authoritative operational manual will find HeavyDB's information spread across at least four sites.

Finally, consider the release cadence. The most recent release in the repository is v9.0.0 from 2025-10-20, and the README announces HeavyAI 10.0 with an ETA of September 2026. The last push to the repository was on 2026-09-04. Anyone planning a production deployment should look at the 10.0 announcement and decide whether to build on the current branch or wait, because the README states a special performance branch previewed at VLDB 2026 will be merged into master after that release.

How HeavyDB differs from BlazingSQL and other GPU SQL engines

BlazingSQL appears in the related searches, and the comparison is instructive. BlazingSQL was a GPU-accelerated SQL engine built on top of cuDF and the RAPIDS ecosystem, oriented toward querying data that already lived in GPU DataFrames and Parquet files. HeavyDB takes the opposite architectural position: it is a database, with its own catalog, its own storage management, its own fragmenter, and its own SQL frontend built on Calcite. It owns the data rather than querying someone else's in-memory representation.

That difference decides adoption. If your data already lives in a RAPIDS pipeline and you want SQL over it, a DataFrame-native engine is a shorter path. If you want a system of record that happens to be GPU-accelerated, with a catalog and a distributed mode, HeavyDB is the closer fit. The cost is that you take on the full operational surface of a database: ingestion, metadata, storage layout, and upgrades. The repository layout confirms that surface exists, with Catalog/, DataMgr/, Fragmenter/, MigrationMgr/, and Distributed/ all present at the top level.

Editorial conclusion

Adopt HeavyDB if you have Nvidia GPUs, a C++ toolchain, and a workload of large columnar scans where you can accept building the engine yourself. Do not adopt it if you need a one-line install, a managed service, or a small operational team without CUDA experience. Before committing, verify three things: that your target platform is covered by the Dependencies section, that a release build with -DENABLE_CUDA=on completes on your hardware, and that the product documentation at docs.nvidia.com/heavyai describes the SQL features you depend on, because the repository README is a developer document and does not cover deployment or rollback.

Frequently asked questions

Which database is best for big data?

That depends on the workload, and the README does not rank databases. HeavyDB is designed for large columnar scans on hybrid CPU and GPU systems, which is one specific shape of big data work rather than a general answer.

Which is the largest database?

The README does not compare database sizes or name a largest one. It states only that HeavyDB targets querying multi-billion row datasets in milliseconds, without indexing, pre-aggregation, or downsampling.

What does "large database" mean?

The README describes the target as multi-billion row datasets queried in milliseconds. It does not state a maximum table size, a memory requirement, or a threshold at which a database counts as large.

What is a big data database?

The README does not define the term. It describes HeavyDB as a SQL-based relational columnar database engine that uses CPUs and Nvidia GPUs to query multi-billion row datasets without indexing, pre-aggregation, or downsampling.

Official sources

  1. heavyai/heavydb on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/heavyai-heavydb.svg)](https://hysenlabs.com/projects/heavyai-heavydb)