Open-source project
prestodb/presto avatar
prestodb/presto

PrestoDB: the distributed SQL engine you build from source

The official home of the Presto distributed SQL query engine for big data

16,748 stars5,542 forksJavaApache-2.0

At a glance

What is it?
PrestoDB is a Java-based distributed SQL query engine for big data, distributed as source rather than a binary. This article covers what it solves, how to build and start it, and where it stops being the right tool.
Who is it for?
Adopt PrestoDB if you already run Hive, HDFS or another catalog-backed warehouse and need one SQL layer across several systems, and if your team is comfortable running Java services and building from source. Do not adopt it if you want a single-binary install, if your data fits in one Postgres instance, or if you have no Hive metastore to point the default configuration at.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Java, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

What PrestoDB solves, and who ends up running it

PrestoDB is a distributed SQL query engine for big data. The problem it addresses is the split between where data lives and where people want to query it: Hive tables, HDFS files, and the long list of connectors visible in the repository layout (presto-hive, presto-delta, presto-cassandra, presto-elasticsearch, presto-clickhouse, presto-bigquery, presto-druid, presto-accumulo, presto-atop) all become queryable through one SQL surface. The engine itself is Java, licensed Apache-2.0, and the README points deployment questions at the installation documentation rather than shipping a packaged server.

The audience is therefore narrower than the phrase "big data" suggests. You need a team that can build and operate JVM services, and you need at least one backend worth connecting. The README's requirements list is blunt about this: Mac OS X or Linux, Java 17 64-bit (Oracle JDK or OpenJDK), Maven 3.6.3+ for building, and Python 2.4+ for the launcher script. There is no Windows entry. If your team's idea of adopting a query engine is downloading an installer, this is the wrong shape of project.

How PrestoDB is put together: coordinator, workers, plugins

The repository is a multi-module Maven build, and the module names describe the architecture. presto-main holds the server entry point, com.facebook.presto.server.PrestoServer, and the sample configuration under presto-main/etc. presto-cli is the client. presto-spi and the connector modules are loaded as plugins, and the README says you modify plugin.bundles in config.properties to change which ones load. The UI (presto-ui) is a React and JSX project compiled into JAR resources; the README notes none of the Java code depends on it being compiled, so a missing console does not block the engine.

The data flow that matters for a first run is: the CLI talks to a coordinator, the coordinator reads catalog configuration, and a catalog such as hive resolves through a connector to an external service. In the sample configuration the Hive connector is mounted as the hive catalog, and the Hive plugin needs the location of a Hive metastore Thrift service, supplied through -Dhive.metastore.uri. If the metastore or HDFS is not reachable from your machine, the README offers SSH port forwarding and a SOCKS proxy via -Dhive.metastore.thrift.client.socks-proxy. That detail is worth noting: the project assumes the coordinator sits close to the data, not behind a corporate proxy.

How to install PrestoDB from source and run a first query

The README's install path is a Maven build from the project root. The first build downloads all dependencies into ~/.m2/repository and, in the README's words, "can take a considerable amount of time"; later builds are faster. The test suite takes several minutes, so the documented way to skip it during a build is:

bash
./mvnw clean install -DskipTests

If you build several Presto projects on one machine, the README warns that they may write to the shared M2 cache and cause build problems, and suggests a project-local cache through .mvn/maven.config:

bash
-Dmaven.repo.local=./.m2/repository

For running the server, the README gives an IntelliJ run configuration rather than a shell script: main class com.facebook.presto.server.PrestoServer, working directory presto-main, and a VM options string that includes -Dconfig=etc/config.properties, -Dlog.levels-file=etc/log.properties, -Xmx2G and a set of GC flags. The Hive plugin additionally needs a metastore address:

bash
-Dhive.metastore.uri=thrift://localhost:9083

Java 17 adds a second requirement. The README lists a block of --add-opens flags for internal JDK modules, starting with java.base/java.io=ALL-UNNAMED and java.base/java.lang=ALL-UNNAMED, and states plainly that the list is not comprehensive: more flags may be needed depending on which catalogs you configure. That is a real operational detail, not boilerplate. Reflective access failures show up as startup errors, and the fix is adding another --add-opens line.

Once the server is up, the CLI is a built artifact rather than an installed command:

bash
presto-cli/target/presto-cli-*-executable.jar

Two queries confirm the setup. SELECT * FROM system.runtime.nodes lists the nodes in the cluster, and SHOW TABLES FROM hive.default lists tables in the Hive database default, which works because the sample configuration mounts the Hive connector in the hive catalog. If the first query returns your node and the second returns tables, the coordinator, the plugin loading and the metastore connection are all working.

Where PrestoDB stops being the right tool

PrestoDB is a query engine, not a storage system and not a scheduler. The README and architecture document describe query execution; nothing in them describes ingestion, table maintenance, compaction or orchestration. If your problem is "get data in and keep it tidy", PrestoDB is a component you add after that problem is solved, not a solution to it.

The Java 17 --add-opens list is the clearest limitation in the README, and it is self-declared as incomplete. Every catalog you add can require another flag, which means the startup configuration is coupled to the connector set. That is manageable in a controlled deployment and annoying in a developer laptop setup that keeps growing connectors.

The build itself is a second constraint. Maven, Java 17, a first build that downloads a large dependency tree, and a test suite measured in minutes. For a team evaluating several engines in an afternoon, that friction matters. And the README does not document a packaged binary distribution, so anyone expecting a tarball and a systemd unit will not find one here. The installation documentation is where deployment instructions live, and the README defers to it.

PrestoDB compared with Trino, and what the fork changed

The obvious alternative is Trino, which began as a fork of Presto and now develops independently. The difference is not a feature checklist; it is governance and release cadence. PrestoDB lives under the Linux Foundation, as the LFX health score badge at the top of the README indicates, while Trino is governed by its own foundation. Practically, that means connector coverage and SQL behaviour can drift: a connector present in one project's module list may be missing or differently implemented in the other.

There is a naming trap worth stating because it costs people time. Searching for "presto" returns a payment card, a pressure cooker, a band and a dry cleaner. The engine's own identity is PrestoDB, and the repository is prestodb/presto. If you are comparing engines and land on a page about a transit card, you are on the wrong subject entirely.

Maintenance, releases and the Apache-2.0 licence

The repository is not archived, and the last push was on 2026-09-19, so the codebase is being pushed to regularly. Releases follow a numbered scheme rather than semantic versioning: 0.299 on 2026-08-28, 0.298.1 on 2026-06-17, and 0.298 on 2026-06-12. The patch release between 0.298 and 0.299 suggests fixes are backported rather than waiting for the next minor number.

Upgrade cost is where the versioning scheme bites. With 0.x numbering there is no compatibility promise encoded in the version, so a jump from 0.298 to 0.299 is not obviously safer than 0.298 to 0.299 plus a patch. The README does not document a rollback procedure or an upgrade path, so plan to read the release notes for each step and to test connectors you depend on, since plugin compatibility is where a distributed engine usually breaks.

The licence is Apache-2.0, which is permissive and includes a patent grant. That is a statement about the licence text, not legal advice; if you redistribute PrestoDB or bundle it into a product, have counsel review the NOTICES file and your own obligations.

Editorial conclusion

Adopt PrestoDB if you already run Hive, HDFS or another catalog-backed warehouse and need one SQL layer across several systems, and if your team is comfortable running Java services and building from source. Do not adopt it if you want a single-binary install, if your data fits in one Postgres instance, or if you have no Hive metastore to point the default configuration at. Before committing, build the project with ./mvnw clean install -DskipTests on a Java 17 machine, start the server from presto-main with the VM options in the README, and confirm that SELECT * FROM system.runtime.nodes returns your node and that SHOW TABLES FROM hive.default works against your metastore.

Frequently asked questions

What does PrestoDB mean and what is it?

PrestoDB is the name of the project: a distributed SQL query engine for big data, written in Java and licensed Apache-2.0. The README describes it as "a distributed SQL query engine for big data" and points to the installation documentation for deployment.

Does PrestoDB still exist?

Yes. The repository is not archived, the last push was on 2026-09-19, and the most recent release listed is 0.299 on 2026-08-28.

How do I install PrestoDB?

The README gives no packaged installer. It describes building from the project root with ./mvnw clean install, requiring Java 17 64-bit, Maven 3.6.3+ and Python 2.4+ for the launcher script, and then running com.facebook.presto.server.PrestoServer from the presto-main module with the documented VM options.

How do I use PrestoDB after starting the server?

The README starts the CLI from the built artifact presto-cli/target/presto-cli-*-executable.jar, then suggests running SELECT * FROM system.runtime.nodes to list cluster nodes and SHOW TABLES FROM hive.default to list tables in the Hive database default.

Official sources

  1. License: Apache-2.0
  2. prestodb/presto on GitHub
  3. Project website
  4. README
  5. Releases
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/prestodb-presto.svg)](https://hysenlabs.com/projects/prestodb-presto)
Community notes

Community notes