# Trino: a distributed SQL engine for querying data where it already lives

> Trino is the Apache-2.0 distributed SQL query engine formerly known as PrestoSQL. This review covers what it does, how the coordinator and workers split a query, how to build and run it from source, and where it stops being the right tool.

**trinodb/trino** — Official repository of Trino, the distributed SQL query engine for big data, formerly known as PrestoSQL (https://trino.io)

- Repository: https://github.com/trinodb/trino
- Website: https://trino.io
- Stars: 13,281 · Forks: 3,795
- Language: Java
- License: Apache-2.0
- Published: 2026-09-21 · Updated: 2026-09-21 · Language: en
- Canonical page: https://hysenlabs.com/projects/trinodb-trino

## What Trino solves, and who ends up running it

Trino is a distributed SQL query engine for big data analytics. The problem it addresses is the one that appears once an organisation has data spread across several systems: a Hive warehouse, an object store holding Iceberg or Delta Lake tables, a relational database, and a few files nobody wants to move. Without a query engine, each of those needs its own client, its own dialect and its own copy of the data. Trino presents them through one SQL interface, with connectors mapping each source into catalogs.

The audience follows from that. Trino is for data platform teams and analytics engineers who already have storage and compute separated, and who want the query layer to sit on top rather than inside. It is not aimed at application developers looking for an embedded database, and it is not aimed at teams who want a single binary that manages its own storage. The README points readers at the User Manual for deployment instructions, which tells you the project expects you to make deployment decisions yourself.

The repository layout reflects the split: core/ holds the engine, plugin/ holds connectors, client/ holds the CLI and JDBC pieces, service/ holds the server, and testing/ holds the development harness. That structure is worth reading before you plan an integration, because it shows where a new connector belongs and how much of the system is engine versus adapter.

## Coordinator, workers, and how a query actually moves

Trino is a distributed system with a coordinator and a set of workers. The coordinator parses and plans the SQL, then hands fragments of the plan to workers that execute them in parallel and exchange intermediate results. Connectors live in plugin/ and are loaded into the server; each catalog in a deployment is backed by one connector configuration.

The README gives a direct way to observe this. The development server is configured with the TPCH connector, and the README suggests two queries: one against system.runtime.nodes to list the nodes in the cluster, and one against tpch.tiny.region. The first query is the useful one for understanding the architecture, because it exposes the cluster membership that the coordinator manages. The second shows the catalog naming convention: catalog.schema.table.

Two design consequences are visible in the documentation. First, plugins are configured through a plugin.bundles property in config.properties, and each entry can be a path to a pom.xml, Maven coordinates, or a directory of JAR files. That makes the plugin set a deployment-time decision rather than something baked into the engine. Second, adding a catalog means adding a <catalog_name>.properties file to the catalog directory. There is no central registry; catalogs are files on disk. If you have run a system where connectors are configured through an API, this will feel deliberately plain.

## Building Trino from source and running a first query

The README describes Trino as a standard Maven project and gives a single build command. Run it from the project root. The first build downloads dependencies into ~/.m2/repository, which the README notes can take a while depending on connection speed; later builds are faster. Tests are skipped by default because the full suite takes considerable time.

```bash
./mvnw clean install -DskipTests
```

Build requirements are specific. Mac OS X or Linux, Java 25.0.1+ 64-bit, and Docker. The README warns that some npm packages used to build the web UI are only available for x86 architectures, so building on Apple Silicon requires Rosetta 2. It also warns that SELinux or other systems blocking write access to the local checkout must be turned off, because containers mount parts of the source tree.

For a first run, the README recommends the TpchQueryRunner class, which starts a development server configured with the TPCH connector. The VM option generally required is --add-modules jdk.incubator.vector, though other *QueryRunner classes may need more; the README says to check the air.test.jvm.additional-arguments property in the relevant module's pom.xml.

Once a server is up, start the CLI and run the query the README gives for inspecting the cluster:

```bash
client/trino-cli/target/trino-cli-*-executable.jar
```

```sql
SELECT * FROM system.runtime.nodes;
```

You should see the nodes that make up the cluster. Then query the TPCH connector to confirm a catalog is reachable:

```sql
SELECT * FROM tpch.tiny.region;
```

If you prefer to run the full server rather than a query runner, the README gives the main class io.trino.server.DevelopmentServer with the config and log properties files under etc/, a working directory of the trino-server-dev subdirectory, and the trino-server-dev module on the classpath. To change which plugins the development server loads, edit plugin.bundles in config.properties.

## Where Trino is the wrong tool

The clearest limitation is that this is a query engine, not a store. Nothing in the README describes Trino owning durable data; it reads through connectors. If your workload is transactional, with frequent small writes and strict latency guarantees, a distributed query engine designed around analytical scans is the wrong layer. You would be adding a coordinator and a network hop in front of a database that already answers those queries.

The development path has its own friction. Building requires Java 25.0.1+ and Docker, and the README explicitly says SELinux or similar write-blocking systems must be turned off for containers to mount parts of the source tree. On Apple Silicon, the web UI build needs Rosetta 2 because some npm packages are x86-only. These are not incidental details; they change who can build the project on a laptop.

Operationally, the plugin model cuts both ways. Plugins are resolved from local or remote Maven repositories, or from directories of JAR files, and catalogs are individual properties files. That is flexible, but it means the set of connectors a cluster exposes is a deployment artifact you maintain. The README does not document rollback or upgrade procedures for a running cluster; the User Manual is where deployment instructions live. If you need a documented upgrade path before you start, that is something to confirm first.

## Trino compared with Spark and DuckDB

The two comparisons people search for most are against Spark and DuckDB, and they differ in ways that matter more than feature lists.

Spark is a general data processing framework. It runs jobs, it has its own execution model and its own APIs beyond SQL, and it is commonly used to transform data as well as query it. Trino is narrower: a SQL query engine. If your work is a pipeline that reads, transforms and writes data on a schedule, Spark's model fits that shape. If your work is interactive SQL against sources that already exist, Trino's model fits that shape, and you avoid writing a job for every question.

DuckDB is an embedded analytical database. It runs in-process, which makes it excellent for local analysis and for a single machine. Trino is a distributed server with a coordinator and workers. The trade is scale against setup: DuckDB needs no cluster, Trino needs one, and Trino is the one that spreads a query across nodes. If your dataset fits comfortably on one machine, the distributed design is overhead you are paying for nothing. The README's own development setup, with a CLI connecting to a server, illustrates the difference immediately: there is a server to run.

A note on the Presto comparison, since it appears in the search data. The repository description states that Trino is the engine formerly known as PrestoSQL. The project has its own release line, currently at 483 as of 2026-07-18, and its own documentation site.

## Maintenance cadence, licence, and what upgrades cost

The repository is not archived, and the last push was on 2026-09-21. Releases are frequent: 483 on 2026-07-18, 482 on 2026-06-25, and 481 on 2026-05-12. That cadence is the practical upgrade cost. A project shipping a release roughly every six to eight weeks means a cluster that stays current is being upgraded several times a year, and the README does not describe an in-place upgrade procedure; it points to the User Manual for deployment instructions. Plan for the release notes to be part of your routine rather than an annual event.

Trino is licensed under Apache-2.0. The LICENSE file sits at the repository root. For most users this is a permissive licence with no copyleft obligation on your own code, but the connectors you enable may carry their own terms, and the licence of the engine does not settle the licence of the systems it connects to. That is a question for your own legal review, not something the repository answers.

One maintenance detail worth noting: the README states that Trino supports reproducible builds as of version 449, and the repository carries a badge for reproducible builds. Reproducibility is a supply-chain property, not a feature you use, but it does mean a build you produce can in principle be compared against another. The README does not explain how to verify it.

## Conclusion

Adopt Trino when you need one SQL surface across data that already sits in Hive, Iceberg, Delta Lake or a relational database, and you have the operational capacity to run a coordinator plus workers. Do not adopt it as an OLTP store or as a replacement for a warehouse that owns its own storage; Trino reads and writes through connectors rather than owning the data. Before committing, verify which connectors your sources need, whether your deployment target is covered by the documented deployment instructions, and whether you can meet the build requirements of Java 25.0.1+ and Docker.

## FAQ

### What is Trino used for?

Trino is a distributed SQL query engine for big data analytics. It presents data held in other systems through connectors and catalogs, so you can query them with one SQL interface rather than moving the data first.

### Is Trino free to use?

The repository is licensed under Apache-2.0, and the LICENSE file is at the project root. That covers the engine itself; the systems you connect to through connectors have their own licences.

### How do I install Trino?

The README gives a source build rather than a binary install: run ./mvnw clean install -DskipTests from the project root. It requires Mac OS X or Linux, Java 25.0.1+, and Docker. The README directs readers to the User Manual for deployment instructions.

### How do I use the Trino CLI?

After building, start the CLI from client/trino-cli/target/trino-cli-*-executable.jar. The README shows running SELECT * FROM system.runtime.nodes to list the cluster nodes, and SELECT * FROM tpch.tiny.region against the TPCH connector.

### How do I set up Trino for development?

The README recommends running the TpchQueryRunner class, which starts a development server configured with the TPCH connector. The VM option generally required is --add-modules jdk.incubator.vector, and other *QueryRunner classes may need additional options listed in their module's pom.xml.

## Sources

- [License: Apache-2.0](https://github.com/trinodb/trino/blob/master/LICENSE)
- [Project website](https://trino.io)
- [README](https://github.com/trinodb/trino/blob/master/README.md)
- [Releases](https://github.com/trinodb/trino/releases)
- [trinodb/trino on GitHub](https://github.com/trinodb/trino)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/trinodb-trino
