Apache Drill: querying NoSQL and Hadoop storage with SQL and no fixed schema
Apache Drill is a distributed MPP query layer for self describing data
At a glance
- What is it?
- Apache Drill is a distributed MPP query engine for SQL and alternative languages over schema-free data stores including Hadoop, Hive and Parquet. The distribution ships with SASL, Kerberos and OpenSSL, requires no predefined schema, and the Dockerfile defaults to OpenJDK 17 for both build and runtime.
- Who is it for?
- Apache Drill suits teams running Hadoop or object storage who want ad-hoc SQL over Parquet, JSON or Hive without first loading data into a relational database. The SASL and Kerberos support makes it viable in environments with enterprise authentication.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 3 days ago.
- What is it written in?
- Mainly Java, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 6, 2026, and from our analysis. They are not legal advice.
Editorial analysis
SQL and alternative languages over self-describing data stores
Apache Drill is a distributed MPP query layer that targets NoSQL and Hadoop data storage systems without requiring a predefined schema. The repository description calls it a query layer for self-describing data, meaning the engine reads the schema from the data itself rather than from a central schema registry. A JSON file on HDFS, a Parquet file in object storage, or a Hive table are all queryable through SQL once Drill has access to the underlying storage.
The primary interface is SQL, but the README describes support for alternative query languages without naming them. JDBC appears in the repository topics, which signals a standards-based driver for connecting business intelligence tools, query clients and application code. Parquet and Hive are also in the topics list, confirming that columnar storage formats and Hive metastores are first-class targets.
This design addresses a gap in many Hadoop deployments: data accumulates in files on HDFS that a relational database cannot reach directly. Drill sits between that storage layer and the SQL client, translating queries into distributed read operations without requiring data movement.
Google's Dremel paper shaped the MPP execution model
The README links directly to Google's Dremel paper and credits it as a partial inspiration. Dremel described a system for interactive ad-hoc analysis over large, read-only nested datasets, using a tree of serving nodes to scatter queries in parallel and gather partial results. Each node processed a slice of columnar data independently, and the coordinator aggregated the pieces into a final result set.
Drill follows that general shape. A query enters the system at a coordinator, gets compiled into a logical plan and then into a physical plan, and is dispatched to execution nodes. The repository top level has a logical/ directory and an exec/ directory, which correspond to those two compilation stages, and a protocol/ directory that governs communication between nodes. The distribution/ directory holds the assembled binaries. According to the README's link list, the docs/ directory includes material on submitting logical and distributed physical plans, which confirms the two-plan compilation structure.
One consequence of the two-plan model is that query failures can appear at either the logical planning stage or the physical execution stage, and those are diagnosed differently. Drill's official documentation at drill.apache.org covers the diagnostic distinction between them.
Building with Maven on OpenJDK 17, with a Docker two-stage path
Building Drill from source is a Maven operation. The Dockerfile at the repository root names the default build base image as maven:3-openjdk-17, which means OpenJDK 17 is the current floor for compilation. The build command in that file is:
mvn clean install -DskipTestsSkipping tests with -DskipTests cuts build time on a first run because the test suite includes integration tests that expect a running cluster. The resulting binaries are placed into the distribution/ directory, one of the top-level modules in the repository alongside exec/, logical/ and the other pipeline components.
The Dockerfile builds in two stages. The first stage uses the Maven image to compile and produce the distribution artifacts. The second stage copies those artifacts into a minimal openjdk:17 runtime image, discarding the source code, Maven cache and intermediate build artifacts. That keeps the final image smaller than shipping everything from the build environment.
For a Java 11 target, the Dockerfile comment gives the override:
{docker|podman} build \
--build-arg BUILD_BASE_IMAGE=maven:3.8.2-openjdk-11 \
--build-arg BASE_IMAGE=openjdk:11 \
-t apache/drill-openjdk-11Both docker and podman are accepted as the container runtime, and the braces in the command above denote a choice rather than a syntax requirement. Passing the BUILD_BASE_IMAGE and BASE_IMAGE arguments overrides both stages of the build. Without those arguments, both default to OpenJDK 17. The official documentation at drill.apache.org/docs/running-drill-on-docker/ covers the Docker deployment configuration beyond what the Dockerfile alone explains.
SASL, Kerberos and OpenSSL ship inside every distribution package
The export control section of the README itemizes the cryptographic libraries in the distribution. Java SE Security packages handle authentication, authorization and secure socket communication. The Jetty Web Server provides HTTPS for the web interface. Cyrus SASL libraries, Kerberos Libraries and OpenSSL Libraries provide SASL-based authentication and SSL communication for inter-node and client connections.
These are not optional modules or separate downloads. They are included in the standard distribution, which is why the US Bureau of Industry and Security has classified the package under Export Commodity Control Number ECCN 5D002.C.1. That classification covers information security software using asymmetric cryptographic algorithms.
For deployments outside the United States, the README instructs readers to check the destination country's import and use restrictions on encryption software before proceeding, pointing to wassenaar.org as the reference. The practical side is that a Drill cluster in a regulated environment can use Kerberos for authentication and SSL for in-flight encryption without adding extra libraries, but the export compliance check applies regardless of deployment scale.
YARN integration and a metastore module for production clusters
Two top-level directories point at production deployment concerns. The drill-yarn/ directory holds Apache YARN integration, which lets Drill run as a YARN application on a managed Hadoop cluster. YARN controls resource allocation across a cluster, and deploying Drill as a YARN application means query workers are governed by the same resource policies as other Hadoop jobs on that cluster.
The metastore/ directory manages metadata about data sources. When Drill queries Parquet or Hive data repeatedly, a metastore allows it to cache schema and statistics information between queries rather than discovering that information from the data on each run. Both modules in the repository indicate that Drill targets multi-tenant Hadoop deployments where resource management and metadata persistence both matter.
For a single-node evaluation or a development trial, neither module is necessary, and the YARN and metastore configuration layers add setup cost without benefit outside a real cluster environment. A standalone or Docker-based setup is the documented path for development and trial use, with instructions on the Apache Drill website.
The README delegates nearly every installation detail to the project website
Almost every step after cloning the repository points to drill.apache.org. Installation instructions for remote execution mode are there, not in the README. Docker usage instructions are at drill.apache.org/docs/running-drill-on-docker/. Example queries and sample data explanations are there. The process for submitting logical and distributed physical plans is there. The README names each resource but does not reproduce its content.
The top-level repository carries a sample-data/ directory, so test files are available locally without a separate download. The docs/ directory contains Environment.md for setting up a development environment and DevDocs.md for broader developer documentation. But operational setup, configuration file syntax, cluster sizing and storage plugin configuration all live on the website, not in the repository.
That delegation works when the website is available and current, but the README alone cannot guide a first installation. This matters in air-gapped environments: the repository gives you the source and the build tooling, but the operational documentation requires a separate step to obtain.
Release 1.22.0 in June 2025, with commits continuing into October 2026
The most recent release tag in the repository is drill-1.22.0, published on 2025-06-29. Before that, 1.21.2 appeared on 2024-06-23 and 1.21.1 on 2023-08-14. The pattern is roughly one minor release per year with patch releases between them, and the gap between 1.21.2 and 1.22.0 is about a year. The repository is not archived, and the last push was on 2026-10-05.
Community channels are the Apache Drill mailing list at drill.apache.org/mailinglists/, a Slack channel at apache-drill.slack.com and a Stack Overflow tag. Contributions and bug reports go through Apache Foundation infrastructure. The distribution is available on Maven Central under the group ID org.apache.drill with the artifact ID distribution.
The LICENSE file in the repository is Apache-2.0. The NOTICE and HEADER-2.0.txt files at the repository root carry the attribution requirements for redistribution. No release in 2026 has been tagged yet as of the last push date, so 1.22.0 is the current stable binary available on Maven Central.
Editorial conclusion
Apache Drill suits teams running Hadoop or object storage who want ad-hoc SQL over Parquet, JSON or Hive without first loading data into a relational database. The SASL and Kerberos support makes it viable in environments with enterprise authentication. It is the wrong choice for transactional workloads, sub-second latency on a single node, or any deployment where the team cannot absorb a documentation gap between the repository and actual setup. Before adopting, match the OpenJDK 17 Dockerfile baseline against the JDK on your cluster hosts, and check the export restrictions if the binaries cross a national border.
Frequently asked questions
What is Apache Drill?
Apache Drill is a distributed MPP query engine that supports SQL and alternative query languages against NoSQL and Hadoop data storage systems. It requires no predefined schema, reading the schema directly from data formats such as Parquet and JSON, and exposes a JDBC interface for SQL client connections.
How do I build Apache Drill from source?
Building Drill requires Maven and OpenJDK 17. The Dockerfile in the repository uses the command mvn clean install -DskipTests to compile the project and produce the distribution binaries, with the -DskipTests flag used to skip integration tests on a first build.
Does Apache Drill support Parquet files?
Parquet appears in the repository topics alongside Hive and Hadoop, indicating it is a supported data source. The self-describing data model means Drill reads the Parquet schema from the file itself rather than requiring a separate schema definition before querying.
Can Apache Drill run on Docker?
The repository includes a Dockerfile that builds Drill using a two-stage image, defaulting to OpenJDK 17 for both build and runtime stages. The official documentation at drill.apache.org/docs/running-drill-on-docker/ covers the full Docker deployment configuration.
What is the latest version of Apache Drill?
The most recent release tagged in the repository is drill-1.22.0, published on 2025-06-29. It is distributed on Maven Central under the group ID org.apache.drill with the artifact ID distribution.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/apache-drill)