Self-hosted service
Alluxio/alluxio avatar
Alluxio/alluxio

Alluxio Open Source: a distributed cache between Presto, Spark and your storage

Alluxio, data orchestration for analytics and machine learning in the cloud

7,246 stars2,931 forksJavaApache-2.0

At a glance

What is it?
Alluxio Open Source is a Java distributed caching layer that puts a common filesystem interface in front of many storage systems. It suits analytics teams that keep re-reading the same structured data, and it is not the AI training tier that the enterprise edition targets.
Who is it for?
Adopt Alluxio Open Source if your analytics engines repeatedly scan the same structured datasets and you want one namespace over several storage systems without changing engine code. Do not adopt it for large-scale model training or inference: the README states the enterprise edition uses a decentralized metadata service for that and scales to tens of billions of files.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 28 days ago.
What is it written in?
Mainly Java, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gap Alluxio fills between compute engines and storage

Analytics engines each speak their own dialect of storage. A Presto cluster, a Spark job and a Trino deployment all need to read the same datasets, and each one carries its own connector, its own credential handling, and its own idea of where the data lives. When the same table sits in object storage and is scanned by three engines, the bytes cross the network three times. Alluxio Open Source (formerly Tachyon) is described in the README as a Distributed Caching Platform for large-scale data that bridges the gap between computation frameworks and storage systems, letting computation applications reach numerous storage systems through a common interface. The audience is narrow and specific: teams running structured data analytics on Presto, Spark or Trino, where the same data is read repeatedly and the storage backend is slower than the compute. The README says the open source edition is recommended for testing, development and small-scale production environments and is available for free without support. That last phrase matters more than it looks. You are adopting a cache with no vendor behind the free tier.

How the caching and namespace layers actually fit together

The architecture visible in the repository is a master and worker split. The master owns metadata: the namespace, file-to-block mappings, and the worker registry. Workers own data: they hold cached blocks and serve them to clients. The Docker example in the README launches exactly these two roles as separate containers, with the worker told where the master is through a Java system property, which tells you the coordination path runs from worker to master at startup. The repository layout supports this reading: core/ holds the shared logic, underfs/ holds the connectors to underlying storage systems, job/ holds the job service, and table/ holds the catalog-facing pieces. Clients do not talk to the underlying store directly for cached data; they ask the master where a block lives and then read it from the worker holding it. The underfs/ directory is the seam where the abstraction lives, which is why the same client can sit in front of different backends. The README also notes that Alluxio originated as the Tachyon research project at UC Berkeley's AMPLab and was the data layer of the Berkeley Data Analytics Stack, and it points to Haoyuan Li's dissertation titled Alluxio: A Virtual Distributed File System. That name is the clearest one-sentence description of the design: a filesystem assembled virtually over other filesystems, with caching as the performance mechanism and the namespace as the compatibility mechanism.

Installing Alluxio with Docker and running a master and worker

The README gives three installation routes: prebuilt binaries from the download page, Homebrew on macOS, and Docker. The Docker path is the one spelled out in full, so it is the one to follow first. Create a network so the two containers can resolve each other by name, and a volume so the worker's underlying storage survives a restart.

bash
docker network create alluxio_nw
docker volume create ufs

Then start the master. It publishes port 19999, which is where the web UI answers.

bash
docker run -d --net=alluxio_nw \
    -p 19999:19999 \
    --name=alluxio-master \
    -v ufs:/opt/alluxio/underFSStorage \
    alluxio/alluxio master

Start the worker with a ramdisk size, passing the master hostname as a system property. The README's example uses 1G.

bash
export ALLUXIO_WORKER_RAMDISK_SIZE=1G
docker run -d --net=alluxio_nw \
    --shm-size=${ALLUXIO_WORKER_RAMDISK_SIZE} \
    --name=alluxio-worker \
    -v ufs:/opt/alluxio/underFSStorage \
    -e ALLUXIO_JAVA_OPTS="-Dalluxio.worker.ramdisk.size=${ALLUXIO_WORKER_RAMDISK_SIZE} -Dalluxio.master.hostname=alluxio-master" \
    alluxio/alluxio worker

What you should see is two running containers and a reachable master UI on port 19999. If you prefer a local install, the README offers `brew install alluxio` on macOS, and the binary download page for everything else. The README then defers to the Guide to Get Started for a simple example rather than reproducing one, so treat the Docker pair above as the smoke test and the documentation site as the next step.

For a JVM project that needs the client, the README recommends the shaded artifact because its jar bundles dependencies to avoid conflicts.

xml
<dependency>
  <groupId>org.alluxio</groupId>
  <artifactId>alluxio-shaded-client</artifactId>
  <version>2.6.0</version>
</dependency>

The README notes that alluxio-core-client-fs (the Alluxio Java file system API) and alluxio-core-client-hdfs (the HDFS-compatible API) are both included inside alluxio-shaded-client, so pulling the shaded jar is usually enough.

The 100 million file ceiling and the edition split

The most consequential number in the README is the scale limit: the open source edition scales to manage up to 100 million files. That is a real ceiling, not a soft guideline, and it comes from the centralized metadata design. Every file and block mapping lives in the master, so metadata operations concentrate there. A workload with a few million large Parquet files will be comfortable. A workload with hundreds of millions of small files will not, and the failure mode is metadata pressure on the master rather than a clean error at the client. The second constraint is the workload shape. The README states plainly that the open source edition is purpose-built for analytics workloads, accelerating structured data analytics, and that for AI and machine learning workloads including model training, distribution and inference at scale, the enterprise edition provides a fundamentally different architecture with a decentralized metadata service. It names FUSE-based POSIX integration and compatibility with PyTorch, TensorFlow and Ray as enterprise features. If your plan is to mount Alluxio as a POSIX filesystem for a training loop, the open source edition is the wrong tool by the project's own description. The third constraint is support. Free, no support, recommended for testing, development and small-scale production is the project's own positioning. Running it as the shared cache under a large production warehouse means you are the support contract.

Alluxio compared with JuiceFS, Redis and HDFS

The comparison that comes up most often is Alluxio versus JuiceFS. Both put a filesystem interface in front of object storage and both cache data near compute, so the question is where the metadata lives. Alluxio Open Source keeps a centralized master for the namespace and block locations, which is what produces the 100 million file figure and the scaling story the enterprise edition addresses with a decentralized metadata service. JuiceFS, by contrast, is built around a metadata engine such as Redis or a SQL database that you supply and operate. That shifts the scaling question to the metadata backend you choose and adds an operational dependency Alluxio does not have. The second comparison, Alluxio versus Redis, is a category error worth naming. Redis is an in-memory data structure store; Alluxio is a caching filesystem with a namespace, block management and storage connectors. Redis does not present an HDFS-compatible API or a file namespace, so it does not replace Alluxio for engine integration. The third, Alluxio versus HDFS, is the one the project's own history answers. HDFS is a storage system with its own replication and its own on-disk format. Alluxio sits above storage systems, including HDFS, and caches them. The README's artifact list includes alluxio-core-client-hdfs precisely because Alluxio presents an HDFS-compatible API to engines that expect one. If your data already lives in HDFS and HDFS is fast enough, Alluxio adds a layer without removing one.

Releases, licence and what maintenance actually looks like

The release history is worth reading carefully. The listed releases are v2.9.2 from 2023-03-01, v2.9.3 from 2023-03-24, and v2.9.4 from 2024-06-11. The last push to the main branch was on 2026-09-01, so the repository is not archived and work continues, but the newest tagged release is more than two years older than that push. That pattern is common in projects where the open source edition is a companion to a commercial product: development continues while releases slow. For an adopter, the practical consequence is that you should expect to build from source or track the main branch if you need a fix that landed after v2.9.4, and you should check whether the specific behaviour you depend on is in the tagged release. On licensing, the project is Apache-2.0, which is a permissive licence that allows commercial use and modification, and the repository carries a NOTICE file alongside the LICENSE. The README asks contributors to state that their contribution is original work and that they license it under the project's open source license. None of this is legal advice; if you are embedding Alluxio in a product, read the LICENSE and NOTICE files in the repository root and get your own review. The governance detail that matters operationally is that the Alluxio Open Source Foundation owns the project and a Project Management Committee operates it, with membership details in a wiki page. That is a real governance structure rather than a single-vendor repository, but it does not by itself promise release cadence.

What to check before you put Alluxio in front of a warehouse

Three things are worth verifying in your own environment before you commit. First, count your files. The README states the open source edition scales to manage up to 100 million files; if your dataset is well under that, the centralized master is not a problem, and if it is over, you are looking at the wrong edition. Second, confirm your engine is in the supported set. The README names Presto, Spark and Trino as the data-intensive computation engines the open source edition is widely adopted with. If your engine is not among them, the integration path is the Java file system API or the HDFS-compatible API, and you should prototype that before planning around it. Third, decide whether you need POSIX semantics. The README places FUSE-based POSIX integration in the enterprise edition, so any tool that expects to open files through a mount point is out of scope for the open source edition. On the operational side, the master is a coordination point and the worker ramdisk is memory you are taking from something else; the Docker example sets it to 1G via ALLUXIO_WORKER_RAMDISK_SIZE, and sizing that against your working set is the difference between a cache that helps and a cache that evicts constantly. The README does not document rollback or downgrade procedures, so plan an upgrade path from v2.9.4 deliberately rather than assuming one exists.

Editorial conclusion

Adopt Alluxio Open Source if your analytics engines repeatedly scan the same structured datasets and you want one namespace over several storage systems without changing engine code. Do not adopt it for large-scale model training or inference: the README states the enterprise edition uses a decentralized metadata service for that and scales to tens of billions of files. Before committing, verify your file count against the stated 100 million limit, confirm that your engine is among the supported ones, and check the release history, since the newest release listed is v2.9.4 from 2024-06-11 while the last push to main was on 2026-09-01.

Frequently asked questions

How does Alluxio work?

It runs a master that holds the namespace and block metadata and workers that hold cached data, and it presents a common filesystem interface so engines such as Presto, Spark and Trino can read from many storage systems without changing their connectors.

Is Alluxio open source?

Yes. This repository is the open source edition, licensed under Apache-2.0, owned by the Alluxio Open Source Foundation and operated by a Project Management Committee. The README states it is available for free without support and is recommended for testing, development and small-scale production environments.

What is Alluxio?

Alluxio Open Source, formerly Tachyon, is a Distributed Caching Platform for large-scale data that bridges computation frameworks and storage systems through a common interface. It originated as a research project at UC Berkeley's AMPLab and was the data layer of the Berkeley Data Analytics Stack.

How does Alluxio compare with JuiceFS?

Both sit in front of object storage and cache near compute, but Alluxio Open Source uses a centralized master for namespace and block metadata, which the README ties to a limit of up to 100 million files. JuiceFS relies on a metadata engine you supply and operate, which moves the scaling question to that backend.

How does Alluxio compare with HDFS?

HDFS is a storage system with its own replication and on-disk format, while Alluxio caches storage systems including HDFS and exposes an HDFS-compatible API through the alluxio-core-client-hdfs artifact. If HDFS is already fast enough for your scans, Alluxio adds a layer rather than replacing one.

Official sources

  1. Alluxio/alluxio on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/alluxio-alluxio.svg)](https://hysenlabs.com/projects/alluxio-alluxio)