# Fluid: a Kubernetes-native layer for datasets and the cache runtimes behind them

> Fluid turns a pile of object storage into something a Spark job or TensorFlow process can treat as a local filesystem, then scales the cache that makes it fast. It graduated from CNCF Sandbox to Incubating in January 2026.

**fluid-cloudnative/fluid** — Fluid, elastic data abstraction and acceleration for BigData/AI applications in cloud. (Project under CNCF) 

- Repository: https://github.com/fluid-cloudnative/fluid
- Website: https://fluid-cloudnative.github.io/
- Stars: 1,986 · Forks: 1,262
- Language: Go
- License: Apache-2.0
- Published: 2026-10-06 · Updated: 2026-10-06 · Language: en
- Canonical page: https://hysenlabs.com/projects/fluid-cloudnative-fluid

## Two concepts carry the whole design

Most Kubernetes operators introduce a handful of custom resources and hope the relationship between them is obvious. Fluid explains its two in the README, and the explanation is worth reading closely.

A Dataset, per the README, is a set of logically related data that computing engines such as Spark for analytics or TensorFlow for AI applications can use. Managing datasets, the README argues, needs more than access: security, version management, and data acceleration are all separate dimensions. Fluid deliberately starts with data acceleration and treats the rest as later concerns.

A Runtime is the enforcement layer. It handles dataset isolation and sharing, provides version management, and enables acceleration by defining a set of interfaces that the dataset lifecycle runs through. The significance is that the management and acceleration logic lives behind those interfaces rather than being hardcoded, which is what allows several different cache engines to sit underneath the same abstraction.

So a Runtime is not a cache itself. It is the thing that decides which cache runs, where it runs, and how a pod's filesystem view gets wired up to it. That indirection is the project's main design bet, and it is also why the repository carries one controller directory per engine.

## One controller per cache engine, all behind the Runtime interface

The Makefile at the repository root is the clearest inventory of what the project actually ships. Each engine gets its own controller image under the fluidcloudnative organisation:

```bash
IMG_REPO ?= fluidcloudnative
DATASET_CONTROLLER_IMG ?= ${IMG_REPO}/dataset-controller
APPLICATION_CONTROLLER_IMG ?= ${IMG_REPO}/application-controller
ALLUXIORUNTIME_CONTROLLER_IMG ?= ${IMG_REPO}/alluxioruntime-controller
JINDORUNTIME_CONTROLLER_IMG ?= ${IMG_REPO}/jindoruntime-controller
JUICEFSRUNTIME_CONTROLLER_IMG ?= ${IMG_REPO}/juicefsruntime-controller
THINRUNTIME_CONTROLLER_IMG ?= ${IMG_REPO}/thinruntime-controller
```

The list continues with cacheruntime-controller, efcruntime-controller, and vineyardruntime-controller. Alluxio is the engine the repository topics list, and JuiceFS, Jindo, Cache, EFC, Vineyard, and Thin each get a dedicated sample directory under samples/, alongside samples for dataload, hpa, cronhpa, co-locality, knative, and dawnbench.

That ThinRuntime entry is worth pausing on, because it represents a different approach from the others. The specific engines are integrations with a known project. ThinRuntime is a general path for bringing your own client, and version 1.0.8 added 3FS and Curvine support through it rather than as new first class engines.

Two other controllers are worth noting because they are not cache engines at all. The dataset controller manages Dataset resources, and the application controller manages the pods that consume them. Everything else in the tree, including csi/, charts/, api/, and sdk/, exists to support those.

The samples directory also tells you what problems the project is aimed at. samples/co-locality is about placing compute next to the data it needs, and samples/knative targets serverless Kubernetes.

## Data affinity is the feature, not just the cache

It would be easy to describe Fluid as a caching layer for object storage, and it is partly that. But the feature list in the README names elasticity and scheduling as one bullet, and that combination is the more interesting claim.

The README describes the goal as enhancing data access performance by combining data caching with elastic scaling, portability, observability, and data affinity-scheduling. Read together, those mean the project wants to know where your data physically lives and where your job runs, then move or scale the job to match. A Spark executor reading from a bucket on the other side of the continent is not a storage problem, it is a placement problem, and a cache that does not scale to the right node does not fix it.

Two other features complete the picture. Dataset abstraction gives a unified abstraction over multiple storage sources with observability built in, described as helping users evaluate whether they need to scale the cache system at all. That is a deliberately modest framing for a caching project: the observability exists so you can decide not to use the cache.

The scalable cache runtime then offers a unified access interface across different runtimes so third-party storage systems can be reached through it, and automated data operations provide several data operation modes intended to plug into automated operations systems.

The last feature on the list is runtime platform agnostic. The README names native, edge, serverless Kubernetes clusters, and multi-cluster Kubernetes environments as supported targets. That claim is mostly about FUSE sidecar patterns rather than a cluster-wide cache, which is consistent with how the project actually works.

## CNCF Incubating as of January 2026, on a v1.0 line

The timeline in the README is short and recent. Fluid was accepted as a CNCF Sandbox project on 2021-04-27 after a majority vote of the Technical Oversight Committee. On 2026-01-08 the TOC accepted it as an incubating project, which the README marks as a major milestone in its maturity and community adoption.

The releases since then are steady rather than dramatic. v1.0.6 on 2025-07-12, v1.0.7 on 2025-09-21, and v1.0.8 on 2025-10-31, each with details in CHANGELOG.md. The repository's Makefile is already stamped one version ahead:

```bash
VERSION := v1.1.0
BUILD_DATE := $(shell date -u +'%Y-%m-%d_%H:%M:%S')
GIT_COMMIT := $(shell git rev-parse HEAD)
PACKAGE := github.com/fluid-cloudnative/fluid
```

The project also publishes the research behind it, which is unusual and useful. Two papers are cited: the ICDE 2022 conference paper on dataset abstraction and elastic acceleration for cloud native deep learning training jobs, and a journal version in IEEE TPDS volume 34 issue 11 from 2023 on high level data abstraction and elastic data caching for data intensive AI applications. Reading those is a faster way to understand the design intent than the README manages.

The project runs bi-weekly community meetings with notes and recordings published in a separate community repository, and the governance files in the root, including GOVERNANCE.md, OWNERS, and MAINTAINERS_COMMITTERS.md, show a mature CNCF layout. Apache-2.0 is the licence, and the last push was on 2026-09-16.

## Sidecar lifecycle and prefetching are where the recent work went

Version 1.0.8 is the release worth reading in full, because its four new features all come from real operational friction.

The first is support for Kubernetes Native Sidecar mode. In the FUSE Sidecar deployment model, Fluid starts a sidecar pod that mounts the accelerated dataset path into your workload. The release notes describe the traditional sidecar as limited in lifecycle management, startup sequencing, and resource isolation, and the native sidecar injection scheme as the fix. If you have run into a sidecar starting before its dataset is ready, or surviving past its consumer, this is the change you were waiting for.

The second is a KubeSphere extension adding a web dashboard for Fluid, covering dataset, runtime, and dataload management. Given how many CRDs this project involves, a dashboard is not decoration.

The third opens PodMetadata configuration so labels and annotations on FUSE pods can be set dynamically, which is what you need when cluster policy depends on pod metadata. The fourth adds 3FS and Curvine storage support through ThinRuntime.

The enhancements are more telling about the project's current pressure points. Node scheduling synchronization gained a switch to disable it, because the scheduling work was costing real time. The ThinRuntime controller's reconcile rate limit now defaults to 0, which means it reconciles without throttling and reacts faster at the cost of more API traffic. Mount point detection was made more reliable, and there is an explicit note about mount point detection reinforcement.

Earlier releases show the same direction. v1.0.7 added custom Pod lifecycle hooks to ThinRuntimeProfile and let multiple sidecar-mounted pods coexist on one node by having the webhook allocate unique hostpaths. v1.0.6 added a prefetch sidecar to accelerate model file loading for inference services, and a JuiceFS FUSE seamless upgrade that cleans up idle pods when the FUSE image changes under a cleanPolicy of onFuseChange.

## Helm 3 installs it, and go.mod pins Kubernetes to 0.29

The prerequisites are short: Kubernetes greater than 1.16 with CSI support, Golang 1.18 or later, and Helm 3. Installation runs through Helm, with the chart published to Artifact Hub, and the README points at a Get Started guide for standing up a test cluster.

There is a mismatch worth knowing about before you build. The README says Golang 1.18+, and the actual module requirement is much higher:

```bash
module github.com/fluid-cloudnative/fluid

go 1.25.0

toolchain go1.25.12
```

If you are contributing rather than installing, use the version in go.mod. The README's prerequisite list reads more like a consumer's requirement than a developer's.

The more consequential thing in go.mod is the block of replace directives that pins every Kubernetes library to the same version:

```bash
replace k8s.io/api => k8s.io/api v0.29.15
replace k8s.io/apiextensions-apiserver => k8s.io/apiextensions-apiserver v0.29.15
replace k8s.io/apimachinery => k8s.io/apimachinery v0.29.15
replace k8s.io/client-go => k8s.io/client-go v0.29.15
```

That means Fluid builds against Kubernetes 0.29 libraries regardless of what else your cluster or your Go project resolves, so it will not conflict with a newer client-go in the same build, but it also means you do not get newer library behaviour. A vendor directory is committed as well, which is consistent with pinning this tightly and with the enterprise provenance of the project.

The documentation lives in the repository under docs/, with separate English and Chinese tables of contents, plus a project site at fluid-cloudnative.github.io. The README's own quick start is a link rather than a set of commands, which is the right call for something this configuration heavy.

## Conclusion

Fluid is at its best when a Spark or TensorFlow job spends more time waiting on object storage than computing, because the value it adds is making that storage behave like a local directory and then scaling the cache that makes it fast. The Dataset and Runtime split is the part worth understanding before anything else, since everything else follows from it. Version 1.0.8 adds Kubernetes Native Sidecar support, a KubeSphere dashboard extension, 3FS and Curvine clients, and a reconcile rate limit defaulting to zero for faster response, and the repository's Makefile is already stamped for v1.1.0. The realistic caveats are the Kubernetes prerequisite above 1.16 with CSI support, Helm 3 as the install path, and the fact that go.mod pins every k8s.io dependency to v0.29.15, so you are adopting a specific dependency baseline rather than tracking the newest libraries.

## FAQ

### What is Fluid in CNCF?

Fluid is a Kubernetes-native Distributed Dataset Orchestrator and Accelerator for data intensive applications such as big data and AI workloads. The CNCF Technical Oversight Committee accepted it as a Sandbox project on 2021-04-27 and promoted it to Incubating status on 2026-01-08. It is written in Go and licensed under Apache-2.0.

### What is the difference between a Dataset and a Runtime in Fluid?

A Dataset is the custom resource describing a set of logically related data that engines like Spark or TensorFlow consume. A Runtime is the layer that enforces isolation and sharing for that dataset, handles version management, and wires up the interfaces that a specific cache engine implements. Multiple runtimes such as Alluxio, JuiceFS, Jindo, and Thin back the same Dataset abstraction.

### How do you install Fluid?

You need Kubernetes greater than 1.16 with CSI support plus Helm 3, and you install the chart from the Artifact Hub repository. The README's Get Started guide covers standing up a test cluster first. Runtime images are published under the fluidcloudnative organisation, one controller image per cache engine.

### Which cache engines does Fluid support?

The repository ships controllers for Alluxio, Jindo, JuiceFS, Cache, EFC, Vineyard, and Thin runtimes, each with its own sample directory. ThinRuntime is the general path for other clients, and version 1.0.8 added 3FS and Curvine storage support through it rather than as dedicated engines.

## Sources

- [fluid-cloudnative/fluid on GitHub](https://github.com/fluid-cloudnative/fluid)
- [License: Apache-2.0](https://github.com/fluid-cloudnative/fluid/blob/master/LICENSE)
- [Project website](https://fluid-cloudnative.github.io/)
- [README](https://github.com/fluid-cloudnative/fluid/blob/master/README.md)
- [Releases](https://github.com/fluid-cloudnative/fluid/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/fluid-cloudnative-fluid
