CubeFS: a CNCF-graduated distributed file and object store for Kubernetes data
cloud-native distributed storage
At a glance
- What is it?
- CubeFS is a CNCF-graduated distributed storage system written in Go that exposes POSIX, HDFS, S3 and a REST API over the same cluster. This review covers what it is for, how the metadata and data nodes fit together, how to install it, and where it stops being the right tool.
- Who is it for?
- Adopt CubeFS when you need one cluster to serve POSIX mounts, an S3 endpoint and HDFS-style access to the same data, and you are willing to run master, metanode, datanode and objectnode roles yourself or via the Helm chart. Do not adopt it for a single-node home lab or a workload that only needs an S3 bucket, since a plain object store is simpler.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 2 days ago.
- What is it written in?
- Mainly Go, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The gap CubeFS fills: one storage cluster, several access protocols
Most storage decisions force a split. A Kubernetes cluster wants a POSIX filesystem for legacy applications, a data lake wants an S3 endpoint, and an analytics engine may want HDFS semantics. Running three separate systems means three sets of credentials, three capacity pools and three failure domains. CubeFS is positioned against exactly that split. The README describes it as an open-source cloud-native distributed file and object storage system, and the feature list names POSIX, HDFS, S3 and its own REST API as access protocols over the same deployment.
The intended audience is infrastructure teams running container platforms, not individual developers. The README lists datacenter filesystem, data lake storage infrastructure, and private or hybrid cloud storage as the target uses, and calls out separation of storage and compute for databases, search systems and AI/ML applications. Hybrid cloud is a named scenario: the project says it can run in public cloud services and provide cache acceleration and file system semantics on top of public cloud storage such as S3. That is a specific claim about a tiering role, not a general-purpose object store pitch.
The project is hosted by the Cloud Native Computing Foundation as a graduated project, and it carries a reference to a SIGMOD 2019 paper titled CFS: A Distributed File System for Large Scale Container Platforms. The paper is the strongest signal of what the design was built for. If your workload does not look like a large container platform, the fit is weaker.
How the pieces fit: master, metanode, datanode, objectnode and the client
The repository layout is the clearest description of the architecture. Top-level directories include master/, metanode/, datanode/, objectnode/, blobstore/, lcnode/, authnode/, raftstore/, remotecache/, sdk/, cli/, shell/, console/ and client/. Each maps to a role in the cluster, and the architecture diagram in the README shows the same decomposition.
The master coordinates the cluster. The metanode holds the metadata service, which the README describes as highly scalable with strong consistency. The datanode stores the actual data and is where the storage policy decision lands: the README lists high-performance replication and low-cost erasure coding as the two flexible storage policies. The objectnode provides the S3-compatible and REST interface, which is why the same cluster can answer both a FUSE mount and an S3 request. The lcnode and remotecache directories point at the multi-level caching and lifecycle work behind the hybrid cloud acceleration claim. Persistent state for the metadata path relies on raftstore, and the go.mod file pins gorocksdb, which is the RocksDB binding used underneath.
The client side is not one binary. The Makefile builds separate targets: server, client, cli, libsdk, fsck, fdstore, bcache, blobstore and deploy. There is also a java/ directory and an sdk/ directory, so non-Go applications are expected to talk to the cluster through a library rather than the filesystem alone. The go.mod module path is github.com/cubefs/cubefs and the declared Go version is 1.18.
Installing CubeFS: build targets, Helm chart and a first mount
The README does not carry installation steps. It points to the documentation site at cubefs.io, and the repository ships INSTALL.md and HELM.md at the top level, plus deploy/ and docker/ directories. The README badge for Artifact Hub links to the Helm chart repository cubefs, so a Kubernetes install is the documented path. For a source build, the Makefile defines the targets and build/build.sh does the work.
To build the server binaries from a checkout, run the default target. It fans out to the server, client, cli, libsdk, fsck and other targets listed in the Makefile.
make buildOutput lands in build/bin, which is also what the clean target removes. If you only need the command line tool, the Makefile exposes it separately.
make cliThe README states that the master branch may be in an unstable or even broken state during development, and asks users to take releases instead of the master branch to get a stable set of binaries. Treat that as an instruction, not a formality: check out the v3.6.0 tag before building if you want the release the maintainers shipped on 2026-08-12.
For Kubernetes, HELM.md and the Artifact Hub chart are the entry points. The chart name published on Artifact Hub is cubefs. The repository does not spell out the chart install command in the README, so read HELM.md and the chart's own values file rather than guessing flags. The same applies to the console and dashboard components: console/ exists in the tree, but the README does not document a UI login flow, so confirm what the release notes for v3.6.0 cover before planning around it.
Where CubeFS is the wrong choice
The strongest limitation is stated by the project itself: the master branch may be unstable or broken, and releases are the supported artifact. That is normal for a large Go codebase, but it means you cannot track master in production and expect a stable binary set. Plan around tagged releases and read CHANGELOG.md and UPDATE.md before moving between them.
The operational surface is large. A working deployment involves master, metanode, datanode and objectnode roles, with raftstore and RocksDB underneath the metadata path. The go.mod replace directives, including a fork of gorocksdb, show that the build depends on pinned third-party code rather than upstream releases. That is a maintenance cost you inherit. If your team has never operated a distributed filesystem, the learning curve is real, and the README offers no quick-start that shortens it.
CubeFS is also the wrong tool when the requirement is narrow. If you only need an S3-compatible bucket for application backups, a single-purpose object store is less to run. If you only need a shared POSIX volume inside one Kubernetes cluster, a CSI driver backed by a simpler system is enough. CubeFS earns its complexity when you genuinely need several protocols over one dataset, or when the hybrid cloud caching path matters to you.
The README does not document rollback procedure for a failed upgrade. UPDATE.md exists, but the README itself is silent on reverting a release. Verify that before you upgrade a production cluster.
CubeFS compared with Ceph and SeaweedFS
Ceph is the obvious comparison and the difference is architectural. Ceph builds everything on RADOS and layers RADOSGW, CephFS and RBD on top of that single object layer. CubeFS instead separates a metadata service (metanode) from the data path (datanode) and gives the object interface its own component (objectnode). The practical consequence is that CubeFS can tune metadata and data independently, which is what the README means by a highly scalable metadata service with strong consistency. Ceph's reputation for metadata-heavy workload pain is the trade-off CubeFS is designed around.
SeaweedFS takes the opposite route. It is built around a simple filer over many small volume servers, which makes it easy to start and pleasant for object workloads. CubeFS is a heavier, more opinionated system with a formal governance model, a TSC, maintainer and committer roles, and a CNCF graduation. If you want the smallest thing that stores files, SeaweedFS is the lighter answer. If you want a project with published governance and a SIGMOD paper behind its design, CubeFS is the more institutional choice.
Neither comparison is settled by a feature list. The README's own framing is that CubeFS targets datacenter filesystem, data lake and hybrid cloud storage. Judge the alternatives against that scope, not against a generic object store checklist.
Maintenance, releases and the Apache-2.0 licence
The last push to the repository was on 2026-09-22, and the most recent release is v3.6.0 from 2026-08-12. Before that, v3.5.3 landed on 2025-12-23 and v3.5.2 on 2025-07-31. The cadence is roughly two releases a year, with patch releases between them. The project is not archived.
Upgrade cost is where the release cadence matters. The repository carries UPDATE.md, CHANGELOG.md and RELEASE.md, which is more documentation than many projects provide for the upgrade path. The README does not describe rollback, so the practical approach is to read UPDATE.md for the target version and treat the upgrade as a one-way operation until the docs say otherwise. The bi-weekly community meeting and the meeting minutes wiki are the places to ask about upgrade experience.
CubeFS is licensed under the Apache License, Version 2.0, with details in LICENSE and NOTICE. Apache-2.0 is a permissive licence with an explicit patent grant and no copyleft obligation on your own code. The NOTICE file matters if you redistribute the software: Apache-2.0 requires you to carry attribution notices forward. This is a description of the licence terms, not legal advice; have your own counsel review redistribution plans.
Editorial conclusion
Adopt CubeFS when you need one cluster to serve POSIX mounts, an S3 endpoint and HDFS-style access to the same data, and you are willing to run master, metanode, datanode and objectnode roles yourself or via the Helm chart. Do not adopt it for a single-node home lab or a workload that only needs an S3 bucket, since a plain object store is simpler. Before committing, verify the artifacthub Helm chart version against the v3.6.0 release, and check whether the console and dashboard components you need are covered by the documentation for that release.
Frequently asked questions
What is CubeFS used for?
The README lists datacenter filesystem, data lake storage infrastructure, and private or hybrid cloud storage as its uses, and says it enables separation of storage and compute for databases, search systems and AI/ML applications. It exposes POSIX, HDFS, S3 and its own REST API over the same cluster.
How do I install CubeFS on Kubernetes?
The repository ships HELM.md and a Helm chart published on Artifact Hub under the name cubefs, and the README links to that chart. The README itself does not give install commands, so read HELM.md and the chart values for the release you are deploying.
Is the CubeFS master branch safe to build from?
No. The README states that the master branch may be in an unstable or even broken state during development and asks users to take releases instead of master to get a stable set of binaries. Check out a tagged release such as v3.6.0 before building.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/cubefs-cubefs)