gcsfuse: a user-space file system over Google Cloud Storage buckets
A user-space file system for interacting with Google Cloud Storage
At a glance
- What is it?
- Google's FUSE adapter turns object storage into something a training job can read like a directory, and v3 rewrites the write path to stop paying for the difference.
- Who is it for?
- gcsfuse is at its best on the workload the README keeps coming back to: a machine learning job that reads large checkpoints and model weights many times, where the alternative is either a copy into local disk or a rewrite of the loader. The file cache with parallel downloads is the feature that changes model load times rather than shaving them, and streaming writes stop a large checkpoint from occupying local disk twice.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 8 days ago.
- What is it written in?
- Mainly Go, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Mounting a bucket with one docker invocation
Cloud Storage FUSE is an open source FUSE adapter that lets you mount and access Cloud Storage buckets as local file systems. It runs in user space and it is Generally Available, supported by Google starting with v1.0, provided it stays inside the documented supported applications, platforms and limits. The project is Apache licensed and the last push was 2026-09-28.
The Dockerfile in the repository is the shortest complete example of how the pieces fit. It builds the binary from source, drops it into an alpine image with fuse installed, and sets the entrypoint so the container starts a foreground mount:
# Mount the gcsfuse to /mnt/gcs:
# > docker run --privileged --device /fuse -v /mnt/gcs:/gcs:rw,rshared gcsfuseENTRYPOINT ["gcsfuse", "-o", "allow_other", "--foreground", "--implicit-dirs", "/gcs"]That last line is worth reading closely, because three of its four options encode the design decisions the project keeps running into. `allow_other` is needed when more than one user inside the container has to see the mount. `--foreground` keeps the process in the foreground so a container runtime can supervise it rather than losing the process to a background fork. `--implicit-dirs` is the one that changes how object names appear: Cloud Storage has no real directories, so gcsfuse synthesises them from the slash prefixes in object names, and `implicit-dirs` makes those synthesised directories show up in listings.
The build image does not copy a mount helper by accident. It installs both binaries, the daemon at /usr/local/bin/gcsfuse and the sbin helper at /usr/sbin/mount.gcsfuse, so that `mount.gcsfuse` style invocations work inside the container the same way they do on a host.
What streaming writes changed in v3
Streaming writes is the new default write path, and it uploads data directly to Cloud Storage as it is written. The previous path staged the entire write in a local file and only uploaded to GCS on close or fsync. That sounds like a small implementation detail until you are writing a multi-gigabyte checkpoint: the old behaviour needed local disk equal to the size of the file being written, and the upload latency was paid at the end rather than spread across the write.
The README claims checkpoint writes can be up to 40 percent faster under the streaming path, observed in training runs. The direction of that number is the interesting part. Sequential large writes are the case that benefits, because there is a natural stream to send and no need to seek backwards and patch a staged file.
The second v3 change is that file cache parallel downloads are now the default. Parallel downloads use multiple workers to download a file in parallel, using the file cache directory as a prefetch buffer. The README recommends it for single-threaded read scenarios that load large files, naming model serving and checkpoint restores, with up to 9x faster model load times claimed. The mechanism is straightforward once you see the problem: a framework that reads a checkpoint through a single thread cannot saturate a network link on its own, so gcsfuse does the reading in parallel and hands over bytes in order.
The third v3 item is quieter and possibly the most useful. GCSFuse automatically optimises its configuration when running on specific high-performance Google Cloud machine types, and manually set values at mount time override those defaults. That is a deliberate decision to have the adapter read the machine it landed on.
The file cache and what it actually stores
The file cache arrived in v2 and is the feature the rest of the design leans on. It allows repeat file reads to be served from a local, faster cache storage of choice, such as a Local SSD, Persistent Disk, or even in-memory tmpfs. The README cites up to 2.3x faster training time and 3.4x higher throughput observed in training runs, and points at multi-epoch training as the case where it pays off most, since a second epoch reads the same bytes again.
Two details are easy to miss. First, the cache is disabled by default and is enabled by passing a directory to `cache-dir`. Second, the directory you pass determines the whole character of the system. Pointing it at a Local SSD gives you fast reads and real capacity. Pointing it at tmpfs gives you read speed that never touches a disk but puts your working set in RAM, which is a bad trade for a training run that streams a large dataset through.
The other small optimization in the release history points the same way. v3.8.4, published 2026-09-02, carries a rapid bucket startup optimization that skips redundant DirectPath connectivity checks. That is a startup path fix rather than a data path fix, and it matters most for short-lived jobs where mount time is a real fraction of the run.
Because the cache is a local directory with a real file layout, cache size becomes a tuning decision with a real failure mode: a full cache directory. The README links out to the caching overview on cloud.google.com for the eviction behaviour rather than describing it here.
Where bucket semantics refuse to look like a filesystem
Cloud Storage has objects, not files. It has no rename, no partial in-place write, no hard link, and no directory inode. gcsfuse closes that gap with conventions, and conventions are where surprises live. The project keeps the full discussion in docs/semantics.md, splitting the explanation between behaviour with and without streaming writes, and the README only links to it.
The practical consequences show up in ordinary workloads. A program that writes a file, seeks back, and patches bytes in the middle is not writing the way gcsfuse uploads. A program that renames a temporary file into place is relying on an atomic operation that object storage does not have, so the naming convention has to stand in for it. Directory listings come from prefix delimitation, which means a listing can be slow and is not a free operation the way a local readdir is.
There is a second class of difference that has nothing to do with operations and everything to do with throughput. Every metadata call is an API call. A build system that stats thousands of files will feel that. The same is true of small appends to a single object, which is why the streaming write path matters more than the headline percentage suggests.
The README handles this the way it handles everything else operationally: a sentence saying the differences from POSIX file systems are documented, and a link. Anyone adopting gcsfuse for a new pipeline should budget time for that page, because it is where the project explains what it will not do.
A Go codebase built around an open source FUSE layer
The dependency list in go.mod says more about the project than the feature list does. The module path is github.com/googlecloudplatform/gcsfuse/v3 and it targets Go 1.27.0. The FUSE layer itself comes from jacobsa/fuse, plus a family of small jacobsa packages for daemonizing, matching, mocking, sync primitives and time handling, which is the standard shape for a project built on that library.
Cloud access goes through cloud.google.com/go/storage v1.65.1 with cloud.google.com/go/auth for credentials and cloud.google.com/go/iam alongside it. Configuration and flags come from spf13/cobra with spf13/viper on top, which is the pair most Go CLIs pick and also the reason flag and config file behaviour here will feel familiar to anyone else who has used that stack.
Observability is not an afterthought in the dependency list. prometheus/client_golang and the OpenTelemetry exporters for metrics and traces are both present, alongside go.opencensus.io and the OpenTelemetry log bridge, which suggests a transition between two generations of instrumentation rather than a single choice. The repository tree backs this up: there are top-level metrics/, perfmetrics/ and tracing/ directories, plus a benchmarks/ directory that suggests performance is tracked as part of the project rather than left to anecdote.
Test dependencies sit alongside the real ones in the same file. github.com/fsouza/fake-gcs-server and github.com/stretchr/testify both appear in the require block, which is the practical signal that the test suite runs against an in-process GCS emulator instead of a real bucket. There is also a flaky_tests.lst file at the root, a refreshing admission that some tests are known to be unreliable and are tracked by name.
The module and version lines read like this:
module github.com/googlecloudplatform/gcsfuse/v3
go 1.27.0The /v3 suffix on the module path matters if you import it, since Go treats major versions after v1 that way. The configuration samples and Kubernetes manifests live under samples/gcsfuse_config/ and samples/gke-csi-yaml/, and there is a DEBIAN/ directory for packaging.
Releases, the CSI driver, and where to start reading
Three of the recent releases are worth reading as a group. v3.12.0 on 2026-09-17 has an empty body, which is what a routine release looks like in a repository that keeps its change notes elsewhere. v3.11.4 on 2026-09-10 carries a single dependency fix, moving google.golang.org/grpc from v1.82.1 to v1.83.1 for CVE-2026-84304, an uncontrolled resource consumption and remote denial of service through HTTP/2 receive buffer memory exhaustion. v3.8.4 on 2026-09-02 is the Rapid bucket startup fix. The version numbers are not in step with the dates, so the numbering is tracking branches rather than a single line, and a patch release can land out of order.
For Kubernetes, the README points at the Cloud Storage FUSE CSI driver in a separate repository. The division of labour is clean: gcsfuse does the mounting, and the CSI driver lets GKE handle deployment and management of the Cloud Storage FUSE instance on nodes, which the README describes as a turn-key experience. If your storage problem is a pod that needs a bucket, you are probably installing the driver, not gcsfuse directly.
There is also a docs/ directory with the semantics page and a troubleshooting page linked from the README, and a cfg/ directory that suggests configuration is a first-class part of the design rather than an afterthought bolted onto flags.
The honest summary is that this README is an index, not a manual. It states the status, describes what v3 changed, and links to the technical overview, the caching overview, the semantics documentation, the pricing page and the list of supported operating systems and validated machine learning frameworks. That last link is the one most evaluators skip and most need: the supported application list is what tells you whether your pipeline qualifies.
Editorial conclusion
gcsfuse is at its best on the workload the README keeps coming back to: a machine learning job that reads large checkpoints and model weights many times, where the alternative is either a copy into local disk or a rewrite of the loader. The file cache with parallel downloads is the feature that changes model load times rather than shaving them, and streaming writes stop a large checkpoint from occupying local disk twice. What it does not do is make buckets POSIX, and docs/semantics.md is where the honest limits live. The last push was 2026-09-28 and v3.12.0 shipped 2026-09-17, so the line is moving quickly. Start with one bucket, one mount point, and a cache directory on local SSD, then read docs/semantics.md before trusting anything a POSIX application does with a rename.
Frequently asked questions
How do you use gcsfuse to mount a bucket?
You mount the bucket with the gcsfuse binary or the mount.gcsfuse helper, pointing it at a mount point and passing options such as `allow_other`, `--foreground` and `--implicit-dirs`. The repository's Dockerfile shows the container form, where the entrypoint is `gcsfuse` with those options and /gcs as the mount point. The full flag list and per-platform install steps are documented on cloud.google.com rather than in the README.
Does gcsfuse behave like a real POSIX filesystem?
No. Cloud Storage has objects rather than files, so gcsfuse emulates directories from slash prefixes in object names and adopts conventions for operations object storage does not have. Renames, in-place partial writes and metadata-heavy workloads all behave differently, and every metadata call becomes an API call. docs/semantics.md in the repository is where the project spells out these differences.
How do you speed up repeated reads of the same objects?
Enable the file cache by passing a directory to `cache-dir`, and choose where that directory lives based on what you are optimising for. A Local SSD gives fast reads with real capacity, and tmpfs keeps reads in memory. In v3 the cache directory also doubles as the prefetch buffer for parallel downloads, which the README recommends for single-threaded readers loading large files such as checkpoints.
What changed between gcsfuse v2 and v3?
v3 makes streaming writes the default write path, so data uploads as it is written instead of being staged in a local file until close, and turns on parallel downloads from the file cache by default. It also adds automatic configuration tuning on specific high-performance machine types, with manual mount-time values taking precedence. v2 introduced the file cache itself, disabled by default.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/googlecloudplatform-gcsfuse)