# JuiceFS keeps your files in numbered blocks, so you cannot browse them in the object storage console

> JuiceFS presents object storage as a POSIX file system, with metadata held in a separate database engine and file data split into three levels of block before it reaches a bucket. It is Apache-2.0 licensed, actively developed, and at version 1.3.3 while the 1.4 series also exists. The two consequences worth reading before you deploy it are in the front page: your objects stop being files, and the performance claims have no numbers attached.

**juicedata/juicefs** — JuiceFS is a distributed POSIX file system built on top of Redis and S3.

- Repository: https://github.com/juicedata/juicefs
- Website: https://juicefs.com
- Stars: 14,497 · Forks: 1,302
- Language: Go
- License: Apache-2.0
- Published: 2026-08-21 · Updated: 2026-08-21 · Language: en
- Canonical page: https://hysenlabs.com/projects/juicedata-juicefs

## Your files become numbered blocks and the front page says not to panic about it

There is a paragraph in the architecture section that explains what your bucket will look like after you use this, and it is unusually blunt about the consequence.

It says files stored through the file system cannot be found in the object storage platform's own file browser. Instead there is a chunks directory and a collection of numerically named directories and files in the bucket. Then: don't panic, this is just the secret of the high-performance operation.

So the object you can point a console at is not a file. It is a block.

That is inherent to the design rather than a shortcoming, and the fix is that blocks are what you want in a bucket: they compress, they replicate, and they do not need rewriting when one byte of a large file changes. The cost is paid everywhere else.

So the consequence is that anything in your estate which reads the bucket directly will break, and not loudly. Lifecycle rules that expire by prefix, cross-region replication configured by prefix, a second application reading the same objects, and any tooling that lists what is in the bucket, will all see blocks instead of the files you put there.

## Data goes to three levels before it reaches storage, with two of the sizes fixed

The front page describes the storage layout in three named units, and two of them have stated default sizes.

A file is split into chunks at a fixed size with a default upper limit of sixty-four mebibytes. Each chunk is composed of one or more slices, and the slice length varies depending on how the file was written. Each slice is composed of fixed-size blocks, four mebibytes by default. The blocks are what land in object storage, while the metadata describing the file and its chunks, slices and blocks goes to the metadata engine.

So there are three layers and one of them is variable. A slice's length depends on the write pattern, which means the same file written twice can produce a different number of slices and therefore a different set of objects.

The variable middle layer is what makes append-heavy workloads efficient, because you do not rewrite what you already wrote. It is also why the object count in a bucket will not match what you expect.

Alongside the layout there are two optional transforms applied to the data. It can be compressed, with a choice of two algorithms, and it can be encrypted in transit and at rest, with the details deferred to a guide.

So the consequence is that the block size and the chunk ceiling are the two numbers that determine your bucket's object count and your rewrite behaviour, and they are configuration rather than constants.

## Metadata and data live in different systems, so the mount needs both

The architecture section names three components, and the interesting one is the split.

The client coordinates the object storage and the metadata engine and implements the file system interfaces. The data storage holds the blocks, across local disk, public or private cloud object storage, or a distributed file system. The metadata engine holds everything about the files themselves: names, sizes, permissions and groups, timestamps and directory structure.

Four metadata engines are named, one of which is an in-memory key-value store described as particularly suitable for metadata, and the others a relational database, an embedded one and a distributed transactional key-value store.

Now set that against the consistency claim in the features list, which says a confirmed modification is immediately visible on all servers mounting the same file system. That guarantee is about clients, and it is bought with the metadata engine being a single source of truth.

Which means there are two systems to run, two to back up, and two to reason about when something is inconsistent. The documentation has an advanced-topics section with a page on best practices for the in-memory store and a separate page on fault diagnosis, which tells you both of those are places people have had trouble.

So the consequence is that the metadata engine is the critical component, not the bucket, and the front page's own advanced-topics list is where to look before you pick one.

## The newest release is a patch on the older line while a newer line exists

Read the release dates against the version numbers and there are two lines in maintenance at once.

The three most recent releases are a patch on the one point three series dated August 2026, and two releases on the one point four series dated July 2026 and earlier in July.

So the newest release by date is numerically behind the two before it. A tool that resolves dependencies by version rather than by date will pick the one point four release, and a tool that picks by recency will pick the older line.

This is a deliberate and ordinary pattern: a maintenance release went to the one point three line after something needed fixing there, while the one point four line continued separately. It is the same shape of choice you saw in the sibling files in this batch, and it is worth the same attention.

The default branch took a push on 2026-09-29, so work is continuing on top of all of it.

So the consequence is that pinning by a floating version is ambiguous here, and if you need one line rather than the other you have to say so explicitly.

## The performance claims are adjectives with one figure attached to a benchmark link

One feature bullet makes three performance statements and supports none of them with a measurement in the text.

It says the latency can be as low as a few milliseconds. It says the throughput can be expanded nearly unlimitedly, with a parenthetical that qualifies it by saying this depends on the size of the object storage. And it links to a benchmark page rather than publishing numbers.

As low as and nearly unlimitedly are both unfalsifiable as written. The latency figure has no workload, no block size, no metadata engine and no cache configuration attached to it. The throughput figure is explicitly bounded by something outside the software, which is honest but means the number is really a property of your bucket.

What is stated precisely, by contrast, is the storage layout, and that is where the performance argument actually lives. Four-mebibyte blocks and a sixty-four-mebibyte chunk ceiling are concrete and they determine object count and rewrite behaviour.

There is a caching page in the advanced-topics list, which is the other half of any latency claim and the part you configure.

So the consequence is that the front page tells you the design that produces the performance and not the performance itself, which is a defensible choice for a filesystem whose speed depends almost entirely on the two systems underneath it.

## Compression, encryption and both flavours of file lock are on the same code path

Three of the ten feature bullets are about what happens to the bytes and the locks rather than about storage, and together they describe the options you have on the data path.

Two compression algorithms are named, one of which is a general-purpose fast codec and the other a modern one with a higher ratio. The choice is per deployment rather than per file, and it applies to all data, which means it is not free and it is not reversible without rewriting.

Encryption is supported in transit and at rest, and the front page defers entirely to a guide for how it works. That deferral is the right call for a scheme with key management, and it also means the key handling is not documented where you would look for it.

Then locking, and this one is specific: both BSD advisory locks and POSIX record locks are supported. That is the pair that matters for a shared mount, because a database or an editor may take one flavour and another tool the other, and a filesystem that silently supported only one would produce locking bugs that are extremely hard to trace.

Supporting both means implementing both, which is more work and more state per open file, and it means the semantics of the two are not identical. Record locks are per byte range; advisory locks are per file. Code written for one does not always behave correctly on the other.

So the consequence is that a mount looks ordinary from inside and behaves correctly for the common cases, and the cases where it does not are the ones where two tools disagree about which kind of lock they are taking.

## One repository holds four interfaces and a documentation toolchain with its own version number

Two structural details are worth naming.

The first is scope. The client implements the POSIX interface, a Hadoop interface, a Kubernetes interface and an S3-compatible gateway. Each of those has its own documentation page on the project site, and the features list calls out the Hadoop compatibility, the object-storage-compatible gateway and the Kubernetes driver as three separate headline items alongside the POSIX claim.

So this is four ways to reach the same data rather than one file system with one mount.

The second detail is in the repository itself. There is a Node package manifest at the root of a Go project, named the same as the project, and its declared version is one point zero while the language releases are in the one point three and one point four series. Its dependencies and its five scripts are all documentation tooling: two linters, a link checker, and two rules that enforce trailing-slash and proper-name conventions.

So the repository carries two version numbers that do not correspond to each other, and the Node manifest exists purely to lint documentation.

The lint setup also explains the prose. One of the scripts runs a Chinese-language proofreading tool over the documentation and both readme files, and there are two readmes and two adapter documents, one pair in each language.

So the consequence is that the English front page is machine-checked by a tool built for Chinese text, which is why some sentences read as translated rather than written.

## Conclusion

This is the right tool if you have data in object storage and want a POSIX mount over it, because the metadata lives in a real database and the data stays in the bucket you already have, which means you are not copying it anywhere. The S3 gateway and the Hadoop and Kubernetes integrations mean it also fits existing estates without rewiring them. Two things to check before you commit. That nothing in your estate needs to read the objects directly, because a mounted file is a set of numbered blocks and a lifecycle rule or a second tool pointed at that prefix will not see a file. And that the metadata engine and the bucket are both on the failure path, since a mount needs both and there is no degraded read mode described.

## FAQ

### what is juicefs

A high-performance POSIX file system released under the Apache 2.0 licence, designed for cloud environments. It persists file data in object storage and the corresponding metadata in a separate database engine, so cloud storage can be mounted as a local file system without changing application code. It has a single runtime dependency layer for expression binding, and the client also implements Hadoop, Kubernetes and object-storage-compatible interfaces.

### How do JuiceFS work?

Files are split into chunks at a fixed size with a default upper limit of sixty-four mebibytes, each chunk into one or more slices whose length varies with how the file was written, and each slice into fixed-size blocks of four mebibytes. The blocks go to object storage and the metadata about the file and its chunks, slices and blocks goes to the metadata engine. The consequence is that your files are not visible as files in the object storage console.

### Is JuiceFS compatible with Kubernetes?

Yes. A container storage interface driver is provided, described on the front page as a headline cloud-native feature for using the file system in Kubernetes, with its own documentation page. It can also be used as a persistent volume for two container runtimes, with a separate page for that.

### is juicefs open source

Yes, under the Apache 2.0 licence, which the front page states in its opening paragraph and links. The repository also carries a code of conduct, a contributing guide, a changelog and separate documentation for English and Chinese readers. Its development is coordinated through a chat server and a community introduction page.

### juicefs vs minio

The front page does not compare the two. It describes itself as a POSIX file system that stores data in object storage and metadata in a database engine, and separately notes that it provides an object-storage-compatible gateway so that existing object-storage clients can read through it. A comparison against a specific alternative would have to come from their documentation rather than this repository.

## Sources

- [Official documentation](https://juicefs.com)
- [Official README](https://github.com/juicedata/juicefs#readme)
- [Project repository](https://github.com/juicedata/juicefs)
- [Release notes](https://github.com/juicedata/juicefs/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/juicedata-juicefs
