3FS: DeepSeek's Distributed File System for AI Workloads
A high-performance distributed file system designed to address the challenges of AI training and inference workloads.
At a glance
- What is it?
- Fire-Flyer File System (3FS) is a C++ distributed file system built by DeepSeek to address the storage demands of large-scale AI training and inference, combining RDMA networks with CRAQ consensus and a FoundationDB-backed metadata layer to deliver a shared file interface that scales across hundreds of storage nodes.
- Who is it for?
- Teams running large-scale AI training clusters on InfiniBand-equipped hardware will find 3FS a technically serious option for shared storage, particularly for workloads that require random-access data loading without prefetching, high-throughput parallel checkpointing, or KV cache offloading for inference.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 145 days ago.
- What is it written in?
- Mainly C++, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
DEEP OPEN-SOURCE ANALYSIS
The Storage Problem in Large-Scale AI Training
Training large neural networks across many compute nodes creates storage patterns that general-purpose distributed file systems were not designed for. Data loading for training typically requires random access to individual training samples across a large dataset, not sequential reads. Checkpointing requires writing large model state snapshots at high throughput, often while training continues. KV cache for inference requires a storage tier that is cheaper than DRAM but faster than typical object storage, with high IOPS for small reads.
Network file systems designed for enterprise workloads, such as NFS or Lustre, face a different access pattern. Lustre is optimized for large sequential reads common in scientific computing. NFS has metadata bottlenecks at scale. Object storage systems like S3 can provide capacity but impose latency overhead for fine-grained random access.
3FS was built by DeepSeek to address these patterns specifically. The README describes four workload categories it targets: data preparation, dataloader access during training, checkpointing, and KV cache for inference. Its disaggregated architecture spreads storage across many nodes while presenting a standard POSIX file interface to application code, so existing training frameworks that read files from disk work without modification.
CRAQ Consensus and the Metadata Architecture
3FS implements data consistency using Chain Replication with Apportioned Queries (CRAQ). CRAQ is a protocol that places replicas in a chain: writes go to the head node and propagate down the chain, while reads can be served from any node in the chain. The README describes this as providing strong consistency, which simplifies application code because clients do not need to handle stale reads or conflict resolution.
The metadata layer uses stateless metadata services backed by FoundationDB, a transactional key-value store. FoundationDB provides ACID transactions at the metadata level, so file system operations like directory creation, rename, and attribute updates are atomic. The README notes that stateless metadata servers are an explicit design choice: they hold no durable state themselves, which makes horizontal scaling and crash recovery straightforward.
The storage layer combines the throughput of the SSDs across all storage nodes and the bandwidth of the RDMA network connecting them. The README describes the architecture as allowing applications to access storage in a locality-oblivious manner: a compute node can read any block from any storage node without caring which physical node holds it. The network fabric, not a specific storage node, determines effective bandwidth.
Building 3FS from Source
3FS has no pre-built binary releases in the GitHub repository. Building from source requires clang-14, cmake, and a set of system libraries. The README provides separate package lists for Ubuntu 20.04 and Ubuntu 22.04:
# for Ubuntu 20.04.
apt install cmake libuv1-dev liblz4-dev liblzma-dev libdouble-conversion-dev libdwarf-dev libunwind-dev \
libaio-dev libgflags-dev libgoogle-glog-dev libgtest-dev libgmock-dev clang-format-14 clang-14 clang-tidy-14 lld-14 \
libgoogle-perftools-dev google-perftools libssl-dev libclang-rt-14-dev gcc-10 g++-10 libboost1.71-all-dev build-essentialBeyond system packages, three additional prerequisites must be installed manually: libfuse version 3.16.1 or newer (from the libfuse GitHub releases), FoundationDB version 7.1 or newer, and a Rust toolchain with a minimum version of 1.75.0 (recommended 1.85.0 or newer). The Cargo.toml in the repository sets the workspace MSRV at 1.85.0.
To check out the code including all submodules:
cd 3fs
git submodule update --init --recursive
./patches/apply.shThe build itself uses CMake:
cmake -S . -B build \
-DCMAKE_CXX_COMPILER=clang++-14 -DCMAKE_C_COMPILER=clang-14 \
-DCMAKE_BUILD_TYPE=RelWithDebInfo -DCMAKE_EXPORT_COMPILE_COMMANDS=ON \
-DSHUFFLE_METHOD=<method>
cmake --build build -j 32The SHUFFLE_METHOD flag (either g++10 or g++11) must be set explicitly. Binaries compiled with different SHUFFLE_METHOD values are incompatible, so all nodes in a cluster must use the same setting. Docker build images are available for TencentOS-4 and OpenCloudOS-9 if the local environment does not match these prerequisites.
The SHUFFLE_METHOD Compatibility Constraint
The README documents a binary compatibility issue related to how the C++ standard library implements std::shuffle across compiler versions. Binaries built with g++10 and binaries built with g++11 or later use different shuffle algorithms, which makes them incompatible at the cluster level.
To resolve this, 3FS requires all builds for a given cluster to use the same explicitly specified -DSHUFFLE_METHOD flag. For new clusters, either g++10 or g++11 can be chosen, but the choice is permanent: once the cluster is deployed with one setting, all future builds must use that same value. For existing clusters, the flag must match whatever compiler was originally used.
This is a concrete operational constraint. If a team adds new nodes to a cluster and builds from source without verifying the SHUFFLE_METHOD setting, the new binaries will be incompatible with the existing nodes. The README links to a GitHub issue (issue 368) documenting this problem. The Docker images for TencentOS-4 and OpenCloudOS-9 abstract away the compiler environment but do not remove this requirement.
Performance Benchmarks Documented in the Repository
The README documents three specific benchmark results, all measured on real 3FS deployments.
For peak read throughput, the README describes a read stress test on a cluster of 180 storage nodes, each with two 200 Gbps InfiniBand NICs and 16 NVMe SSDs of 14 TiB each. Over 500 client nodes participated, each with a 200 Gbps InfiniBand NIC. The measured aggregate read throughput reached approximately 6.6 TiB/s while background traffic from training jobs was present.
For the GraySort benchmark, a two-phase sort of 110.5 TiB of data across 8,192 partitions completed in 30 minutes and 14 seconds on a cluster of 25 storage nodes and 50 compute nodes. The average throughput was 3.66 TiB per minute.
For KV cache read throughput in inference workloads, the README reports peak read throughput of up to 40 GiB/s across all KV cache clients, each using a 400 Gbps InfiniBand NIC. These numbers come from the 3FS deployment at DeepSeek; the benchmark tooling for reproducing them is in the benchmarks/fio_usrbio directory using the project's fio engine for the USRBIO API.
Where 3FS Does Not Fit
3FS depends on RDMA networking (InfiniBand in the documented configurations) for its throughput characteristics. The README describes InfiniBand NICs at 200 Gbps and 400 Gbps per node. Ethernet-based RDMA (RoCE) is not discussed in the README, and standard TCP/IP networking will not deliver the throughput the benchmarks demonstrate. Teams without InfiniBand infrastructure cannot expect comparable performance.
The build process is complex. The dependency on FoundationDB 7.1 and libfuse 3.16.1 alongside specific compiler versions means that deploying 3FS requires deliberate cluster preparation. There is no package manager installation path and no pre-built container for production use in the repository.
The last push to the repository was on 2026-05-07. New issues and questions should be filed at github.com/deepseek-ai/3fs/issues. The README does not describe a release schedule or a versioning policy, and the setup.py shows the version as 1.2.9 plus the current git commit hash, indicating that semantic versioning is not used.
Comparing 3FS with HDFS
HDFS (Hadoop Distributed File System) is a widely deployed distributed file system designed for batch processing workloads. It uses block replication across commodity hardware, an HDFS NameNode for metadata, and TCP/IP networking. DataVec in the DL4J ecosystem and many Spark-based training pipelines are built around HDFS.
HDFS and 3FS target different network and storage assumptions. HDFS is designed for clusters using standard Ethernet with no RDMA requirement, which makes it accessible on general-purpose cloud VMs. 3FS requires RDMA hardware to achieve its documented throughput figures. HDFS processes reads sequentially in large blocks; 3FS is designed for random access to individual training samples without prefetching or shuffling, which the README highlights as a specific design goal for dataloader workloads.
For teams already running HDFS-based pipelines on standard cloud instances, 3FS is not a direct replacement. It is an architecture for organizations building or operating dedicated AI training clusters with InfiniBand interconnects, where the storage layer is a meaningful bottleneck.
Editorial conclusion
Teams running large-scale AI training clusters on InfiniBand-equipped hardware will find 3FS a technically serious option for shared storage, particularly for workloads that require random-access data loading without prefetching, high-throughput parallel checkpointing, or KV cache offloading for inference. The system is not a drop-in replacement for general-purpose distributed file systems in environments without RDMA hardware: it is designed specifically around InfiniBand bandwidth and NVMe SSD throughput. Before deploying, verify that your cluster meets the hardware requirements and that FoundationDB 7.1 or newer is available, since the metadata layer depends on it directly.
Frequently asked questions
What metadata store does 3FS use?
3FS uses FoundationDB, a transactional key-value store, as the backend for its stateless metadata services. FoundationDB provides ACID transactions for file system operations, and version 7.1 or newer is required.
Does 3FS require RDMA hardware?
The documented performance benchmarks and the system design are based on InfiniBand RDMA networking. The README describes clusters using 200 Gbps and 400 Gbps InfiniBand NICs. Standard TCP/IP networking is not discussed in the README as a supported configuration for the storage layer.
What is the SHUFFLE_METHOD build flag in 3FS?
SHUFFLE_METHOD controls which shuffle algorithm is compiled into the 3FS binaries. Binaries compiled with g++10 and g++11 are incompatible at the cluster level, so all nodes must use the same setting. Once a cluster is deployed with a given SHUFFLE_METHOD value, the same value must be used for all future builds to maintain compatibility.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/deepseek-ai-3fs)
Community notes