Library / SDK
apache/opendal avatar
apache/opendal

Apache OpenDAL: one Rust data access layer for S3, GCS, HDFS and more

Apache OpenDAL: One Layer, All Storage. Object Storage File Storage s3 gcs azblob fs hdfs hdfs-native oss obs cos webhdfs lakefs ipfs tos b2 swift ipmfs azfile azdls upyun vercel-blob alluxio goosefs dbfs gridfs <a href.

5,396 stars828 forksRustApache-2.0

At a glance

What is it?
OpenDAL puts object storage, file systems, databases and key-value services behind a single Operator abstraction, with retry, timeout and metrics added as layers. The API is uniform; the semantics underneath are not, and that is the part worth checking before adoption.
Who is it for?
Adopt OpenDAL when one binary or service has to speak to several storage backends and you want retry, timeout and metrics configured once instead of once per SDK. Do not adopt it if your application is committed to a single backend and already uses that vendor's SDK well, or if you need transactional semantics across keys.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 6 days ago.
What is it written in?
Mainly Rust, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 26, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem OpenDAL solves for multi-backend code

Every storage SDK has its own vocabulary. S3 talks about buckets and objects, HDFS about paths and name nodes, Redis about keys, and a local filesystem about directories. An application that reads from more than one of them ends up with an adapter per backend, and the adapters drift: one retries, another does not; one sets a timeout, another hangs. OpenDAL's answer is a single trait surface called Operator. The README states the vision as One Layer, All Storage, and the service list spans object storage (s3, gcs, azblob, oss, obs, cos, tos, b2, swift, upyun, vercel-blob), file storage (fs, hdfs, hdfs-native, webhdfs, lakefs, ipfs, ipmfs, azfile, azdls, alluxio, goosefs, dbfs, gridfs, opfs, monoiofs), and further entries beyond the truncated table.

The audience is narrower than the tagline suggests. This is for library authors and platform teams who need the same read and write path against several backends, and for people building data systems where the backend is a deployment decision rather than a compile-time one. It is not aimed at someone who will only ever talk to one bucket.

How the Operator, services and layers fit together

Three pieces carry the design. The core is the Rust crate `opendal`, and the main abstraction is `Operator`. A service is a backend implementation; a layer wraps an operator to add behaviour that is not the backend's job. The README lists retry, timeout, logging, tracing, metrics, throttling and concurrency control as common layers, with concrete names such as RetryLayer, TimeoutLayer, LoggingLayer, TracingLayer, MetricsLayer, PrometheusLayer, OtelMetricsLayer, ThrottleLayer, ConcurrentLimitLayer, MimeGuessLayer, RouteLayer and FoyerLayer.

That split is the real mechanism. Retry policy is not reimplemented inside each of the dozens of services; it is a decorator you compose onto whichever operator you built. The same applies to bounding a slow call with TimeoutLayer or exporting counters with MetricsLayer. RouteLayer is the interesting one for multi-backend deployments, since it routes operations by path, which is how you keep one Operator in application code while the bytes land in different places.

Language bindings are the third piece, and the README is explicit that each binding has its own independent version number, which may differ from the Rust core version. That note matters operationally: when you check for updates or compatibility, the binding's version is the one to read, not the core crate's.

Installing OpenDAL and a first S3 read

The README points to a package per binding rather than a single install path, so the first step is picking the language. For Rust, the core package is the crate `opendal`; for Python, Go, Java, Node.js and Ruby the README links the corresponding package registries. The repository also ships a `.env.example` at the top level, which is the quickest way to see which keys a given service expects.

The S3 section of that file is a good template. It uses `OPENDAL_S3_BUCKET`, `OPENDAL_S3_ENDPOINT`, `OPENDAL_S3_REGION`, `OPENDAL_S3_ACCESS_KEY_ID` and `OPENDAL_S3_SECRET_ACCESS_KEY`.

bash
OPENDAL_S3_BUCKET=<bucket>
OPENDAL_S3_ENDPOINT=<endpoint>
OPENDAL_S3_REGION=<region>
OPENDAL_S3_ACCESS_KEY_ID=<access_key_id>
OPENDAL_S3_SECRET_ACCESS_KEY=<secret_access_key>

Those names are the ones the repository uses; the values are placeholders. The same file shows how local testing is meant to work: `OPENDAL_FS_ROOT=/path/to/dir` for the filesystem service, and `OPENDAL_FS_ATOMIC_WRITE_DIR=/path/to/tempdir` for the temporary directory used when atomic writes are required. If you want to try the API without any cloud credentials, the fs service is the cheapest entry point.

A second pattern worth copying from the file is the append flag. HDFS and hdfs-native both expose `OPENDAL_HDFS_ENABLE_APPEND=true` and `OPENDAL_HDFS_NATIVE_ENABLE_APPEND=true`, which tells you that append support is an explicit per-service setting rather than something the layer assumes. The README does not document rollback of a failed write in this article's material, so treat that as something to establish from the service documentation before relying on it.

Where a unified layer leaks backend semantics

One API does not mean one behaviour. The service list itself shows why: object storage, a POSIX filesystem, HDFS, GridFS, Redis and IPFS do not agree on what a directory is, whether a rename is atomic, or whether a write is visible before it completes. OpenDAL can normalize the call signature; it cannot normalize the guarantees underneath, and the README does not claim otherwise.

The append flags make this concrete. `OPENDAL_HDFS_ENABLE_APPEND` and `OPENDAL_HDFS_NATIVE_ENABLE_APPEND` exist because append is a property of the backend, not of the abstraction. Expect the same shape elsewhere: capabilities vary by service, and the documentation states that applications should enable only the backends and capabilities they use.

The second limitation is scope. OpenDAL is a data access layer, not a transaction manager. Nothing in the README describes cross-key atomicity, so a workload that needs multi-object transactions, or a filesystem that must behave exactly like ext4 for an existing application, is the wrong fit. The third is version skew: because each binding carries its own version number, a bug fix or a capability in the Rust core does not automatically appear in the Python or Java package you depend on.

OpenDAL compared with rclone

The comparison people search for is with rclone, and the difference is where the code runs. rclone is a command-line program and a sync engine: you point it at a source and a destination and it moves or mirrors data, with its own config file and its own scheduling. OpenDAL is a library. There is no OpenDAL binary to run a sync with; you embed the Operator in your own service and decide what the read and write path does.

That changes what each tool is good at. If the job is copying a bucket to another bucket on a cron, rclone is the direct answer and OpenDAL would mean writing the program first. If the job is an application that serves reads and writes against a backend chosen at deploy time, and you want retry, timeouts and metrics to be configuration rather than code, rclone is not in that space at all. The layers are the clearest expression of the split: RetryLayer, TimeoutLayer and MetricsLayer are things you compose into your own process, not flags on someone else's CLI.

Maintenance, releases and licence

The repository is not archived, and the last push was on 2026-08-21, the same date as the v0.58.2 release. The two releases before it were v0.58.1 on 2026-07-31 and v0.58.0 on 2026-07-16, so the patch cadence over that window is roughly one release per two to four weeks. That is a real signal about the core crate; it says nothing about the cadence of an individual binding, which the README warns carries its own version number.

Upgrade cost is mostly a function of which binding you use. If you depend on the Rust core, the release notes and CHANGELOG.md in the repository are the place to read before bumping. If you depend on a binding, check that binding's version rather than the core version, because matching them is not the compatibility rule the project defines.

The project is licensed under Apache-2.0, and the repository carries both a LICENSE and a NOTICE file, which is the usual Apache Software Foundation arrangement. Apache-2.0 includes an express patent grant and requires that notices be preserved. That is a description of the licence text, not legal advice; if you redistribute OpenDAL inside a product, have counsel read the NOTICE and the attribution requirements rather than treating this paragraph as clearance.

Editorial conclusion

Adopt OpenDAL when one binary or service has to speak to several storage backends and you want retry, timeout and metrics configured once instead of once per SDK. Do not adopt it if your application is committed to a single backend and already uses that vendor's SDK well, or if you need transactional semantics across keys. Before writing code, confirm that the service you need is listed in the services documentation, that your binding exposes the operations you plan to call, and that the binding's own version is the one you are pinning.

Frequently asked questions

What is Apache OpenDAL?

It is an Open Data Access Layer that gives every language a unified way to access object storage, file storage, cloud SaaS, databases, protocols and key-value services. The core is the Rust crate `opendal`, and the main abstraction is `Operator`.

What is OpenDAL?

The README describes it as an Open Data Access Layer guided by the vision One Layer, All Storage, with retry, timeout, logging, tracing, metrics and throttling offered as reusable layers. Services include s3, gcs, azblob, fs, hdfs, hdfs-native, oss, obs, cos, webhdfs, lakefs, ipfs and more.

How does OpenDAL compare with rclone?

OpenDAL is a library you embed, with no command-line sync program of its own, while rclone is a tool you run to move or mirror data between remotes. The difference shows in the layers: retry, timeout and metrics are composed into your process rather than passed as flags to a CLI.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
  4. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/apache-opendal.svg)](https://hysenlabs.com/projects/apache-opendal)