smart_open defaults to no dependencies and one in practice
Utils for streaming large files (S3, HDFS, gzip, bz2...)
At a glance
- What is it?
- smart_open presents itself as a drop-in replacement for Python's built-in open across S3, GCS, Azure, HDFS, HTTP and SFTP. Its README documents credentials inside URIs, two transports that shell out to a local command, and a compression parameter with six possible values.
- Who is it for?
- smart_open fits a Python codebase that streams large objects through cloud storage and wants one call signature instead of per-client boilerplate, and that is willing to choose extras per backend.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 6 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 5, 2026, and from our analysis. They are not legal advice.
Editorial analysis
A drop-in replacement with a fallback
The pitch is compatibility. The library claims it can do anything the built-in open function can, at full compatibility, and that wherever possible it falls back to the native implementation, which means local paths keep working exactly as before and only remote schemes take a different path. The storages named are S3, Google Cloud Storage, Azure Blob Storage, HDFS, WebHDFS, HTTP, HTTPS, SFTP and the local filesystem, with compression and decompression applied on the fly as the bytes pass. The reason the project exists is stated in the next section and is specific rather than aspirational: the object client's own upload and download methods need file-like wrappers constructed around them, and that boilerplate is where the bugs live.
Credentials appear inside URIs and in the environment
The list of accepted URI shapes is the densest part of the file, and two of the entries put credentials in the string itself:
s3://bucket/key
s3://access_key_id:secret_access_key@bucket/key
gcs://bucket/blob
azure://bucket/blob
hdfs://host:port/path/file
./local/path/file.gz
file:///home/user/file.bz2
[ssh|scp|sftp]://username:password@host/path/fileA later example on the same page does the opposite: it builds a storage client from environment variables, with a comment explaining that the client is thread-safe and a connection configuration that raises the pool size and enables keep-alive. So the documentation shows both conventions, and the choice is yours. What is worth naming out loud is that the first form puts a secret in whatever holds the URI, which means logs, tracebacks and shell history if you pass one on a command line.
Two transports shell out to a local command
Most backends here are a Python client library and a matching extra. Two are not. The HDFS schemes reach the Hadoop command line client instead, and that client has to be installed on the machine separately and be reachable on the path, so this is a transport whose prerequisites live outside Python entirely. The same paragraph lists two more cases of the same kind: Kerberos authentication for the HTTP transport needs a separate package installed, and GSSAPI or Kerberos authentication over SSH needs the SSH library installed with its GSSAPI support. None of those three is an extra you can request from this project, which is the practical difference between them and the seven backends that are extras.
No dependencies by default, and one in the manifest
The installation section explains the default: nothing optional is installed, so that the package stays small, and you ask for what you need by name. The documented form requests a set of extras in one line, and there is a catch-all extra for everything:
pip install 'smart_open[s3,gcs,azure,http,webhdfs,ssh,zst,lz4]'The manifest is more revealing than the prose. It declares a runtime dependency list with exactly one entry, a wrapping library, and then seven optional groups, one per backend, plus a development group. That group is unusually honest in its comments: the stubs for the S3 client are listed there for type checking only, with a note that the real client ships no inline types and that two typing helpers are used under a type-checking guard. So the library itself is thin, and everything heavy is opt-in and commented.
Compression is one parameter with six values
A single top-level parameter controls compression, and it takes six documented values. The default infers from the file extension, which is why the library can be pointed at a file whose name ends in one of the recognised extensions and have it handled without being told. The other values disable compression entirely, or name one of the five algorithms explicitly, the useful case being a file whose extension does not say what it is. Per-call options are forwarded to the underlying library untouched, and the documentation is specific that you must spell each option using that library's own keyword name, which is not one name: a compression level for gzip and bzip, a preset for xz, a level for zstandard, and a differently spelled level again for lz4. Registering a new format is a documented extension point.
The help text is a committed file with a generator
The API reference is unusual in being a plain text file in the repository root rather than a hosted site, reachable either in a browser or from an interpreter through the built-in help function. There is a script at the root whose name says what it does: it regenerates that file. Which means the reference is a build artifact that is committed, and a pull request that changes the API without regenerating it produces a stale document. Three more prose files sit beside it for the things READMEs usually skip: a guide to extending the library, a how-to, and a document specifically about migrating from older versions.
The documentation links a branch that is not the default
The repository's default branch is develop, and the version scheme has already moved past the branching convention some tools still assume. The file's own links do not follow: the licence badge, the reference file, the compression module and the manifest are all linked by paths containing the older branch name, so a reader following any of them lands on a page for a branch that is no longer where development happens. The coverage badge has the same problem in a different form, pointing at a repository under the organisation the project used to live in rather than the one it lives in now. None of this breaks the package, and all of it makes the documentation point at history.
The examples read a fixture named after a year
The compression examples read a small gzipped fixture from the test data directory and print its first thirty-two characters, once decompressed and once raw, where the raw output is a short run of escaped bytes starting with the gzip magic number. The fixture is named after a year, and the decompressed line the documentation prints is a sentence about a bright cold day in April with clocks striking thirteen. A separate example disables compression on the same file, and another passes an explicit algorithm for a file whose name does not end in the right extension. The version itself is not in the manifest: it is declared dynamic and filled in from the repository's tags at build time, which is why the release history can show three releases inside a single season with the manifest never hard-coding them.
Editorial conclusion
smart_open fits a Python codebase that streams large objects through cloud storage and wants one call signature instead of per-client boilerplate, and that is willing to choose extras per backend. Check four things before adopting it: which backends you actually use, since each is a separate extra and two of them need software outside Python entirely, whether you are on the branch you think, because the documentation links assume a branch the repository no longer defaults to, how you pass credentials, since the file shows them both in the URI and in the environment, and whether compression inference from the file extension matches your naming conventions. The single runtime dependency is small enough that you can read it before you install anything.
Frequently asked questions
What is smart_open in Python?
A library for streaming very large files to and from remote storage such as S3, Google Cloud Storage, Azure Blob Storage, HDFS, WebHDFS, HTTP, HTTPS and SFTP, with transparent compression. It presents itself as a drop-in replacement for the built-in open function, falling back to the native one where possible.
How do I install smart_open with cloud support?
By asking for the extras you need, for example an install requesting the s3, gcs, azure, http, webhdfs, ssh, zst and lz4 extras, or the all extra for everything. By default the package installs no optional dependencies, to keep the installation small.
Does smart_open support HDFS?
Yes, through two URI schemes that shell out to the Hadoop command line client, which must be installed separately and available on the path. It is therefore unusable on a machine that does not have that client.
How does smart_open decide whether to decompress a file?
By default it infers from the file extension. The top-level compression parameter can also disable compression or name one of the algorithms explicitly, with documented values for inference, disabling, and the bz2, gz, lz4, xz and zst formats.
What are the dependencies of smart_open?
One at runtime, a wrapping library. Everything else is an optional extra per backend: the S3 client, the Google cloud storage client, the Azure blob client, requests for HTTP, paramiko for SSH and lz4 for that compression format.
How does smart_open compare with fsspec?
The README makes no comparison. What it documents is a drop-in replacement model for the built-in open, a per-backend extras model with nothing installed by default, and two HDFS transports that depend on an external command line client.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/piskvorky-smart-open)