mosaicml/streaming: a PyTorch IterableDataset that reads training shards from object storage
A Data Streaming Library for Efficient Neural Network Training
At a glance
- What is it?
- The MosaicML Streaming library writes datasets into shards, uploads them to S3 or another object store, and serves them to distributed training jobs as a drop-in PyTorch IterableDataset. Here is what it actually does, what it costs you, and when plain local files are the better answer.
- Who is it for?
- Adopt mosaicml/streaming if your training data already lives in S3, GCS, Azure, OCI or a Databricks-backed store and your jobs are multi-node, because the shard index is what makes a shuffled stream reproducible across ranks. Do not adopt it if your dataset fits on the node's local disk or you need per-sample random access, since StreamingDataset is an IterableDataset and the README itself frames it as a replacement for that class.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 98 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What mosaicml/streaming solves, and who it is written for
Training on data that lives in cloud object storage usually means one of two bad options: copy the whole dataset to local disk before the job starts, or write custom code that opens objects during the training loop. The first wastes time and disk on every node, and the second tends to produce something slow and subtly non-deterministic. StreamingDataset is the MosaicML answer to both. The README describes the goal directly: make training on large datasets from cloud storage "as fast, cheap, and scalable as possible," aimed at multi-node distributed training for large models. It supports images, text, video and multimodal data, and the supported backends listed in the README are AWS, OCI, GCS, Azure, Databricks, and any S3-compatible store such as Cloudflare R2, Coreweave or Backblaze B2. The audience is therefore narrow and specific: teams running PyTorch training across more than one node, whose dataset is too large to stage locally, and who already pay for object storage. If you train a single-GPU model on a dataset that fits on an NVMe drive, this library adds a shard conversion step and a cache layer for no benefit.
Shards, a shard index, and a local cache: the mechanism
The data flow has three stages. First, MDSWriter reads your raw samples and writes them out as shards, which the README describes as the Mosaic Data Shard format, capable of encoding and decoding any Python object; CSV/TSV and JSONL are also listed as supported streaming formats. Second, you upload that directory to object storage, for example with the AWS CLI. Third, StreamingDataset is constructed with a remote path (where the shards live) and a local path (a working directory used as cache during operation), plus a shuffle flag. Samples are then addressed by integer index, as in the README's dataset[1337] example, even though the underlying class is an IterableDataset. That last detail is the interesting design choice. A plain IterableDataset gives you a stream and no index, which is why shuffling and resumption across ranks are hard to get right. Streaming keeps a shard-level index so every rank can agree on which samples belong to which step without each worker listing the bucket independently. The local directory is not optional bookkeeping: it is where shard data is cached as the job runs, so the node's disk becomes a buffer rather than a full copy. The README does not document cache eviction policy or how the local directory behaves when it fills, which is the main thing to measure before a long run.
Installing mosaicml-streaming and streaming your first shard
The README gives a single install command for the PyPI package. The distribution name differs from the import name, which trips people up: you install mosaicml-streaming and import from streaming.
pip install mosaicml-streamingConverting raw data means choosing a column schema and a compression codec, then writing samples through MDSWriter. The README's example maps an image field to jpeg and a class field to int, and uses zstd compression. The columns dictionary is the contract: every sample written must match it.
import numpy as np
from PIL import Image
from streaming import MDSWriter
data_dir = 'path-to-dataset'
columns = {
'image': 'jpeg',
'class': 'int'
}
compression = 'zstd'
with MDSWriter(out=data_dir, columns=columns, compression=compression) as out:
for i in range(10000):
sample = {
'image': Image.fromarray(np.random.randint(0, 256, (32, 32, 3), np.uint8)),
'class': np.random.randint(10),
}
out.write(sample)After the writer context manager exits, data_dir holds the shard files. Upload them with whatever CLI you already use; the README shows the AWS CLI form.
aws s3 cp --recursive path-to-dataset s3://my-bucket/path-to-datasetOn the training side you point remote at the bucket prefix and local at a scratch directory on the node, then wrap the dataset in an ordinary DataLoader. Indexing returns a dict with the same keys you declared in columns.
from torch.utils.data import DataLoader
from streaming import StreamingDataset
remote = 's3://my-bucket/path-to-dataset'
local = '/tmp/path-to-dataset'
dataset = StreamingDataset(local=local, remote=remote, shuffle=True)
sample = dataset[1337]
img = sample['image']
cls = sample['class']
dataloader = DataLoader(dataset)If the bucket credentials are wrong or the prefix is empty, the failure surfaces when StreamingDataset is constructed, not when you index, because the shard index has to be fetched first. The README points to the quick start guide and to end-to-end tutorials for CIFAR-10, FaceSynthetics and SyntheticNLP for the full walkthrough.
The shard index is a startup cost you cannot avoid
The design that makes cross-rank shuffling deterministic also creates the library's sharpest edge. Every job start has to obtain the shard index before training can begin, and every worker needs a consistent view of it. That is fine for a dataset with a few thousand shards. It is less fine when the shard count grows into the hundreds of thousands, because the metadata fetch becomes a fixed latency on every restart, including the restarts caused by a preempted node. The README does not document how the index is partitioned or how long it takes to load at scale, so this is something to measure on your own shard count rather than assume. The second limitation is the local cache. Because local is a working directory rather than a full copy, the node needs enough disk to hold whatever the epoch touches, and the repository does not document eviction behavior. A node with a small scratch volume will fail in a way that looks like a data problem but is a disk problem. Third, the format choice locks you in. MDS shards are read by this library; if you later want to hand the same directory to a different loader, you are converting back out of a format the README describes as able to encode any Python object, which is convenient to write and awkward to leave.
How it differs from WebDataset and from a plain IterableDataset
WebDataset is the closest comparison and it takes a different bet. It stores samples as tar archives named by shard, and the loader relies on the tar member naming convention to group related files into one sample; ordering and sharding are handled with a shuffle buffer and a split-by-worker rule that you configure yourself. Streaming instead keeps an explicit shard index that the library reads and uses to assign samples to ranks, which is where its determinism claim comes from. The trade-off is that WebDataset's tar shards are readable by anything that reads tar, while MDS shards require this library. If you need a format other tools can consume without a conversion step, WebDataset is the more portable choice. If your problem is that eight ranks on four nodes disagree about which samples they have seen, Streaming's index is the part doing the work. A hand-written IterableDataset over an fsspec or boto3 client sits at the other end: maximum control, no shard conversion, and you own shuffling, resumption and rank assignment yourself. That is a reasonable place to start and a bad place to stay once the training cluster grows.
Maintenance, releases and what Apache-2.0 means here
The repository is not archived, and the last push was on 2026-06-25, so it is within six months of a current date and the codebase is receiving changes. Release cadence is visible from the tags: v0.11.0 on 2025-01-15, v0.12.0 on 2025-04-03, v0.13.0 on 2025-07-15. That is roughly quarterly, and the version numbers are still pre-1.0, which means minor releases can carry breaking changes and the upgrade cost is real. Budget for reading the release notes before bumping, not for a silent pip upgrade in a training image. The licence is Apache-2.0, which permits commercial and closed-source use and includes an explicit patent grant; it also requires that you preserve notices and state significant changes. That is a permissive posture, and it is the reason a library like this can be embedded in a proprietary training stack. It is not legal advice, and the usual caveat applies: if you redistribute the package or a modified version, read the LICENSE file in the repository rather than this paragraph. Dependencies are pinned with upper bounds in setup.py, including torch, numpy and the cloud SDKs, so an environment that resolves today can stop resolving after an unrelated upgrade elsewhere in the image.
What to check before you convert your dataset
Start with the shard count, because it determines the startup latency of every job and every restart. Then check the local cache volume on the smallest node type you plan to use, since local is a cache and not a mirror. Then confirm that your access pattern is sequential-ish: StreamingDataset is an IterableDataset, and the README presents it as a drop-in replacement for that class, so code that needs random access to arbitrary samples is fighting the design. Finally, decide the column schema before the first conversion. Changing a column type after you have uploaded terabytes of shards means rewriting them, and the columns dictionary you pass to MDSWriter is the schema. The repository ships examples for CIFAR-10, FaceSynthetics, SyntheticNLP and a Spark DataFrame to MDS notebook, so the fastest way to validate the schema decision is to run one of those against a small slice of your own data before committing the full conversion.
Editorial conclusion
Adopt mosaicml/streaming if your training data already lives in S3, GCS, Azure, OCI or a Databricks-backed store and your jobs are multi-node, because the shard index is what makes a shuffled stream reproducible across ranks. Do not adopt it if your dataset fits on the node's local disk or you need per-sample random access, since StreamingDataset is an IterableDataset and the README itself frames it as a replacement for that class. Before committing, verify three things against your own environment: that the install resolves against your torch and Python versions, that the shard index is small enough to fetch at job start, and that the local cache directory you pass as local has enough space for the samples your epoch touches.
Frequently asked questions
How do I use mosaicml/streaming with an existing PyTorch training loop?
Construct a StreamingDataset with a remote bucket path and a local cache directory, then pass it to a standard torch.utils.data.DataLoader. The README shows this exact pattern, including indexing the dataset directly to inspect a single sample.
How do I install mosaicml/streaming?
The README gives one command, pip install mosaicml-streaming. Note that the distribution name is mosaicml-streaming while the import name is streaming.
What data formats can mosaicml/streaming write?
The README lists MDS (Mosaic Data Shard), which can encode and decode any Python object, plus CSV/TSV and JSONL. MDSWriter takes a columns dictionary mapping field names to types and an optional compression codec such as zstd.
Which cloud storage providers does mosaicml/streaming support?
The README lists AWS, OCI, GCS, Azure and Databricks, plus any S3-compatible object store such as Cloudflare R2, Coreweave or Backblaze B2. Uploading is left to your own CLI, and the README shows an aws s3 cp example.
Is mosaicml/streaming still maintained?
The repository is not archived and the last push was on 2026-06-25. Releases are pre-1.0 and roughly quarterly, with v0.13.0 on 2025-07-15 following v0.12.0 on 2025-04-03.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/mosaicml-streaming)