Composer by MosaicML: A PyTorch Training Library for Multi-Node Scale and Algorithmic Speedups
Supercharge Your Model Training
At a glance
- What is it?
- Composer is an open-source deep learning training library by MosaicML that wraps PyTorch with distributed training abstractions, elastic sharded checkpointing, and a callback system for inserting training-time speedup algorithms.
- Who is it for?
- Composer is worth adopting for teams that train large models across multiple GPUs or nodes using PyTorch and want a structured way to apply speedup algorithms without rewriting their training loop. It is not suited for inference workloads, and teams whose bottleneck is data preprocessing or pipeline orchestration rather than the training loop itself will see little benefit.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 155 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Composer Addresses and Who Needs It
Composer targets ML engineers who train large deep learning models on GPU clusters and who want to reduce training cost without fundamentally changing their model architecture. The library wraps PyTorch's training loop with a `Trainer` abstraction that handles distributed data parallelism, sharded model parallelism, data loading, checkpointing, and logging, while exposing extension points for speedup algorithms.
The README lists the model types the library is designed for: large language models, diffusion models, embedding models such as BERT, transformer-based models generally, and convolutional neural networks. Teams training on a single GPU can also use Composer, though the distributed training features are the primary value proposition.
The Trainer and Its Distributed Training Modes
The heart of Composer is the `Trainer` class, which provides a highly optimized PyTorch training loop. It exposes configuration parameters for parallelization scheme, data loaders, metrics, loggers, and callbacks in a single initialization call.
For large models that exceed GPU memory, Composer integrates PyTorch's FullyShardedDataParallelism (FSDP), which shards model parameters, gradients, and optimizer state across GPUs. The README notes that FSDP is competitive in performance with more complex parallelism strategies. Standard distributed data parallelism (DDP) is also supported for models that fit on a single GPU but benefit from data parallelism.
A practical feature is elastic sharded checkpointing: a model saved across eight GPUs can be resumed on sixteen. This removes the constraint that checkpoints are tied to a specific hardware configuration. Combined with auto-resumption (the trainer automatically resumes from the latest checkpoint when a run is restarted), this significantly reduces the operational cost of long training jobs on preemptible cloud instances.
The Trainer also integrates directly with MosaicML StreamingDataset for large-scale datasets that are stored in cloud blob storage. Data is downloaded on the fly during training rather than requiring full local copies, which matters when the dataset is measured in terabytes. This integration is documented separately in the StreamingDataset repository but the Trainer supports it natively.
Installing Composer and Running a Training Job
Composer installs from PyPI:
pip install mosaicmlFor GPU support, ensure your PyTorch installation matches your CUDA version before installing. The setup.py reads the Composer version from `composer/_version.py` at build time.
A basic distributed training run using the launcher:
python -m composer.cli.launcher -n 2 --master_port 26000 -m pytestThe Makefile shows the test patterns for reference:
WORLD_SIZE=1 LOCAL_WORLD_SIZE=1 python3 -m pytestFor distributed GPU tests:
WORLD_SIZE=1 LOCAL_WORLD_SIZE=1 python3 -m pytest -m gpuThe examples directory includes notebooks for common starting points: `checkpoint_autoresume.ipynb`, `finetune_huggingface.ipynb`, `pretrain_finetune_huggingface.ipynb`, and `migrate_from_ptl.ipynb` for teams coming from PyTorch Lightning. The `exporting_for_inference.ipynb` notebook documents how to export a Composer-trained model for inference outside the Trainer, which is a necessary step for teams who train with Composer but serve with a different stack.
Speedup Algorithms and Published Results
Composer includes a collection of training speedup algorithms that can be stacked into recipes. The README publishes the results from MosaicML's own training runs using these algorithms:
For Stable Diffusion, combining speedups reduced training cost from $200,000 to $50,000, an 8x reduction. For ResNet-50 on ImageNet, training time dropped from 3 hours 33 minutes to 25 minutes on 8 A100 GPUs, a 7x speedup. BERT-Base pretraining went from 10 hours to 1.13 hours on 8 A100s, an 8.8x speedup. DeepLab v3 on ADE20K went from 3 hours 30 minutes to 39 minutes on 8 A100s, a 5.4x speedup.
These results come from MosaicML's own experiments and are linked to blog posts in the README. They should be treated as directional figures for the types of models listed, not as guarantees for other architectures or datasets.
The Callback System for Custom Training Logic
Composer's training loop runs through a series of named events: `BATCH_START`, `BEFORE_LOSS`, `AFTER_LOSS`, `AFTER_BACKWARD`, `BATCH_END`, and others at the epoch and training level. Callbacks are Python classes that implement methods named after these events.
The README gives the Learning Rate Monitor Callback as an example: it logs the learning rate at every `BATCH_END` event. MosaicML has written callbacks for memory usage monitoring, image logging and visualization, and estimated time-to-completion. Researchers who want to implement novel training techniques, such as curriculum learning schedules or custom gradient clipping rules, can insert them as callbacks without modifying the core Trainer.
This design contrasts with frameworks that require subclassing a Trainer or overriding forward/backward methods. The event system keeps custom logic isolated and composable.
Limitations and Cases Where Composer Is the Wrong Tool
Composer is built around the Trainer abstraction, which means it works best when your training loop maps cleanly onto the epoch/batch/callback structure. Reinforcement learning training loops, online learning systems, or custom multi-objective training protocols that do not fit the standard epoch structure require significant adaptation.
The library does not handle data pipeline construction, preprocessing, or streaming outside the Trainer. For large datasets, it integrates with MosaicML's StreamingDataset, but that is a separate library. Teams that need a comprehensive end-to-end ML platform, including experiment tracking, model registry, and serving, will need to integrate Composer with other tools.
PyTorch Lightning is a direct alternative. It also wraps PyTorch training with distributed training support and a callback system, and has a larger community and more integrations. The `migrate_from_ptl.ipynb` example in the repository exists specifically to help teams coming from Lightning evaluate whether Composer's speedup algorithms justify migration.
The library also carries a notable version constraint in the build system: `setuptools < 78.0.2` is pinned as a build requirement in pyproject.toml. Teams using very recent Python build toolchains should verify their setuptools version is compatible before building from source. The pre-commit configuration and ruff lint rules are present, suggesting the codebase follows a defined style review process, but they also mean that contributing changes requires matching the project's toolchain.
Maintenance Status and License
The repository is Apache-2.0 licensed and not archived. The last push was on 2026-04-29. The most recent GitHub release is v0.32.1, published on 2025-07-26. The prior releases were v0.32.0 (2025-07-15) and v0.31.0 (2025-05-28). The MosaicML team was acquired by Databricks, and the README links to Databricks career pages under a Mosaic AI department label.
The library uses setuptools with `setuptools < 78.0.2` pinned as a build requirement. The pre-commit configuration and ruff lint rules are included, suggesting the codebase follows a defined style and code review process. The `docker/` directory contains Dockerfiles for setting up a development or training environment, which are useful for teams working on GPU clusters where reproducible environments matter.
Editorial conclusion
Composer is worth adopting for teams that train large models across multiple GPUs or nodes using PyTorch and want a structured way to apply speedup algorithms without rewriting their training loop. It is not suited for inference workloads, and teams whose bottleneck is data preprocessing or pipeline orchestration rather than the training loop itself will see little benefit. The last push to the repository was on 2026-04-29, and the most recent release was v0.32.1 from 2025-07-26; verify the latest release is compatible with your PyTorch version before starting a new project on top of it.
Frequently asked questions
How do I install MosaicML Composer?
Run `pip install mosaicml` to install the core library from PyPI. Ensure your PyTorch installation matches your CUDA version first. The repository also provides setup.py and pyproject.toml for editable installs from source.
Does Composer support training on multiple GPUs and nodes?
Yes. Composer supports both standard PyTorch distributed data parallelism (DDP) and FullyShardedDataParallelism (FSDP) for models that are too large to fit on a single GPU. Elastic sharded checkpointing allows saving on one number of GPUs and resuming on a different number.
What is the difference between MosaicML Composer and PyTorch Lightning?
Both wrap the PyTorch training loop with distributed training support and a callback system. Composer's focus is on training efficiency algorithms and the ability to stack speedup recipes; the repository includes a migration notebook from PyTorch Lightning. Lightning has a larger ecosystem of integrations and plugins. The README includes a `migrate_from_ptl.ipynb` notebook for teams evaluating whether to switch.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/mosaicml-composer)