Model or dataset
mosaicml/llm-foundry avatar
mosaicml/llm-foundry

LLM Foundry: training and finetuning MPT and DBRX models with Composer

LLM training code for Databricks foundation models

4,449 stars589 forksPythonApache-2.0

At a glance

What is it?
LLM Foundry is Databricks Mosaic's Python codebase for training, finetuning, evaluating and deploying large language models. It is worth adopting when your training loop already fits Composer and StreamingDataset, and worth skipping when it does not.
Who is it for?
Adopt LLM Foundry if you are training or finetuning MPT, DBRX or HuggingFace-format models through Composer and are willing to accept the MosaicML platform and MCLI as the supported launch path. Do not adopt it if you want a framework-neutral trainer or a single-GPU recipe that avoids distributed configuration.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Activity is slowing. The repository last received commits 6 months ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 26, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What LLM Foundry is for, and who it is aimed at

LLM Foundry is the training stack Databricks Mosaic used for its own foundation models. The README states the repository contains code for training, finetuning, evaluating and deploying LLMs with Composer and the MosaicML platform. That sentence is the whole scope. It is not a general-purpose trainer that happens to support MPT; it is the codebase that produced MPT and, later, DBRX.

The audience follows from that. You are a team with multi-GPU or multi-node hardware, a dataset large enough that you care about streaming rather than loading into memory, and a reason to prefer YAML configuration over writing a training loop. The scripts directory is organised by stage: data_prep converts text into StreamingDataset format, train handles models from 125M to 70B parameters, inference converts to HuggingFace or ONNX, and eval runs in-context-learning tasks. Someone finetuning a 7B model on one node is inside the target audience. Someone who wants to swap in a custom optimizer written against raw PyTorch is fighting the design.

One thing the README does not do is separate the two products living in the same repository. MPT is a family of GPT-style models with Flash Attention, ALiBi for context length extrapolation and stability changes to reduce loss spikes. DBRX is a separate mixture-of-experts model with 132B total and 36B active parameters, trained with optimised versions of Composer, LLM Foundry and MegaBlocks. If you only care about DBRX, most of the MPT material is noise, and vice versa.

How the training stack is put together

The mechanism is a layer of configuration and callbacks on top of Composer. Composer supplies the training loop, the distributed launcher and the callback system; LLM Foundry supplies the model builders, the dataset wrappers and the YAML configs that wire them together. The pyproject.toml targets Python 3.9 and the setup.py classifiers list Python 3.10, so the supported interpreter range is narrow by modern standards.

Data flow starts in scripts/data_prep, which converts text from original sources into StreamingDataset format. That choice matters: StreamingDataset is designed so the dataset is streamed from object storage rather than materialised locally, which is why the repository can claim training runs at the 70B scale without a separate data pipeline. The cost is that your data has to be converted first, and the conversion step is a real stage in the workflow, not an optional optimisation.

Training is driven by YAML configs under scripts/train. The same directory holds a benchmarking subdirectory for profiling training throughput and MFU. Inference has a parallel structure: conversion to HuggingFace or ONNX, generation, and a benchmarking subdirectory for latency and throughput. The mcli directory launches any of these workloads through MCLI and the MosaicML platform. That last directory is the clearest statement of intent in the repository layout. The supported path for serious runs goes through MosaicML's tooling, and the scripts are written with that assumption.

Installing llm-foundry and generating from an MPT model

The package is published on PyPI as llm-foundry, and the README badge points at that project page. A plain pip install gets you the library and the scripts. The repository also ships a Dockerfile that clones the branch, installs it with pip, and then optionally uninstalls the package and deletes the clone directory when KEEP_FOUNDRY is not true, which is a caching trick for building images that only need the dependencies.

bash
pip install llm-foundry

After installation the README points at scripts/inference/README.md for interactive generation, and names two scripts: hf_generate.py for prompting HuggingFace-format models and hf_chat.py for chat-style prompting. The README describes this as the way to try the MPT models locally. Running the generation script against a downloaded checkpoint is the shortest path to confirming that the install works end to end.

bash
cd scripts/inference
python hf_generate.py --help

The --help invocation is worth doing before anything else, because the flags for model path, prompt and generation settings are documented in the script's own argument parser rather than in the top-level README. If the script imports cleanly and prints its options, the Composer and Transformers dependencies resolved correctly. For a deeper walkthrough, the repository ships TUTORIAL.md, which the README describes as a deeper dive into the repo, example workflows and FAQs. That file is the right second stop, ahead of reading the source.

Where LLM Foundry gets in the way

The tightest constraint is the Composer dependency. Every training, finetuning and evaluation path in the repository runs through it, and the Dockerfile installs a specific branch of the repository rather than a released wheel. If your organisation has standardised on a different training framework, adopting LLM Foundry means adopting Composer as well, and the migration cost is not confined to the training script.

The second constraint is the platform assumption. The mcli directory exists to launch workloads on the MosaicML platform. Nothing in the README claims the scripts only run there, and the Makefile shows a local path using composer.cli.launcher with a configurable WORLD_SIZE and MASTER_PORT. But the documented, maintained route for multi-node work is the platform, and a team running on its own Slurm cluster is largely on its own for scheduling.

The third is the model coverage. The README advertises training and finetuning of HuggingFace and MPT models from 125M to 70B parameters, with DBRX trained using optimised versions of the stack. That is a specific set. If your architecture is not in that set, the model builders are the part you will be extending, and the YAML config format assumes the builders exist.

Finally, the repository is not archived and the last push was on 2026-03-25, roughly six months before this writing. The most recent release listed is v0.22.0 from 2025-07-29. That is a slower cadence than the release history earlier in the list, and it is worth checking whether the specific component you depend on has moved since.

LLM Foundry against a general-purpose finetuning library

The obvious comparison is a finetuning library such as HuggingFace's Trainer or a wrapper like Axolotl. The difference is not features, it is where the abstraction sits. A general-purpose finetuning library treats the training loop as fixed and exposes the model, dataset and hyperparameters. LLM Foundry treats the loop as configurable and exposes the callbacks, the dataset streaming layer and the model builder.

That inversion shows up in the directory layout. Data preparation is a separate script stage rather than a dataset class you pass in. Throughput and MFU profiling live in a benchmarking subdirectory next to the training scripts rather than in a separate tool. Evaluation is a scripts/eval directory running in-context-learning tasks, not a metrics callback you register. If you want to change how data is fed or how a run is profiled, LLM Foundry gives you a place to do it. If you want to finetune a model on a CSV and stop, the extra structure is overhead.

The second difference is the model families. LLM Foundry is the codebase behind MPT and DBRX, and the MPT models carry specific design choices: Flash Attention, ALiBi for context length extrapolation, and stability changes aimed at loss spikes. A general-purpose library supports those models too, but it does not assume them. Choosing LLM Foundry means those choices are the default rather than something you configure.

Licence, releases and the cost of staying current

The repository is Apache-2.0, and the setup.py carries the SPDX identifier Apache-2.0 in its header. That covers the code. It does not cover the weights. The README is explicit that DBRX weights and code are licensed under the Databricks Open Source License with a separate Acceptable Use Policy, and the MPT table has a commercial use column where several chat variants are marked No. If you plan to ship a product on top of an MPT chat model, check that column before you build anything. This is a factual distinction in the documentation, not legal advice, and the licence text is the authority.

Upgrade cost is driven by the release cadence and the Composer coupling. The listed releases run v0.20.0 in April 2025, v0.21.0 in May 2025 and v0.22.0 in July 2025, with no further release listed before the last push on 2026-03-25. Because the Dockerfile installs from a branch rather than a pinned release, a team building images from main absorbs whatever landed since the last tag. Pinning to a release tag and pinning Composer alongside it is the cheaper arrangement, at the price of missing fixes.

There is a test surface worth knowing about. The Makefile defines test, test-gpu, test-dist and test-dist-gpu targets, with test-dist running pytest under composer.cli.launcher at a configurable WORLD_SIZE and MASTER_PORT. If you fork the repository, those targets are the fastest way to see whether your changes break the distributed path, which is the path most likely to break silently.

Editorial conclusion

Adopt LLM Foundry if you are training or finetuning MPT, DBRX or HuggingFace-format models through Composer and are willing to accept the MosaicML platform and MCLI as the supported launch path. Do not adopt it if you want a framework-neutral trainer or a single-GPU recipe that avoids distributed configuration. Before committing, verify that the pinned Composer version in your environment matches the one this repository expects, check whether the model you intend to train has a YAML config under scripts/train, and confirm the licence of the weights you download separately from the Apache-2.0 code.

Frequently asked questions

What is LLM Foundry?

It is the Databricks Mosaic repository containing code for training, finetuning, evaluating and deploying LLMs with Composer and the MosaicML platform. It is the codebase behind the MPT model family and, together with MegaBlocks, DBRX.

What is LLM Foundry used for?

The README lists four uses: training and finetuning HuggingFace and MPT models from 125M to 70B parameters, converting text data into StreamingDataset format, converting models to HuggingFace or ONNX and generating responses, and evaluating models on in-context-learning tasks.

What does LLM stand for?

The repository does not expand the acronym. Its own naming is Mosaic Pretrained Transformers for the MPT family, and it describes DBRX as an open source LLM using a mixture-of-experts architecture.

What is the primary role of LLM Foundry within the enterprise?

The README frames it as the codebase for training, finetuning, evaluating and deploying LLMs with Composer and the MosaicML platform, and the mcli directory launches those workloads on that platform. The repository does not describe a separate enterprise product role.

Official sources

  1. License: Apache-2.0
  2. mosaicml/llm-foundry on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/mosaicml-llm-foundry.svg)](https://hysenlabs.com/projects/mosaicml-llm-foundry)