Marin: an open framework for training foundation models from data curation to evaluation
Open-source framework for the research and development of foundation models. Marin's primary use case is training language model like Llama, DeepSeek, Qwen, etc.
At a glance
- What is it?
- Marin is a Python research platform for training language models, covering tokenization, pretraining, posttraining and evaluation. Its documentation is the main entry point, and its example experiments run through a step graph rather than a single training script.
- Who is it for?
- Marin fits teams that want to reproduce an end-to-end language model pipeline in the open and are comfortable working from the docs/ folder, since the README points there for installation and the first experiment. It is a poor fit if you only need to fine-tune an existing checkpoint, because the framework's centre of gravity is data curation and pretraining.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The gap Marin fills: process knowledge, not just checkpoints
Most open model releases ship weights and a model card. Marin's stated concern is the rest of the pipeline: data curation, transformation, filtering, tokenization, pretraining, posttraining and evaluation. The README describes Marin as a research program, software platform and community, and says its core value is open development, with experiments and decisions documented as they happen, including failed experiments.
That framing sets the audience. It is aimed at people who need to build or reproduce a training run, not people who want a downloadable model to call through an API. The README notes that Marin has also been used for audio-text models, DNA and protein models, and encourages that work through Marin as a library in marin/experiments. So the intended user is a researcher or engineer who wants the machinery, and is willing to read the docs/ folder to get it.
Steps, dependencies and topological order: how a Marin experiment is assembled
The README's clearest architectural statement is that Marin experiments are defined as a set of steps that can depend on each other and are executed in a topological order, like a Makefile. That single sentence explains a lot about the repository layout: experiments/ holds definitions, config/ holds configuration, and the workspace is split into packages such as marin-iris, marin-fray, marin-levanter, marin-core, marin-zephyr and marin-rigging, all listed as dependencies in pyproject.toml.
The example in the README builds a tokenization step and a training step, where training depends on the tokenized dataset. The comment in the sample code is explicit that the tokenized dataset is a lazy handle and nothing downloads yet. That laziness is the mechanism that makes the dependency graph useful: you can declare a pipeline whose later stages are expensive, and the work is deferred until the graph is executed.
One consequence worth naming: because steps are declared before they run, debugging a failure means reasoning about which step produced the bad artifact, not just reading a stack trace from a training loop.
Installing Marin and training a tiny model on TinyStories
The README does not inline installation steps. It points to docs/tutorials/installation.md, and the Makefile shows how the repository expects a development environment to be initialised. The init target installs pandoc through conda, installs pandiff through npm, syncs the workspace with uv, and logs in to Hugging Face.
make initThat runs conda install -c conda-forge pandoc, npm install -g pandiff, uv sync --extra cpu and huggingface-cli login. If you only want the dependencies and not the documentation tooling, the uv sync --extra cpu line is the part that matters. The project requires Python 3.12 or newer according to pyproject.toml.
The README's first real use is a tiny model trained on TinyStories. The excerpt shows the beginning of the script, which imports ResourceConfig, AdamConfig, the step runner, the tokenized dataset helper and the training helper, then declares a tokenized dataset with a sample_count cap of 1000.
from fray.cluster import ResourceConfig
from levanter.optim import AdamConfig
from marin.execution.lazy import lower
from marin.execution.step_runner import StepRunner
from marin.experiment.data import tokenized
from marin.experiment.train import train_lm
from experiments.llama import llama_nano
from experiments.marin_tokenizer import marin_tokenizer
tinystories_tokenized = tokenized(
name="tokenized/tinystories",
source="roneneldan/TinyStories",
tokenizer=marin_tokenizer,
sample_count=1000, # cap
)After that you add a training step that depends on the tokenized dataset, and the README says the training step runs after tokenization completes. The full script lives at experiments/tutorials/train_tiny_model.py, and the README points to docs/tutorials/first-experiment.md for the guided version. Expect the sample_count cap to keep the first run small; the README does not state how long it takes.
Where Marin is the wrong tool
Marin is built around training models, and the README's current focus is a frontier mixture-of-experts model at 5e24 model-FLOPs with 500 billion-plus total parameters, plus the Delphi scaling suite that spans 3e18 to 1e23 FLOPs. That is a different problem from adapting an existing checkpoint to a narrow task.
The README does not document a fine-tuning-only path, a serving stack, or an inference API. If your goal is to take a released model and put it behind an endpoint, the framework's data curation, tokenization and pretraining stages are overhead you would be carrying without using.
There is a second boundary. The Delphi suite was trained on the Google TPU Research Cloud, and the Makefile includes a configure_gcp_registry_all target that applies a 30-day Artifact Registry cleanup policy across regions drawn from config/marin.yaml. The tooling clearly assumes a cloud cluster with credentials and a registry. Running the larger tutorials on a single workstation is not something the README describes.
Marin compared with the Pythia-style scaling suite approach
The README names Pythia as the inspiration for Delphi, and the comparison is instructive. A scaling suite in the Pythia tradition publishes a family of models trained at increasing compute budgets, and the models are the deliverable. Delphi is described as having three parts: a scaling recipe that maps compute budgets to model configurations, a scaling suite trained from that recipe, and a scaling law that uses the smaller models to predict the larger ones.
The difference is what gets released. Alongside checkpoints on Hugging Face, Marin also publishes the training mixture pipelines that deterministically reproduce the mix from Nemotron-CC, StarCoderData and ProofPile 2, a forkable CompletedAdamHParams recipe class, an agent skill named add_scaling_heuristic, and plot-ready data with one config per figure and a wandb_url on every row. The scaling law itself is one artifact among several.
That is a heavier release surface, and it only pays off if you intend to re-run or extend the recipe. If you just want the checkpoints, the extra machinery is inert.
Maintenance, licensing and the cost of staying current
The repository is not archived, and the last push was on 2026-08-13. The most recent release listed is dev-wheels, a set of dev native wheels published the same day; before that, marin-zephyr-latest and marin-rigging-latest were published on 2026-05-29. The gap between those dates is worth noting if you depend on the native components rather than the Python code.
Upgrade cost is shaped by the workspace design. pyproject.toml declares packages such as marin-finelog, marin-ducky and marin-marina as workspace members so that fixes land without a wheel republish, while the native companion marin-finelog-server ships as a pre-built wheel and arrives transitively through marin-iris. The Makefile exposes rust-dev and rust-user targets to switch between building the Rust component from source and installing the pre-built wheel, with rust-status to show the current mode. That switch is the main knob you will touch when a native dependency and the Python code drift apart.
Marin is licensed under Apache-2.0, with the licence declared in pyproject.toml and the LICENSE file at the repository root. Apache-2.0 is a permissive licence with an explicit patent grant. That is a statement about the licence text, not legal advice; if you plan to redistribute a derivative, read the LICENSE file yourself.
Editorial conclusion
Marin fits teams that want to reproduce an end-to-end language model pipeline in the open and are comfortable working from the docs/ folder, since the README points there for installation and the first experiment. It is a poor fit if you only need to fine-tune an existing checkpoint, because the framework's centre of gravity is data curation and pretraining. Before adopting it, read docs/tutorials/installation.md and docs/tutorials/first-experiment.md, and check whether your hardware matches the Google TPU Research Cloud setup the Delphi scaling suite was trained on. Note that the last push to the repository was on 2026-08-13.
Frequently asked questions
What is Marin and what is it used for?
Marin is a research program, software platform and community for the research and development of foundation models, with its primary concern being training large language models. The README states its scope covers data curation, transformation, filtering, tokenization, pretraining, posttraining and evaluation, and that it has also been used for audio-text, DNA and protein models.
How do I install Marin?
The README points to docs/tutorials/installation.md for installation rather than listing steps inline. The Makefile's init target runs conda install -c conda-forge pandoc, npm install -g pandiff, uv sync --extra cpu and huggingface-cli login, and pyproject.toml requires Python 3.12 or newer.
How do I train a tiny language model with Marin?
The README's example trains a tiny model on TinyStories by declaring a tokenized dataset with a sample_count cap of 1000, then a training step that depends on it. The full script is at experiments/tutorials/train_tiny_model.py and the guided tutorial is docs/tutorials/first-experiment.md.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/marin-community-marin)