Model or dataset
NVIDIA/cosmos-framework avatar
NVIDIA/cosmos-framework

NVIDIA Cosmos-Framework: training and serving the Cosmos3 world models

Our inference and training framework to run on the Cosmos Models

552 stars153 forksPythonNOASSERTION

At a glance

What is it?
Cosmos-Framework is NVIDIA's single-package toolchain for fine-tuning and serving the Cosmos3 omnimodal world models. It assumes Linux, an NVIDIA GPU and a willingness to work through 8-GPU recipes before anything runs.
Who is it for?
Adopt Cosmos-Framework if you already have Linux hosts with NVIDIA GPUs and want to fine-tune or serve the Cosmos3 family rather than build a training stack yourself. Do not adopt it if you are looking for a CPU-only or macOS path, or a stable API: the package is classified as Beta and the shipped recipes assume 8 GPUs.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 5 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Cosmos-Framework is for, and who it is not for

Cosmos-Framework is the training and serving half of the NVIDIA Cosmos project family. The README describes it as an end-to-end framework for training and serving world models, including the Cosmos3 model family, and everything lives in one top-level Python package, cosmos_framework. The companion repository at github.com/nvidia/cosmos is where NVIDIA points people who want a guided experience with Cosmos3; this repository is the machinery underneath.

The audience is narrow by design. The package metadata lists Linux as the operating system and NVIDIA CUDA as the environment, and the setup instructions install ffmpeg, git-lfs and X11 development headers before anything Python happens. If you want to generate a video from a prompt on a laptop, this is the wrong entry point. If you have a multi-GPU node and want to fine-tune a Cosmos3 checkpoint on your own video or action data, or stand up an inference service around one, the repository covers both ends: a distributed trainer and Diffusers, Transformers and vLLM backends for generation.

One package, two entry points: training and inference

The architecture is deliberately flat. There is no separate trainer repository and inference repository to keep in sync; cosmos_framework.scripts.train and cosmos_framework.scripts.inference are the two entry points, and the README names both.

On the training side the trainer is distributed across FSDP, tensor parallelism, context parallelism and pipeline parallelism. Checkpoints are native DCP with import and export to HuggingFace safetensors, which matters if you want to move a fine-tuned model into an inference stack that expects safetensors. Dataset adapters cover JSONL, WebDataset and LeRobot, so video-text pairs and robot action trajectories are both first-class inputs. The training guide lives in docs/training.md, with a separate page for the JSONL format.

On the inference side the backends are Diffusers, Transformers and vLLM, with offline batch generation and online serving through Ray and Gradio. The README also mentions shim libraries under packages/, described as lightweight standalone wrappers for downstream projects. That is a reasonable split: the heavy framework stays in one place, and projects that only need to call a model get a thin dependency instead of the whole training stack.

Installing Cosmos-Framework and running a first inference job

The README's setup path assumes a Linux machine with a CUDA toolkit already present. System dependencies come first, and they include ffmpeg and git-lfs, which tells you video decoding and large checkpoint downloads are part of the normal workflow.

bash
sudo apt-get install -y --no-install-recommends curl ffmpeg git-lfs libx11-dev tree wget

The package installs through uv, and the dependency group you pick has to match your CUDA toolkit. The README marks CUDA 13.0 as recommended and gives a 12.8 alternative in a comment.

bash
# CUDA 13.0 (recommended)
uv sync --all-extras --group=cu130-train
# Or, for CUDA 12.8:
# uv sync --all-extras --group=cu128-train
source .venv/bin/activate && export LD_LIBRARY_PATH=

That last line is worth pausing on. The README clears LD_LIBRARY_PATH after activating the environment, and the justfile does the same thing with an explanatory comment about environment variables that interfere with the uv virtual environment. If you have a CUDA setup that relies on LD_LIBRARY_PATH, expect to reconcile that with the project's expectation.

The first real use is a single-GPU inference run. The README gives this command, which reads a JSON input file describing the generation task and writes results to a directory:

bash
python -m cosmos_framework.scripts.inference \
    --parallelism-preset=latency \
    -i "inputs/omni/t2v.json" \
    -o outputs/omni_nano \
    --checkpoint-path Cosmos3-Nano \
    --seed=0

The checkpoint is referenced by name (Cosmos3-Nano), so the download path is handled somewhere in the setup documentation rather than by an explicit URL here. If that command fails on a missing checkpoint, docs/setup.md and the cosmos3-setup agent skill are the places the README points to.

The 8-GPU assumption in the shipped recipes

The examples directory contains the supervised fine-tuning recipes, and the README is explicit that they are 8-GPU configurations tested on 8x H100 80 GB. Launching one is a single shell command:

bash
bash examples/launch_sft_vision_nano.sh

The README says users may adjust the GPU count to match their model and hardware by tuning NPROC_PER_NODE and the parallelism degrees (DP/CP/FSDP shard) in the recipe. That is honest, but it is also the main practical hurdle. A recipe written for eight H100s does not degrade gracefully to one consumer card; you are editing distributed configuration before you see a loss curve. The recipe names themselves map to model sizes and modalities (vision, videophy2, action policy, llava), so picking the right starting point is part of the work.

There is a justfile that wraps the common operations. Running just with no arguments lists the available recipes, and there are targets for setup, install, run, lint and export-schemas. The install target runs uv sync with --all-extras and the CUDA group, then reinstalls. The run target executes a command inside the synced environment with --no-sync. For anyone who prefers a task runner to remembering flags, that is the lower-friction path, though the justfile is not mentioned in the README's setup section.

Licence, packaging and the Beta label

The pyproject.toml carries an SPDX identifier of OpenMDW-1.1, and the same identifier appears in the header comments of the Dockerfile and pyproject.toml. The repository's LICENSE and NOTICE files are the authoritative text; the GitHub metadata reports the licence as NOASSERTION, which means GitHub's detector could not classify it automatically. Anyone planning commercial deployment should read the actual licence file rather than trusting either label. This is not legal advice, and OpenMDW-1.1 is not a licence most engineers will recognise from memory.

The package declares itself as Development Status 4 - Beta, version 1.2.2, requiring Python 3.10 or newer. The dependency list is long and pinned in places: diffusers at 0.39.0 or later, transformers at 4.57.1 or later but below 5.0.0. The train extra pulls in megatron-core, lerobot, jupyterlab, boto3 and a long tail of media and numerics libraries. The serve extra is much smaller (fastapi, httpx, gradio, ray[serve]), which suggests inference serving is the lighter deployment target and training is the heavy one. The Dockerfile uses a CUDA 13.0.2 base image by default and installs uv 0.12.2 from its container image, so container builds are a supported path rather than an afterthought.

Agent skills and the repo map

One unusual choice: the repository ships instructions for coding agents. AGENTS.md is described as the canonical repo map that agents load first, and five skills live under .agents/skills/ with mirrors under .claude/skills/ for Claude Code. The skills cover setup, codebase navigation, inference, post-training and environment troubleshooting, and each is a self-contained SKILL.md that activates when a request matches its description.

This is a documentation decision as much as a tooling one. A codebase with distributed parallelism, checkpoint conversion and dataset adapters is hard to navigate from a README alone, and the skills encode the answers to questions like where a parameter is set. If you do not use an AGENTS.md-aware tool, the same content is still useful as a structured index. What the README does not say is how these skills are kept in sync with the code, and the mirroring between .agents/skills/ and .claude/skills/ means two copies exist.

Alternatives and where Cosmos-Framework sits

The closest comparison is the HuggingFace Diffusers and Transformers stack on its own. Cosmos-Framework depends on both and uses them as inference backends, but it adds the distributed trainer, the DCP checkpoint format with safetensors import and export, the LeRobot and WebDataset adapters, and the Ray plus Gradio serving layer. If you only need to run a Cosmos3 checkpoint for generation, the backends underneath may be enough and you avoid the training dependency group entirely. If you need to fine-tune on your own action data or serve at throughput, the framework is doing work the base libraries do not.

The other reference point is NVIDIA's own NeMo and Megatron tooling. The train extra already depends on megatron-core, so Cosmos-Framework is not competing with that ecosystem; it is a Cosmos-specific layer on top of it. Teams already standardized on NeMo for other model families will recognise the parallelism vocabulary, but the recipes, dataset adapters and checkpoint conversion here are specific to Cosmos3.

Who should adopt it, who should not, and what to check first

Adopt it if you have Linux hosts with NVIDIA GPUs, you intend to fine-tune or serve a Cosmos3 model, and you would rather start from NVIDIA's distributed recipes than assemble FSDP, checkpoint conversion and a serving layer yourself. The single-package layout and the paired launch shells lower the cost of the first experiment.

Do not adopt it if your hardware is anything other than Linux plus NVIDIA CUDA, if you need a CPU or macOS path, or if you need API stability. The Beta classifier and the 8-GPU recipe defaults are both stated plainly, and the README's own guidance is that GPU count depends on the recipe. A team without multi-GPU capacity will spend its time editing parallelism configuration rather than training.

Three things to verify before you commit. First, read docs/setup.md for system requirements and the CUDA variant that matches your toolkit, because the dependency group is not interchangeable. Second, read LICENSE and NOTICE for the OpenMDW-1.1 terms, since GitHub reports the licence as NOASSERTION and the SPDX text is the only reliable statement. Third, run the single-GPU inference command against Cosmos3-Nano before touching a training recipe, so that checkpoint download, ffmpeg and the environment variable handling are proven on your machine first.

Editorial conclusion

Adopt Cosmos-Framework if you already have Linux hosts with NVIDIA GPUs and want to fine-tune or serve the Cosmos3 family rather than build a training stack yourself. Do not adopt it if you are looking for a CPU-only or macOS path, or a stable API: the package is classified as Beta and the shipped recipes assume 8 GPUs. Before committing, check docs/setup.md for your CUDA variant, confirm the OpenMDW-1.1 terms in LICENSE against your intended use, and try the single-GPU inference command against a Cosmos3-Nano checkpoint to see whether your hardware and the checkpoint download actually work.

Frequently asked questions

What is the Cosmos model?

According to the README, Cosmos 3 is a suite of omnimodal world models built on a unified Mixture-of-Transformers architecture that jointly processes and generates language, images, video, audio and action sequences. Cosmos-Framework is the training and serving repository for that family.

Is NVIDIA Cosmos a world model?

Yes. The README describes Cosmos 3 as a suite of omnimodal world models and says the framework is for training and serving world models, including the Cosmos3 family.

Is NVIDIA Cosmos free to use?

The repository is public and the package declares an OpenMDW-1.1 licence in pyproject.toml, but GitHub reports the licence as NOASSERTION. The LICENSE and NOTICE files in the repository are the authoritative terms, and the README does not discuss pricing.

Official sources

  1. Issues
  2. NVIDIA/cosmos-framework on GitHub
  3. Project website
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/nvidia-cosmos-framework.svg)](https://hysenlabs.com/projects/nvidia-cosmos-framework)