Model or dataset
NVIDIA/cosmos avatar
NVIDIA/cosmos

NVIDIA Cosmos: An Omnimodal World Model Platform for Physical AI

NVIDIA Cosmos is an open platform of world models, datasets, and tools that enables developers to build Physical AI for robots, autonomous vehicles, smart infrastructure, and more.

11,943 stars894 forksJupyter NotebookNOASSERTION

At a glance

What is it?
Cosmos 3 unifies a reasoning transformer and a diffusion generator over text, image, video, audio and action. Here is what the repository actually documents, who the hardware targets fit, and where the gaps are.
Who is it for?
Adopt Cosmos if you already have H100, H200, B200, GB200, RTX Pro 6000 or Jetson AGX Orin class hardware and a Physical AI workload that needs either video and action generation or video-grounded reasoning. Do not adopt it if you need a CPU-only or small-GPU path, or if you need documented post-training today, since the README marks the post-training recipes as coming soon.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Jupyter Notebook, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

What Cosmos 3 Is For, and Who Should Care

Cosmos is a platform of world models, datasets and tools aimed at Physical AI: robots, autonomous vehicles and smart infrastructure. The current family, Cosmos 3, is described as a suite of omnimodal world models that jointly process and generate language, images, video, audio and action sequences inside a unified Mixture-of-Transformers architecture. The README frames this as subsuming vision-language models, video generators, world simulators and world-action models into one framework.

The practical split is between two runtime surfaces. Reasoner takes text and vision as input and returns text, covering world understanding, grounding, physical reasoning, task planning, action forecasting and embodied agent reasoning. Generator takes text, vision, sound and action as input and returns vision, sound and action, covering world generation, simulation, future prediction, synthetic data generation, policy learning and robot training.

That split matters for adoption. If your problem is perception or planning on top of camera streams, you want Reasoner. If your problem is producing training data or rollouts, you want Generator. The repository is not a single model you call one way; it is two modes with different serving stacks and different hardware expectations.

Mixture-of-Transformers: One Backbone, Two Attention Regimes

The architecture is a unified Mixture-of-Transformers combining an autoregressive transformer for reasoning with a diffusion transformer for multimodal generation. In Reasoner Mode, language and visual understanding tokens pass through causal self-attention for next-token prediction. In Generator Mode, noisy image, video, audio and action tokens are denoised through full attention, so the model produces coherent multimodal output jointly rather than one modality at a time.

The two modes share the transformer architecture, the multimodal attention layers and a unified 3D multi-dimensional rotary position embedding (mRoPE) that encodes spatial and temporal structure across modalities. That shared representation is the design bet: the same positional scheme has to describe an image grid, a video timeline, an audio stream and an action trajectory without modality-specific hacks.

The trade-off is visible in the mode split. Causal attention and full attention are different computation patterns, so a single checkpoint serves two very different inference paths. The README does not document how much of the weights are shared between modes, nor what the memory penalty is for loading one model that must serve both. Anyone sizing a deployment should treat that as an open question rather than assume the modes are cheap to co-host.

Model Family Sizes and the Hardware They Assume

Three checkpoints are listed. Cosmos3-Super is 64B and targets data center hardware: H200, B200 or GB200. Cosmos3-Nano is 16B and targets both data center and workstation: RTX Pro 6000, H100 or B200. Cosmos3-Edge is 4B and targets edge and on-device: Jetson AGX Orin, Thor or RTX Pro 6000.

Input coverage is text, image, video and action for all three. Output coverage differs. Super and Nano emit text, image, video, sound and action. Edge emits text, image, video and action, with no sound, and the README marks its video input with a footnote. So the smallest model is not a smaller version of the same thing; it drops audio generation and carries a caveat on video input.

That is the first real constraint to check before committing. A team that prototypes on Edge and then expects parity with Super will find a different output surface, not just different quality. The hardware list is also narrow in a specific way: there is no documented CPU path and no consumer-GPU tier below RTX Pro 6000, so a laptop or a single mid-range card is not a supported target for any listed size.

Installing Cosmos 3 and Running a First Generator Call

The README points to the Cosmos Framework repository for tooling and to a Hugging Face collection for the Cosmos 3 weights, and it documents several integration paths rather than one canonical install. For a Python-first start, the Diffusers path is the one the README lists first under Quickstart for the Generator surface. It does not give a pinned version string for the package, so treat the install line below as the package name only and resolve the version yourself.

bash
pip install diffusers

After installing, the documented pattern is to load a Cosmos 3 checkpoint from the Hugging Face collection and drive it through the Diffusers pipeline interface. The README does not print the full pipeline constructor, so check the collection card for the exact model identifier before running anything. What you should see is a generated video or image artifact for the Generator surface, or text for the Reasoner surface.

If you prefer an OpenAI-compatible server instead of in-process Python, the README lists vLLM-Omni and SGLang for Generator serving, and vLLM and TensorRT-LLM for Reasoner serving. For a turnkey container, NIM is the documented option for both Reasoner serving and Generator deployment for text-to-video and image-to-video generation.

bash
pip install vllm

The README does not document a port number, a launch flag or an environment variable for the vLLM path, so do not assume a default endpoint. The troubleshooting section does cover the failure you are most likely to hit first: if torch.cuda.is_available() returns False, the README attributes it to an NVIDIA driver that is too old. Two other documented failures are an import error for libxcb.so.1 and errors from uv on install or sync. The README also has entries for which CUDA version and which base container to use, which suggests the dependency matrix is the part new users trip over.

Where Cosmos 3 Is the Wrong Tool

The README's own Limitations section is the place to start, and the post-training recipes are marked as coming soon. That means the documented surface today is inference and integration, not a finished fine-tuning workflow, even though the table of contents lists Finetune, Export and Convert Checkpoints, and Distill sections. A team whose plan depends on adapting the model to a proprietary action space should treat that as unproven until those recipes ship.

Hardware is the second boundary. Every recommended configuration is an NVIDIA accelerator at workstation class or above. There is no documented path for CPU inference, and no listed support for older or smaller cards. If your deployment target is a fleet of modest embedded boards, the Edge variant names Jetson AGX Orin and Thor, which are not the low end of the Jetson line.

Modality gaps are the third. Edge omits sound output entirely and carries a footnote on video input. If synchronized audio is part of your simulation requirement, Edge is disqualified on the model card alone.

Finally, the repository is a platform rather than a library. The README routes you to a separate framework repository, a model collection and NIM containers. If you wanted a single pip-installable package with one API, this is a different shape of project, and the integration choice in the README's Choosing an Integration section is a decision you have to make rather than one the project makes for you.

How This Differs from a Video Generation Model or a VLM

The closest familiar alternative is a standalone video diffusion model paired with a separate vision-language model. In that setup you run a VLM to caption or reason about frames, then feed text into a video generator. The two models have separate weights, separate tokenizers and separate serving stacks, and nothing enforces that the reasoning model and the generator agree about space or time.

Cosmos 3's difference is architectural rather than packaging. One Mixture-of-Transformers backbone holds both an autoregressive reasoning path and a diffusion generation path, and both use the same 3D mRoPE representation across images, video, audio and action. The claim in the README is that this unifies the modalities into a single framework rather than gluing outputs together.

The second difference is action. A video generator produces pixels; Cosmos 3 also takes and emits action sequences, with the README listing policy actions, inverse dynamics and forward dynamics for robotics, camera motion, egocentric motion and autonomous driving. That is a capability a generic text-to-video model does not have at all.

The cost of that unification is the mode split described earlier and the hardware floor. A smaller, task-specific video model will run on hardware Cosmos 3 does not support. The question is whether you need the joint representation, because you pay for it in memory and in integration choices.

Licence, Maintenance Signals and Upgrade Cost

The repository reports its licence as NOASSERTION, and a LICENSE file sits at the top level. NOASSERTION means the automated classifier could not map the file to a recognized SPDX identifier, so the terms are whatever that file says, plus whatever terms attach to the model weights on the Hugging Face collection and to any NIM container you pull. Those are three separate artifacts and they may not carry the same terms. Read the LICENSE file and the model card before you build a product around this; this is a description of what the repository reports, not legal advice.

The last push to the default branch was on 2026-09-20, so the repository is being updated. The only listed release is Cosmos3, dated 2026-06-01. For upgrade cost, the practical exposure is the integration path you pick. If you build against the Diffusers pipeline, an upstream change in that library can affect you independently of Cosmos. If you build against a NIM container, you inherit the container's release cadence instead. The README's Choosing an Integration section exists precisely because these paths have different maintenance profiles, and the repository does not promise API stability across them.

Editorial conclusion

Adopt Cosmos if you already have H100, H200, B200, GB200, RTX Pro 6000 or Jetson AGX Orin class hardware and a Physical AI workload that needs either video and action generation or video-grounded reasoning. Do not adopt it if you need a CPU-only or small-GPU path, or if you need documented post-training today, since the README marks the post-training recipes as coming soon. Verify first that your target model size matches your memory budget, that your chosen serving path is one of the documented integrations, and that the licence terms in the LICENSE file cover your intended use, because the repository reports NOASSERTION rather than a named licence.

Frequently asked questions

How do I install NVIDIA Cosmos?

The README does not give one canonical install. It points to the Cosmos Framework repository for tooling and to a Hugging Face collection for the Cosmos 3 weights, and it documents several integration paths including Diffusers for Python, vLLM-Omni and SGLang for Generator serving, vLLM and TensorRT-LLM for Reasoner serving, and NIM containers.

How do I use NVIDIA Cosmos 3?

You pick a runtime surface first. Reasoner takes text and vision and returns text for understanding, grounding and planning. Generator takes text, vision, sound and action and returns vision, sound and action for world generation, simulation and synthetic data. Then you pick an integration such as Diffusers, vLLM, TensorRT-LLM, SGLang or NIM.

What does Cosmos 3 mean in this context?

In the repository, Cosmos 3 is the newest NVIDIA model family: a suite of omnimodal world models that jointly process and generate language, images, video, audio and action sequences within a unified Mixture-of-Transformers architecture. It is not a reference to the flower or to any unrelated product.

How do I use Cosmos?

The README separates the two runtime surfaces: Reasoner for text and vision input returning text, and Generator for text, vision, sound and action input returning vision, sound and action. After choosing a surface, you choose an integration from the README's list, which includes Diffusers, vLLM-Omni, vLLM, TensorRT-LLM, SGLang and NIM.

Official sources

  1. Issues
  2. NVIDIA/cosmos on GitHub
  3. Project website
  4. README
  5. Releases
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/nvidia-cosmos.svg)](https://hysenlabs.com/projects/nvidia-cosmos)
Community notes

Community notes