# Four agents in one world, at 24 frames per second

> Gamma-World is NVIDIA's generative world model for interactive video with more than one controllable agent, built on a simplex vertex encoding that gives every agent a distinct rotary phase with no fixed ordering and no learned identities, a hub-token attention scheme that makes cross-agent cost linear rather than quadratic, and a distilled causal student that streams at 24 FPS while generalising from two players to four without retraining.

**nv-tlabs/Gamma-World** — Implementation of Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players

- Repository: https://github.com/nv-tlabs/Gamma-World
- Website: https://research.nvidia.com/labs/sil/projects/gamma-world/
- Stars: 666 · Forks: 14
- Language: Python
- License: Apache-2.0
- Published: 2026-09-20 · Updated: 2026-09-20 · Language: en
- Canonical page: https://hysenlabs.com/projects/nv-tlabs-gamma-world

## The hard part is identity, not pixels

The problem statement is narrow and it explains the whole design.

World models for interactive video have concentrated on single-agent settings, where the future is generated from one control signal. But many generated environments need several actors at once: multiple players, several robots, or embodied agents moving in the same space.

Scaling to that case imposes requirements that a single-agent model never has to satisfy. Agents must stay independently controllable, so one player's input cannot change another's. They must be permutation-symmetric, so swapping two players produces a valid, equivalent situation. Inference has to stay efficient as the count grows. And the rollout has to remain consistent across time and across viewpoints, because two agents watching the same room have to see the same room.

That last requirement is what makes this hard. A model can generate a plausible scene per agent and still be wrong, because two independently plausible scenes are not one shared world.

The claimed reach is from multiplayer virtual games to real-world multi-robot environments, and the gallery shows two-agent interaction, four-agent generalisation and robotics coordination. The paper, the project page and the code were released in May and June 2026.

## Simplex vertices instead of slot indices

Agent identity is handled by an encoding rather than by a convention, and the encoding is described as parameter-free.

Simplex Rotary Agent Encoding extends three-dimensional rotary position encoding. Instead of giving agents scalar indices, or learned per-slot identity vectors, the agents are placed at the vertices of a regular simplex in rotary angle space.

A regular simplex has the property the design needs: all vertices are equidistant from each other. So every pair of agents is equivalent under permutation, which removes the need for a fixed agent ordering, and each agent still receives a distinct rotary phase, which keeps them distinguishable to the network.

Compare that with the usual alternatives. Scalar indices impose an order that means agent one and agent two are treated differently by construction. Learned identity vectors cost parameters and have to be trained, and they do not extend to an agent count you did not train on. A simplex has neither problem: adding a player adds a vertex.

That is the mechanism behind the four-agent result, which is claimed to work without any additional training, since the encoding can represent four agents as readily as two.

## Hub tokens make cross-agent attention linear

Even with a clean identity scheme, letting every agent attend to every other agent costs a quadratic number of interactions, and the number of agents is exactly the axis you want to grow.

Sparse Hub Attention replaces pairwise attention with a small set of learnable hub tokens. Agent tokens attend to their own stream and to the hubs. The hubs aggregate information from across the agents and broadcast it back out.

So the communication path becomes two hops instead of a complete graph, and the cross-agent cost drops from quadratic to linear in the number of agents. The shared-world consistency requirement is still met, because the hubs are the shared channel: what one agent learns can reach another through them.

The efficiency claim follows from that structure: lower latency and fewer floating point operations as agent count increases, which is the regime that matters for four or more agents rather than for the two-agent case.

One design consequence is worth naming. The hubs are a bottleneck by construction, a fixed number of vectors through which all inter-agent information passes. The documentation does not quantify how much detail survives that squeeze, so the scaling win and the fidelity limit come from the same mechanism.

## Twenty-four frames per second comes from the student, not the teacher

Real-time generation is handled by distillation, and the two models do different jobs.

The teacher is a full-context diffusion model. The student is a causal model that generates temporal blocks sequentially and keeps KV caches for past visual tokens and for hub states. That caching is what preserves block-causal generation while remaining efficient as the number of agents grows.

The student is the one that runs in real time, and the claimed figure is 24 FPS of action-responsive generation. So when you read a latency number for this project, the model doing the generating matters: the teacher sees full context and the student sees a cache.

The documentation also names three inference modes you can run on your own initial frames and actions: bidirectional, causal, and causal few-step. The few-step variant is the one that trades quality for latency, and the name suggests it is the mode to reach for when the frame budget is the binding constraint.

The training pipeline has three stages, which the training guide describes as teacher, causal, and a DMD stage, and it starts from a released checkpoint that has to be converted to DCP first. That conversion step is documented separately, which is a good sign that it is not trivial.

## Two players to four, with no retraining

The headline generalisation result is that the model goes from two players to four without additional training.

That claim is exactly what the simplex encoding is for. Because identity is positional in rotary angle space rather than tied to a learned slot, four agents are four vertices of a tetrahedral arrangement in the same space where two agents were two vertices of a line. Nothing about the representation has to be extended or fine-tuned.

If the claim holds, it changes what a world model is for. A model trained on two-player interaction can be used for a four-player scenario without collecting four-player data, which is the expensive part of multi-agent simulation.

The same section extends the reach to real-world multi-robot coordination, demonstrated beyond virtual environments. That is a much bigger claim than the four-agent result, because it involves transferring from generated or game-like data to physical robots, and the evidence offered for it is qualitative results rather than a benchmark table.

So read the two claims at different weights. The scaling result is a structural consequence of the encoding and should reproduce. The robotics transfer is a demonstration, and demonstrations of physical systems are where the evaluation is hardest.

## CUDA 12.8 or 13.0, picked by CPU architecture

The packaging is unusually small and the two decisions in it are the ones to get right.

There is exactly one runtime dependency, pinned to an exact version, and the optional extras are two variants of it. One extra pulls a CUDA 12.8 stack with Torch 2.7, the other a CUDA 13.0 stack with Torch 2.9. The tool configuration lists those two extras as conflicting, so they are mutually exclusive by design and choosing one is the install decision.

Python 3.10 or newer is required, the operating system classifier is Linux only, and the audience classifier is science and research rather than end users.

The part that will surprise you is in the task runner. The install recipe picks the CUDA extra from the CPU architecture: on aarch64 it selects the 13.0 stack, and on everything else the 12.8 stack. The same file sets a default GPU count for the same reason, four on aarch64 and eight elsewhere.

So a single install command produces a different toolchain depending on the machine it runs on. If you are on ARM hardware and expected the same stack as your x86 colleague, this is where that happens, and the choice is written to a small file during install so later commands reuse it.

## A CUDA devel image with uv and just preinstalled

The container image is a developer image rather than a runtime image, and its contents say what the workflow expects.

The base is a CUDA devel image with cuDNN on Ubuntu, addressed through a build argument for the target platform so the same file can be built for more than one architecture. The environment is set to non-interactive so apt calls do not hang waiting for a prompt.

The installed system packages are the ones a video project needs: ffmpeg, which is the non-obvious and most important of them for anything that decodes or encodes frames, git with large-file support for checkpoints, and curl, wget and tree.

Two tools are installed by copying binaries in rather than by package manager. uv is pulled from a pinned container image tag and placed in the binary directory, and just is fetched by a curl-to-shell installer pinned to a specific release tag. Both are pinned exactly, which is the right call for reproducibility and means upgrades are deliberate.

Two environment settings exist for a reason worth knowing. One switches uv to copy instead of link mode, with the comment explaining that it is because of a mounted volume, and another points uv's tool binaries at the same system directory so installed tools are on the path out of the box.

If you are not using the container, you will need ffmpeg and git-lfs on the host anyway.

## The task file still globs for a package named cosmos

A few things in the task runner show the project's origin, and one of them will bite a contributor.

Three helper recipes compute names by globbing. They echo a pattern matching cosmos-prefixed directories, strip a prefix from them, and convert the underscore form to a hyphenated form. This repository's package is named gamma_world, not cosmos-something, so those recipes produce nothing useful here.

That is a leftover from whatever the repository was derived from, and unlike a stale comment it sits in executable recipes. Harmless for the documented setup path, wrong for any task that relies on those names.

The type checking is also worth reading. There are two pyrefly configurations, a default one and one pointed at a separate source config, and a recipe that runs both. One recipe suppresses unused ignores automatically, which keeps the file from accumulating stale suppressions. Another whitelists every error, which is a blunt instrument that belongs in a local experiment rather than in a default.

The test-install recipe is the tidiest thing here: it removes the recorded CUDA choice, syncs without an extra, and then runs Python expecting a specific message about the CUDA extra not being installed. That is an install smoke test that asserts the failure mode rather than the success one.

And lint is wired as pre-commit, which runs twice, the second run as a fallback. Running twice is what fixes files that a formatter changes in a way that another hook rejects.

## Conclusion

It fits a research team working on interactive simulation or multi-robot planning who needs agent identity to scale without reordering assumptions, and who can afford a CUDA workstation. Three things to check first. The repository carries a single runtime dependency pinned to a specific version, and the two CUDA extras are declared mutually exclusive, so your install is a choice between toolchain versions rather than a default. The justfile picks that CUDA extra from CPU architecture, which means the same command gives you a different stack on an ARM machine than on x86. And the efficiency claim is about FLOPs and latency as agent count rises, while the 24 FPS figure comes from the distilled student, so the teacher and the student are not the same thing to benchmark.

## FAQ

### What is Gamma-World?

It is a generative multi-agent world model from NVIDIA Labs for interactive simulation, released in 2026 under Apache-2.0. It rolls out a single shared environment from several independently controllable agents, aimed at multiplayer games and real-world multi-robot coordination, with the paper on arXiv and the project page hosted on the NVIDIA research site.

### How does Gamma-World give agents distinct identities?

With Simplex Rotary Agent Encoding, a parameter-free extension of three-dimensional rotary position encoding that places agents at the vertices of a regular simplex in rotary angle space. All agents are equidistant, so every pair is permutation-equivalent, while each keeps a distinct rotary phase. No learned per-slot identities or fixed agent ordering are needed.

### How does Gamma-World keep attention cost low as agents increase?

Through Sparse Hub Attention, where a small set of learnable hub tokens mediates communication. Agent tokens attend to their own stream and to the hubs, and the hubs aggregate across agents and broadcast back, reducing cross-agent attention from quadratic to linear in the number of agents while keeping a shared channel for world consistency.

### How does Gamma-World generate video in real time?

It distills a full-context diffusion teacher into a causal student that generates temporal blocks sequentially with KV caching for past visual tokens and hub states, which is what enables action-responsive generation at 24 FPS. Inference can be run in bidirectional, causal, or causal few-step modes on your own initial frames and actions.

### What does installing Gamma-World require?

Python 3.10 or newer on Linux, a CUDA GPU, and uv for dependency management. The package has a single pinned runtime dependency with two mutually exclusive extras, one for a CUDA 12.8 stack with Torch 2.7 and one for CUDA 13.0 with Torch 2.9, and the task runner selects between them from the CPU architecture, choosing 13.0 on aarch64 and 12.8 elsewhere.

## Sources

- [Issues](https://github.com/nv-tlabs/Gamma-World/issues)
- [License: Apache-2.0](https://github.com/nv-tlabs/Gamma-World/blob/main/LICENSE)
- [nv-tlabs/Gamma-World on GitHub](https://github.com/nv-tlabs/Gamma-World)
- [Project website](https://research.nvidia.com/labs/sil/projects/gamma-world/)
- [README](https://github.com/nv-tlabs/Gamma-World/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/nv-tlabs-gamma-world
