Open-source project
amap-cvlab/ABot-World avatar
amap-cvlab/ABot-World

ABot-World: Running an Interactive World Model on One Desktop GPU

Infinite Interactive World Rollout on a Single Desktop GPU

2,513 stars129 forksPythonApache-2.0

At a glance

What is it?
ABot-World is a Python world-model project from amap-cvlab that turns a single NVIDIA RTX 5090 into a real-time, action-conditioned world simulator. Its setup path, its hardware ceiling and its Apache-2.0 terms decide whether it fits your machine.
Who is it for?
Adopt ABot-World if you have a single NVIDIA RTX 5090 desktop and want to study action-conditioned world rollout at 720p without a cluster, since the README states 16 FPS, 1.2s latency and 19GB GPU memory on that card. Do not adopt it if your hardware is a laptop GPU, a multi-GPU server, or a non-NVIDIA accelerator, because the README names only the RTX 5090 and the requirements.txt pins xformers and torchao builds that assume a matching stack.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 17 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem ABot-World targets: video models that stop when the clip ends

Most video generation models produce a fixed-length clip. You send a prompt, you get seconds of footage, and the model has no notion of what happens if you keep going. ABot-World attacks that boundary directly. The README describes it as turning a single NVIDIA RTX 5090 desktop GPU into a real-time interactive world simulator, with infinite action-conditioned world rollout at 720P, 16 FPS, 1.2s latency and 19GB GPU memory. The audience is narrow and specific: a researcher or engineer who wants to drive a world model with actions, watch the scene respond, and keep rolling past the point where a normal clip would have ended, all on one desktop card rather than a multi-GPU node. The project also ships a local gradio demo and an online playground called ABot World Studio, so the same model can be inspected without writing an inference loop first.

LongForcing and the causal student model behind open-ended rollout

The mechanism the README names is LongForcing, described as a training approach that lets the world expand with new scenes and dynamics during rollout, avoiding scene lock-in, without prompt switching. That is the part worth reading carefully, because it is the difference between a demo that loops and a rollout that keeps producing new content. The released artifact is a causal student model, ABot-World-0-5B-LF, which the news section lists alongside inference code, a local gradio demo and the online playground. The repository layout supports that reading: pipeline/, quantizer/, wan/, configs/ and utils/ sit next to web_client/, which suggests a generation pipeline, a quantisation step, a Wan-family backbone, configuration files and a browser-facing client. The 5B parameter size in the model name, combined with the 19GB memory figure, is the constraint that makes the desktop claim plausible. The README does not explain how LongForcing is implemented, so anyone planning to retrain rather than run inference will be reading the arXiv report and the training code, not the README.

Installing ABot-World and running a first rollout

The README's Setup section is truncated in the repository text, so the safest entry point is the Docker image announced on 2026-09-08, which the news section says is available on Docker Hub and Alibaba Cloud Container Registry. If you prefer a local environment, the repository ships requirements.txt at the top level. Note that the file pins specific builds, including xformers==0.0.32.post2 and torchao==0.17.0, so a mismatched CUDA or PyTorch install is the most likely first failure.

bash
pip install -r requirements.txt

The requirements file also pins gradio==6.9.0 and gradio-client==2.3.0, which is consistent with the local gradio demo the news section mentions. After installing, the model weights come from the HuggingFace page acvlab/ABot-World-0-5B-LF or the ModelScope mirror amap_cvlab/ABot-World-0-5B-LF. The README does not print the exact download command, so use the model page's own instructions rather than guessing a filename.

The fastest first use is the hosted Space at huggingface.co/spaces/acvlab/abot-world-interactive, or the online playground at abot-world.amap.com. Both let you drive the world with actions before you spend time on a local environment. For the local gradio demo, the top-level repository entries include web_client/, css/, js/ and index.html, so the demo is a web front end rather than a terminal-only script. The README does not document the launch command for that demo, which is a real gap if you are trying to reproduce the 16 FPS figure on your own machine.

The RTX 5090 is the ceiling, not a suggestion

Every performance number in the README is tied to one card: a single NVIDIA RTX 5090 desktop GPU, 720p, 16 FPS, 1.2s latency, 19GB GPU memory. That is a strong claim and a narrow one. If your workstation has a 24GB card of an older generation, the memory figure alone does not tell you whether you are fine, because the pinned xformers and torchao builds in requirements.txt are version-sensitive and the README does not list a supported matrix. If you are on a laptop GPU, an AMD or Intel accelerator, or a CPU-only machine, the README offers nothing: it names no fallback path, no reduced-resolution mode and no CPU inference option. The honest reading is that ABot-World is a single-target project. It is the wrong tool when you need to serve many concurrent users, when you need determinism across hardware, or when your goal is video editing rather than action-conditioned simulation. It is also the wrong tool if you want a stable API: the project is a research release with a 0.x model name, and the news timeline shows the model, datasets and Docker image arriving in separate steps across July, August and September 2026.

ABot-World versus a general-purpose video diffusion pipeline

The closest comparison is a general-purpose text-to-video diffusion pipeline, the kind you build from the diffusers library that appears in requirements.txt. The difference in approach is the conditioning signal. A text-to-video pipeline takes a prompt and returns a clip; the prompt is fixed for the duration of the generation, and the output length is baked into the model's training. ABot-World takes actions as input and produces a rollout, which is why the README frames it as continuous exploration instead of passive video playback, and why LongForcing exists to avoid scene lock-in without prompt switching. The trade-off is scope. A general video pipeline can be fine-tuned for many styles and resolutions and runs on a wide range of accelerators; ABot-World is tuned for interactive rollout at 720p on one specific desktop GPU, and its released weights are a 5B causal student model rather than a general-purpose generator. If your task is producing a polished 30-second clip, the general pipeline is the better fit. If your task is studying how a world model behaves when a user keeps acting on it, ABot-World is aimed at exactly that.

Licence, maintenance and what an upgrade costs you

The repository is Apache-2.0, and it carries NOTICE and THIRD_PARTY_NOTICES.md files at the top level, which is the normal signal that some bundled components are under different terms. Read both before you redistribute anything; the README does not summarise which parts are covered by which licence, and I am not going to guess. On maintenance, the last push was on 2026-09-14, and the repository is not archived, so the project is being touched. The release cadence visible in the news section is roughly monthly: the model and inference code on 2026-07-09, the 500-hour dataset on 2026-08-03, the 24-hour rollout demo on 2026-08-15, and the Docker image on 2026-09-08. That cadence is the upgrade cost. The requirements.txt pins exact versions across roughly thirty packages, including diffusers==0.37.0, transformers==5.4.0 and gradio==6.9.0, so a new checkpoint can drag a new dependency set with it. There are no retrieved releases, so there is no changelog to diff against; you will be comparing commits and model cards. Budget for a container rebuild, not a pip upgrade.

Editorial conclusion

Adopt ABot-World if you have a single NVIDIA RTX 5090 desktop and want to study action-conditioned world rollout at 720p without a cluster, since the README states 16 FPS, 1.2s latency and 19GB GPU memory on that card. Do not adopt it if your hardware is a laptop GPU, a multi-GPU server, or a non-NVIDIA accelerator, because the README names only the RTX 5090 and the requirements.txt pins xformers and torchao builds that assume a matching stack. Before you commit, open the Docker image announced on 2026-09-08, confirm the checkpoint you intend to load from the HuggingFace or ModelScope model page, and read THIRD_PARTY_NOTICES.md to see which bundled components carry terms other than Apache-2.0.

Frequently asked questions

What hardware does ABot-World need to run?

The README states it runs at 720p and 16 FPS with 1.2s latency and 19GB GPU memory on a single NVIDIA RTX 5090 desktop GPU. No other accelerator or a CPU fallback is documented.

How do I install ABot-World?

The repository ships a requirements.txt at the top level, and the news section says a Docker environment image is available on Docker Hub and Alibaba Cloud Container Registry as of 2026-09-08. The README's Setup section is truncated, so the Docker image is the better-documented path.

Where can I try ABot-World without installing it?

The README links a HuggingFace Space at acvlab/abot-world-interactive and an online playground called ABot World Studio at abot-world.amap.com. Both are listed alongside the local gradio demo.

Which model checkpoint does ABot-World release?

The news section names ABot-World-0-5B-LF as the causal student model, hosted on HuggingFace at acvlab/ABot-World-0-5B-LF and mirrored on ModelScope at amap_cvlab/ABot-World-0-5B-LF.

What is LongForcing in ABot-World?

The README describes LongForcing as the training approach that expands the world with new scenes and dynamics during rollout, avoiding scene lock-in, without prompt switching. The README does not explain how it is implemented.

Official sources

  1. amap-cvlab/ABot-World on GitHub
  2. Issues
  3. License: Apache-2.0
  4. Project website
  5. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/amap-cvlab-abot-world.svg)](https://hysenlabs.com/projects/amap-cvlab-abot-world)