LingBot-World: an open-source world model you should probably read before you clone
Advancing Open-source World Models
At a glance
- What is it?
- LingBot-World is Robbyant's Apache-2.0 video-generation world simulator, released in January 2026 and superseded by LingBot-World-Infinity. Here is how its camera and action models work, how to run the 480P inference script, and why the repository itself tells you to use the v2 code instead.
- Who is it for?
- Adopt LingBot-World only if you are reproducing the January 2026 technical report or need the Base (Cam) and Base (Act) checkpoints for a camera- or action-conditioned experiment, and clone lingbot-world-v2 instead if you want anything that will still receive fixes. Skip it if you need a supported pipeline, a documented API, or a single-GPU path, because the README names no such path and the repository now points elsewhere.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 83 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 20, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What LingBot-World actually simulates, and who the repository is for
LingBot-World is a world simulator built on video generation. The README describes three properties it aims at: high fidelity across realistic, scientific and cartoon environments; minute-level horizons with contextual consistency over time, which the project calls long-term memory; and real-time interactivity at 16 frames per second with latency under one second. The intended audiences named in the README are content creation, gaming and robot learning.
The practical audience is narrower than that list. Inference is driven by torchrun across multiple processes, the codebase is a fork of Wan2.2, and the checkpoints are conditioned on either camera poses or discrete actions. That is a research and prototyping tool for people who already run diffusion video models on multi-GPU machines. If you want a hosted demo, the README points to an online demo hosted by Reactor rather than to a local one-command path.
Camera poses, action tokens and the Wan2.2 lineage
The repository ships two conditioning schemes. LingBot-World-Base (Cam) takes camera poses; LingBot-World-Base (Act) takes actions; LingBot-World-Fast also takes camera poses and exists to shorten inference. All three are listed at 480P and 720P.
Camera control is expressed as two NumPy arrays. The README specifies intrinsics.npy with shape [num_frames, 4] holding [fx, fy, cx, cy], and poses.npy with shape [num_frames, 4, 4] holding one transformation matrix per frame in OpenCV coordinates. The README notes these can be produced from an existing video using ViPE, an external project by nv-tlabs. Nothing in the repository generates them for you.
The codebase is built on Wan2.2, and the README defers to Wan2.2's documentation for installation. The wan/ directory in the repository is where that inherited code lives, and pyproject.toml declares it as the only packaged module. The generation entry points sit at the top level: generate.py, generate_fast.py, and shell wrappers run_fast.sh, run_act2cam.sh and run_act2cam_string.sh. The act2cam scripts are the interesting ones, since they suggest a path from action-conditioned output to camera-conditioned output; the README does not explain what they do beyond their names.
Installing LingBot-World and running the 480P Base (Cam) script
The README gives installation steps directly. Clone the repository, then install the pinned requirements. The comment in the README states that torch must be at least 2.4.0, and requirements.txt agrees.
git clone https://github.com/robbyant/lingbot-world.git
cd lingbot-world
pip install -r requirements.txtflash_attn is installed separately, without build isolation, because it compiles against your installed torch:
pip install flash-attn --no-build-isolationWeights come from HuggingFace or ModelScope. The README shows the Base (Cam) download for both:
pip install "huggingface_hub[cli]"
huggingface-cli download robbyant/lingbot-world-base-cam --local-dir ./lingbot-world-base-camFor a first real run, the README's 480P example for Base (Cam) uses eight processes, FSDP for the DiT and T5, an Ulysses sequence-parallel size of 8, and 161 frames. The command below is the README's own example with the prompt shortened; the full prompt in the README is a paragraph describing a flight through a fantasy jungle.
torchrun --nproc_per_node=8 generate.py --task i2v-A14B --size 480*832 \
--ckpt_dir lingbot-world-base-cam --image examples/00/image.jpg \
--action_path examples/00 --dit_fsdp --t5_fsdp --ulysses_size 8 \
--frame_num 161 --prompt "The video presents a soaring journey through a fantasy jungle."Note the two path arguments. --image points at one file, --action_path points at a directory, and the README's example uses examples/00 for both, which is where the intrinsics.npy and poses.npy for that sample are expected to be. The README does not document the output filename or output directory, so check the script before assuming where the video lands.
The maintenance problem is stated in the README itself
The most important line in this repository is a notice at the top: this repository is no longer actively maintained, and readers are directed to https://github.com/Robbyant/lingbot-world-v2 under the name LingBot-World-Infinity, where future updates, models and features will be released. The last push to this repository was on 2026-07-09.
That is not a footnote. It means bug reports against generate.py or the act2cam scripts have no owner, and the model weights published here are the ones the project has moved past. The news list ends in April 2026 with the Base (Act) scripts, so the repository's own history stops before the handoff notice does.
There is a second, quieter constraint. The README claims sub-second latency at 16 frames per second, but the documented path to that is the Fast variant, whose model weights arrived on 2026-04-02 and whose inference scripts arrived on 2026-04-07. The Base (Cam) example above is a 161-frame, 8-process job, which is a batch generation shape rather than an interactive one. If interactivity is what you need, the Fast scripts are the ones to read, and the README does not describe how they differ internally.
How LingBot-World compares with Genie 3
Genie 3 is the obvious comparison, and the search data suggests people make it. The difference that matters is access. Genie 3 is a closed model; LingBot-World publishes code, weights and a technical report under Apache-2.0, with checkpoints on both HuggingFace and ModelScope. If your requirement is to inspect the sampler, change the conditioning representation, or fine-tune on your own camera trajectories, the open release is the only one of the two where that is possible at all.
If your requirement is a working interactive demo without owning eight GPUs, the comparison inverts. The README's answer to that case is not a local install but the third-party demo at Reactor. Against other open video models, the distinguishing feature is the control interface rather than the backbone: most image-to-video projects accept an image and a prompt, while LingBot-World additionally accepts per-frame camera intrinsics and extrinsics, which is what makes it a world model in the project's sense rather than a clip generator.
Licence and the cost of staying on this branch
The repository is Apache-2.0, declared in LICENSE.txt and referenced from pyproject.toml. That permits commercial use and modification, and it is a permissive licence rather than a copyleft one, so derivative work does not have to be released. Model weights are distributed separately through HuggingFace and ModelScope collections; the README does not state a separate licence for the weights, so check the model card before you assume the Apache-2.0 terms carry over to them.
Upgrade cost here is unusual: the correct upgrade is a migration. Because the README directs users to lingbot-world-v2, any work built against this repository's generate.py, its checkpoint names, or the act2cam wrappers will need to be re-checked against the v2 code. There are no releases in this repository to pin against, only the main branch, and the last commit was 2026-07-09.
Editorial conclusion
Adopt LingBot-World only if you are reproducing the January 2026 technical report or need the Base (Cam) and Base (Act) checkpoints for a camera- or action-conditioned experiment, and clone lingbot-world-v2 instead if you want anything that will still receive fixes. Skip it if you need a supported pipeline, a documented API, or a single-GPU path, because the README names no such path and the repository now points elsewhere. Before committing, verify the flash_attn build against your torch version and check that examples/00 ships the intrinsics.npy and poses.npy files your run expects.
Frequently asked questions
What is LingBot-World?
It is an open-source world simulator from the Robbyant Team built on video generation, released under Apache-2.0 with a technical report, code and model weights. The README describes it as supporting camera-pose and action control, minute-level horizons, and real-time interactivity.
How do I install LingBot-World?
Clone the repository, run pip install -r requirements.txt with torch at 2.4.0 or newer, then install flash_attn with --no-build-isolation. The README notes the codebase is built on Wan2.2 and defers to Wan2.2's documentation for installation details.
How do I use LingBot-World?
Prepare an input image, a text prompt, and optionally control signals as intrinsics.npy and poses.npy, then run the reference script for your model variant. The README's Base (Cam) 480P example uses torchrun with eight processes, --task i2v-A14B, --size 480*832 and --frame_num 161.
How can I try LingBot-World without installing it?
The README thanks Reactor for providing an online LingBot-World demo and links to https://www.reactor.inc/ for it. The repository itself documents no hosted endpoint of its own.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/robbyant-lingbot-world)