# HelixWorld: An Open-Source Interactive Audio-Visual World Model from Noiz AI

> HelixWorld is an open-source audio-visual world model that takes a single starting image and a text prompt, then lets you walk through the resulting scene in real time: camera navigation updates both the video and the spatial sound field simultaneously. It is built for researchers and engineers working on interactive world models, generative environments, and spatial audio-visual generation. The Preview v1 checkpoint requires Linux, Python 3.11, CUDA 12.x, and approximately 80 GB of VRAM.

**NoizAI/HelixWorld** — 🪐 HelixWorld: real-time interactive audio-visual world model.

- Repository: https://github.com/NoizAI/HelixWorld
- Website: https://helixworld.org
- Stars: 616 · Forks: 43
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/noizai-helixworld

## What HelixWorld Does and Who It Is For

Most video generation models produce a fixed output from a text prompt: the viewer watches but cannot navigate. HelixWorld is designed for interactive navigation. The README describes the core loop as: give it an image and a prompt, then walk forward or turn around, and the picture and sound update together. The spatial sound field follows the camera rather than being a fixed soundtrack added after the fact.

The project is from Noiz AI and was released as Preview v1 on 2026-09-03, with inference code and a preview model checkpoint published on Hugging Face. An online interactive demo runs at helixworld.org, allowing browser-based exploration without local setup. The repository is the offline inference path for researchers who want to generate their own navigation sequences or study the model's behavior outside the hosted demo.

The target audience is researchers in generative models, computer vision, and spatial audio, and engineers building interactive simulation or virtual environment systems who want to study a joint audio-video generation approach.

## How the Model Works: Joint Audio-Video Generation

The README describes the key design property directly: audio is not a soundtrack laid on afterwards. The spatial sound field turns with the camera. This distinguishes HelixWorld from systems that generate video and then separately add ambient audio; here the two modalities are generated as a unified output tied to the camera's position and orientation.

Navigation is specified through an actions string. The supported actions are W (forward), A (left strafe), S (backward), D (right strafe), left, right, up, down, and stop. Actions can be combined with a plus sign (W+D:8 for diagonal movement). The number after the colon is the count of latent transitions for that action segment. At the standard output size of 121 frames, there are 15 latent transitions. The last action segment may omit its duration.

The model uses a Gemma 3 12B text encoder (downloaded into models/text_encoder/gemma-3-12b/ by the download script) alongside the main model weights (models/weights/model.safetensors). Output is an MP4 file written to the output directory under release/native/.

The perspective parameter accepts first_person as the documented value in the README example.

## Installing HelixWorld and Generating a First Navigation Clip

The README specifies Linux with Python 3.11 and CUDA 12.x. A reviewed BF16 single-GPU run uses approximately 70 GB of VRAM; the README recommends 80 GB. ffmpeg and ffprobe must be installed system-wide as separate requirements.

Clone the repository, create the conda environment, install dependencies, and download the model weights:

```bash
git clone https://github.com/NoizAI/HelixWorld.git
cd HelixWorld
conda create -n helixworld-preview python=3.11 -y
conda activate helixworld-preview
python -m pip install -r requirements.txt
python download_models.py
```

After download, the models directory contains the text encoder and the main checkpoint:

```text
models/
├── text_encoder/gemma-3-12b/
└── weights/
    └── model.safetensors
```

To generate a navigation sequence with a walk-forward and turn-right pattern:

```bash
CUDA_VISIBLE_DEVICES=0 ./run.sh \
  --image /path/to/first_frame.png \
  --prompt-file examples/prompt.json \
  --actions "W:5,right:5,stop:5" \
  --perspective first_person \
  --num-frames 121 \
  --output-dir outputs/demo
```

Edit examples/prompt.json to change the scene description, or pass --video-prompt, --audio-prompt, and --av-prompt flags directly. The output MP4 appears in outputs/demo/release/native/.

## Limitations and Gaps in the Preview v1 Release

The 80 GB VRAM requirement is a hard constraint. This rules out consumer GPUs; the model requires an A100 or H100 class card. The README does not document a quantized or reduced-precision path that would run on smaller GPUs, and there is no mention of multi-GPU support in the Preview v1 documentation.

Linux is the only documented operating system. The README does not mention Windows or macOS support. The system dependencies on ffmpeg and ffprobe add another installation step beyond the conda environment.

The README explicitly states that the full model, training code, and technical report are forthcoming. Preview v1 is an inference-only release with a single checkpoint. Researchers who want to fine-tune the model on their own data, study the training objective, or reproduce the results from a paper cannot do so from this release alone.

The code licence is Apache 2.0. The weight licence is published separately with the Hugging Face checkpoint and should be reviewed before commercial use, as research model licences often restrict commercial deployment even when the code is permissively licensed.

## Comparison with Sora and Video Generation Models

Sora (OpenAI's video generation model, released publicly in 2024) generates video from text or image prompts. It produces high-quality video clips without requiring user navigation inputs during generation.

The functional gap is interactivity. Sora generates a fixed sequence; the model decides camera movement and scene evolution. HelixWorld takes a camera action string and generates what the camera would see at each step as the user navigates. This makes it a world model in the interactive sense rather than a video generator.

The second gap is audio. Sora does not generate spatial audio that follows the camera. HelixWorld's design couples audio and video output, with the spatial sound field updating as the camera turns.

The practical access gap runs in the opposite direction. Sora is available through a consumer subscription and runs on OpenAI's infrastructure without local GPU requirements. HelixWorld requires approximately 80 GB of VRAM on a Linux machine and is a research-grade tool. The trade-off is that HelixWorld is open-source with a downloadable checkpoint, while Sora operates only through OpenAI's API.

## Licence, Maintenance, and Citation

The code is Apache 2.0 licensed. The model weights are published on Hugging Face at NoizAI/HelixWorld-preview under a separate licence stated on the Hugging Face page. Before deploying the weights commercially, review that licence specifically.

The repository was last pushed on 2026-09-03, the same date as the Preview v1 announcement. There are no GitHub releases and no changelog file. The repository includes a BibTeX citation entry in the README for academic use, citing the project as a 2026 misc entry from Noiz AI with a note that the technical report and training code are forthcoming.

The NOTICE.md file in the repository indicates third-party components are acknowledged. The pyproject.toml configures ruff for linting with Python 3.11 as the target version and a line length of 120. The requirements.txt pins specific versions of all major dependencies including torch 2.9.1, torchaudio 2.9.1, and transformers 4.57.6.

## Conclusion

HelixWorld is aimed at researchers working on interactive world models and teams exploring spatial audio-visual generation. A developer who wants to generate a short navigation clip will need a machine with roughly 80 GB of GPU memory, which limits the audience to labs or cloud environments with high-end A100 or H100 instances. Before running, verify that your NVIDIA driver supports CUDA 12.x, that ffmpeg and ffprobe are installed system-wide, and that conda is available to create the isolated Python 3.11 environment the README requires. The full model, training code, and technical report are listed as forthcoming as of the 2026-09-03 release.

## FAQ

### How many frames can HelixWorld Preview v1 generate in one run?

The README example uses 121 frames, which corresponds to 15 latent transitions. The num-frames parameter controls the output length. At 15 latent transitions for 121 frames, each transition spans roughly 8 frames.

### What GPU hardware is needed to run HelixWorld?

The README states that a BF16 single-GPU run uses approximately 70 GB of VRAM and recommends a GPU with 80 GB. This requires an NVIDIA A100 80 GB or H100 80 GB class card. CUDA 12.x and a Linux operating system are also required.

### Is the full HelixWorld model and training code available for download?

Preview v1 releases the inference code and a preview checkpoint. The README states that the full model, training code, and technical report are forthcoming. As of the 2026-09-03 release, no training code is available in the repository.

## Sources

- [Issues](https://github.com/NoizAI/HelixWorld/issues)
- [License: Apache-2.0](https://github.com/NoizAI/HelixWorld/blob/main/LICENSE)
- [NoizAI/HelixWorld on GitHub](https://github.com/NoizAI/HelixWorld)
- [Project website](https://helixworld.org)
- [README](https://github.com/NoizAI/HelixWorld/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/noizai-helixworld
