# JoyAI-Echo: Long-Horizon Audio-Visual Generation from JD.com's Research Team

> JoyAI-Echo is a research repository from JD.com holding two independent projects: Echo-LongVideo for multi-shot video generation beyond ten minutes, and Echo-WM, an omnimodal world model that lets navigation control video, sound and speech together. Both are for academic and non-commercial use only.

**jd-opensource/JoyAI-Echo** — JoyAI-Echo-1.5: Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds

- Repository: https://github.com/jd-opensource/JoyAI-Echo
- Website: https://echo-team-joy-future-academy-jd.github.io/Echo-1.5-Page/
- Stars: 2,041 · Forks: 169
- Language: Python
- License: NOASSERTION
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/jd-opensource-joyai-echo

## Two Independent Projects Under One Repository

JoyAI-Echo is organized as two self-contained projects in separate directories:

```text
JoyAI-Echo/
├── echo_longvideo/   # long-video generation: inference.py, configs/, prompts/, ltx-*
└── echo_wm/          # world model: inference_wm.py, Gradio demo, bundled ltx-*
```

Echo-LongVideo handles multi-shot audio-visual generation with stated support for sequences longer than ten minutes, using a paired audio-video memory bank to carry continuity across shots. Echo-WM is described as an omnimodal world model for generative media that responds to continuous navigation while video, environmental sound, music and speech evolve together.

The two projects do not share a Python environment or a checkpoint directory. The `echo_wm/` directory bundles its own copy of `ltx-core` and `ltx-pipelines`, so installing one project does not affect the other. This separation is intentional but it also means researchers working with both projects must maintain two separate environments.

## Release History and the LTX Backbone

The repository documents four milestones. JoyAI-Echo 1.0 was released on 2026-06-22 and is preserved on the `echo1.0` archive branch. Echo-WM was released on 2026-08-26. JoyAI-Echo 1.5 (Echo-LongVideo) followed on 2026-08-28, including long-horizon generation, consumer-GPU inference profiles and a Director Agent. On 2026-09-04, a UE simulation pipeline for the Echo-WM world data engine was released, covering physics-based trajectory generation, Movie Render Queue rendering and distributed scheduling.

Both projects are built on LTX-2 by Lightricks Ltd. Echo-WM is currently on LTX-2.3. The roadmap lists planned upgrades to LTX-2.5 for both the Base and Causal variants, along with sparse attention work using SageAttention and similar kernels, a paged KV-cache for bounded rollout memory, and FP8 or TensorRT compilation for throughput improvements. These items are marked as not yet complete.

## Setting Up Echo-LongVideo and Echo-WM

Both projects require separate environments. For Echo-LongVideo:

```bash
cd echo_longvideo
conda env create -f environment.yml && conda activate echo-long
```

For Echo-WM:

```bash
cd echo_wm
conda create -n echo-wm python=3.11 -y && conda activate echo-wm
pip install -r requirements.txt
```

After environment setup, checkpoints must be downloaded separately in both cases. The README does not list the exact checkpoint filenames in the top-level document; each subdirectory's own README specifies the files and paths. A ComfyUI integration is available as a separate repository at github.com/zhuang2002/ComfyUI_JoyAI_Echo for users who prefer a node-based interface. Model weights for Echo-LongVideo are hosted on Hugging Face under jdopensource/JoyAI-Echo, and Echo-WM weights are at Echo-Team/Echo-WM.

## Echo-WM: Causal Mode and the Rollout Mechanism

Echo-WM's roadmap describes two backbone variants under LTX-2.3. The Base variant is a bidirectional audio-visual DiT producing approximately ten-second segments. The Flash Preview and Causal variant uses chunk-causal attention, a KV-cache rollout and four-step inference; this is the current public preview and its details are in `echo_wm/README_CAUSAL.md`.

The planned LTX-2.5 work includes loading the newer Lightricks weights with Gemma 4 TE and a 2.5 VAE and DiT into the existing bidirectional path, followed by adapting the causal recipe to the 2.5 architecture. The acceleration roadmap lists SageAttention for sparse and low-bit kernels across video, audio and UCPE branches, FlashAttention for long causal windows without blowing up HBM, a paged KV-cache for variable-length cache (with rebase of RoPE and UCPE when tokens evict), and FP8 or TensorRT compilation. All acceleration items are listed as future work and are not currently implemented.

The UCPE branch in the roadmap is specific to Echo-WM's architecture and is not a standard term in the broader literature; the README does not define the abbreviation in the top-level document. Researchers who want to contribute to the acceleration work should consult the echo_wm subdirectory documentation and the causal README for the precise tensor flow before attempting to integrate new attention kernels.

## Licence Constraints and the Non-Commercial Boundary

The project derives from LTX-2 by Lightricks Ltd and operates under the LTX-2 Community License Agreement. The README states explicitly: this project is not intended for commercial use. For commercial use of LTX-2 or its derivatives, contact Lightricks directly. All original copyright, licence, patent, trademark and attribution notices from LTX-2 are retained. A THIRD_PARTY_NOTICES.md file in the repository records the specific upstream dependencies and their terms.

This is a meaningful restriction for anyone considering productizing the output. Academic researchers and non-commercial experimenters are within scope. Companies that want to incorporate the technology into a product need to go through Lightricks regardless of what modifications JD.com applied. The licence restriction applies to the model weights and the code that processes them; the BibTeX citation entries in the README suggest the primary intended users are researchers who would reference the work in publications. The repository has no GitHub releases; access to weights is through Hugging Face, not through GitHub release assets, which means version tracking relies on Hugging Face repository history rather than on Git tags.

## Limitations and Cases Where JoyAI-Echo Is the Wrong Choice

The repository does not document minimum GPU memory requirements for either project in the top-level README. The main README refers users to each subdirectory for the exact checkpoint files and paths, which means a researcher cannot determine from the top-level page alone whether their hardware meets the requirements. The echo_longvideo/README and echo_wm/README hold the specifics, so reading both subdirectory documents before any environment setup is the practical first step.

The non-commercial restriction excludes all production deployment scenarios. The two projects cannot share a Python environment, which adds overhead for researchers exploring both. Both projects require conda; the Echo-LongVideo setup uses `environment.yml` while Echo-WM uses `requirements.txt` with a manually created conda environment, so the setup procedures differ.

Another constraint is that Echo-LongVideo's Director Agent and consumer-GPU inference profiles are new as of the 2026-08-28 release, and the README describes them without giving specific GPU model recommendations or VRAM figures. Researchers who want to run long-horizon generation on a laptop or modest workstation will need to consult the subdirectory documentation to confirm what is feasible.

A direct alternative for long-form video generation is Wan2.1 by Alibaba, which is also based on diffusion transformers and targets research applications. The difference in approach is that JoyAI-Echo explicitly pairs audio with video across long horizons using a memory bank, while most open video generation models produce video without synchronized environmental sound and speech. Wan2.1 does not include a world model component analogous to Echo-WM.

## Citing the Work and Academic Context

The repository cites three papers. The JoyAI-Echo 1.5 paper is available as arXiv preprint 2608.23383, authored by Duan, Huang, Jin and others at JD.com, with the full author list including Xiaoxiao Ma, Yanwen Ma, Yaofeng Su, Yilang Sun, Haoyu Wang, Zeyue Xue, Songchun Zhang and Junhao Zhuang. A separate Echo-WM paper is at arXiv 2608.23189. The UE simulation pipeline paper is at arXiv 2609.03557. The 1.0 paper was published through ResearchGate.

The project page for Echo 1.5 is hosted at echo-team-joy-future-academy-jd.github.io/Echo-1.5-Page/. The README provides a BibTeX entry for the 1.5 paper under the key `duan2026joyaiecho15`. The repository was last pushed on 2026-09-20. The repository has no GitHub releases; all weights are distributed through Hugging Face rather than through GitHub release assets.

The UE simulation pipeline, released on 2026-09-04, covers physics-based trajectory generation, Movie Render Queue rendering and distributed scheduling. It is documented as a separate paper at arXiv 2609.03557 and is intended to support the Echo-WM world data engine with synthetic training data. This makes JoyAI-Echo more than a simple inference repository: the trajectory generation pipeline provides a path to producing training data at scale for the world model, which is an important distinction from repositories that only supply inference code.

## Conclusion

JoyAI-Echo is suited to computer vision and audio-visual generation researchers who want to run long-horizon video experiments or explore world-model generation at the intersection of video, environmental sound, music and speech. The non-commercial restriction is a firm boundary: the project is based on Lightricks' LTX-2 under the LTX-2 Community License Agreement, and commercial use requires contacting Lightricks directly. Before starting, confirm that your hardware can run either project: Echo-LongVideo targets consumer GPUs with specific inference profiles, while Echo-WM's causal mode requires a reasonably capable GPU for chunk-causal attention and KV-cache rollout. Checkpoints must be downloaded separately in both cases, and the two projects use different Python environments that must not be mixed.

## FAQ

### What is JoyAI-Echo?

JoyAI-Echo is a research repository from JD.com containing two projects: Echo-LongVideo, which generates multi-shot audio-visual sequences longer than ten minutes, and Echo-WM, an omnimodal world model where navigation continuously controls video, environmental sound, music and speech together. Both are for non-commercial use only.

### Can JoyAI-Echo be used for commercial products?

No. The project is based on LTX-2 by Lightricks Ltd and is explicitly for academic and non-commercial use only. The README states that commercial use of LTX-2 or its derivatives requires contacting Lightricks directly.

### Do Echo-LongVideo and Echo-WM share a Python environment?

No. The two projects use different Python environments and do not share a checkpoint directory. The echo_wm/ directory bundles its own copy of ltx-core and ltx-pipelines, so installing one project does not affect the other.

### Where are the model weights for JoyAI-Echo hosted?

Echo-LongVideo weights are on Hugging Face at jdopensource/JoyAI-Echo, and Echo-WM weights are at Echo-Team/Echo-WM. Checkpoints must be downloaded separately after setting up the relevant environment.

## Sources

- [Issues](https://github.com/jd-opensource/JoyAI-Echo/issues)
- [jd-opensource/JoyAI-Echo on GitHub](https://github.com/jd-opensource/JoyAI-Echo)
- [Project website](https://echo-team-joy-future-academy-jd.github.io/Echo-1.5-Page/)
- [README](https://github.com/jd-opensource/JoyAI-Echo/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/jd-opensource-joyai-echo
