# DreamX-Creator ships two sub-repositories and three checkpoint sets

> DreamX-Creator is a research framework for joint audio-video generation: a 7B generator that takes a first frame and a prompt and produces synchronized sound and picture, plus a 5B refiner that upscales the result to 2K. Only those two stages are in the repository, and the weights live outside git entirely.

**AMAP-ML/DreamX-Creator** — Democratizing Native Audio-Video Generation at 2K Resolution

- Repository: https://github.com/AMAP-ML/DreamX-Creator
- Stars: 312 · Forks: 14
- Language: Python
- License: Apache-2.0
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/amap-ml-dreamx-creator

## Two stages are implemented and the rest is described as a system

The repository contains the base generator and the refiner, and the wording around them is careful about that boundary.

What is implemented is a base generator that takes a first frame and a text prompt and jointly models modality-specialized video and audio streams. Two named techniques sit underneath: gated cross-modal attention, and progressive joint training. Their stated purpose is bidirectional audio-video interaction, which is the point of generating both streams in one model rather than generating video and scoring sound afterwards.

The reinforcement-learning stage and the multimodal feedback stage are described as part of the broader system that improves visual and audio quality, semantic consistency, and fine-grained synchronization. They are not among the released components.

The release plan makes the same split legible. Three items are marked done: initializing the repository, releasing the technical report, and releasing validated weights, inference code, and configurations. One item is left open, and it is the distilled, faster models with fewer sampling steps for reduced latency.

The news entries match: the project was initialized on 2026-09-01 and the weights and inference code for the 7B generator and the refiner followed on 2026-09-03.

## Three checkpoint sets are required and one dependency is optional

Weights are not in git. They are published on a model hub and on ModelScope, and the expected on-disk layout is described in a README inside the checkpoints directory.

The first entry is the creator directory for the joint generator, at 7B with LoRA already merged. Inside it sit a video model directory holding the video diffusion transformer shards and their config, an audio model directory holding the audio transformer and config, and a single separate file for the gated audio-to-video and video-to-audio cross-modal attention weights.

The second entry is an audio VAE, named for the codec the audio stream runs through.

The third is the refiner directory, containing the super-resolution transformer checkpoint at 5B and two latent upsamplers: a flash variant marked as the default and a causal two-dimensional one as the alternative.

A fourth directory is available but explicitly not required. The Wan 2.2 text-and-image-to-video 5B weights can be downloaded from their own model repository, and the documentation is clear that only the three entries above are needed.

The word merged in that first entry is worth pausing on. The published generator weights are LoRA-merged, so there is no adapter to attach or detach, and nothing to tune the merge from.

## A 7B generator and a 5B refiner, each self-contained

The repository is two directories with no shared code, and that is stated as a feature of the layout rather than a limitation.

The first holds the 7B native joint audio-video generator, and it is described as running on a single GPU. Its README covers usage, input overrides, and memory options, and the top-level README points at CPU offload and multi-GPU sequence-parallel inference as further topics in the same place.

The second holds the autoregressive 1-step 2K refiner, which is a super-resolution transformer at 5B. Its README is described as carrying the full list of inference knobs.

Each directory has its own requirements file, so the two stages can be installed separately. That matters practically: the refiner does not need the generator's dependencies, and someone who only wants to upscale an external video does not pay for the 7B stack.

The split also explains the sizes involved. The generator produces the first output and the refiner upscales it in one autoregressive step, so a full run touches both a 7B model and a 5B model. The refiner stage takes an input path and leaves the audio unchanged, which is what makes it usable on video the system did not generate.

## The default run is a benchmark case, not your own prompt

The quickstart for the generator is three commands, and the third one carries the detail.

```bash
cd audio_video_generation
pip install -r requirements.txt
./inference.sh                        # runs the default Verse-Bench case (case1)
```

So an install followed by the script runs a prepared benchmark case rather than a text prompt of your own. The input override mechanism is documented in that subdirectory's README, along with the CPU-offload options and the output details.

Verse-Bench appears here as a named case set, with case one as the default. That is a reasonable default for a research release, since it makes a first run comparable to published results, but it does mean the first thing you see is not your own input.

The refiner's quickstart is the mirror image, and it takes its input through an environment variable rather than a flag.

```bash
cd ../video_refiner
pip install -r requirements.txt
INPUT=/path/to/video.mp4 bash run_inference.sh
```

That difference in style between the two stages is small but real: one is configured by editing or passing overrides to a script, the other by setting a variable before invoking the runner.

## Two latent upsamplers ship and the flash one is the default

The refiner directory contains a choice that the README does not explain in prose, and it is easy to miss.

Both upsampler checkpoints are present. One is a flash variant, marked in the layout as the default. The other is described as a causal two-dimensional latent upsampler, which is the alternative for anyone who needs that behavior.

Which one is right depends on the trade-off the refiner README is said to cover: KV cache, window attention, and speed versus quality. Those three are the knobs, and they are the whole tuning surface for the 2K stage.

The two-dimensional causal upsampler is the more interesting of the two names. A causal constraint means each output position depends only on preceding input, which is what allows a long video to be processed without the whole sequence resident, at the cost of not using future context the way the flash path can.

The refiner is described as autoregressive in one step, meaning the upscaling happens in a single pass rather than through a chain of diffusion steps. That is what keeps the 2K stage cheap enough to be worth running on top of the generator.

## Cross-modal attention is a third file beside the two modality models

The creator checkpoint layout explains how the joint model is split, and the split is not arbitrary.

The video model and the audio model are separate modality-specialized components, each with its own transformer weights and config. Neither one is a unified backbone.

The interaction between them lives in its own file, holding the gated audio-to-video and video-to-audio cross-modal attention weights. Gating means those connections are learned to be selective rather than always on, and the two directions are named separately, which is consistent with the claim of bidirectional interaction rather than one stream conditioning the other.

Progressive joint training is the other half of the mechanism and is described in the overview rather than in the layout, so there is nothing in the checkpoint tree that corresponds to it directly.

For a user, the practical consequence is that the cross-modal file is not optional. A checkout with the video and audio directories but without it has two independent generators and none of the audio-video behavior the project is about. The default inference case would still run and would not produce synchronized sound.

## Three upstream teams are credited and one item is still open

The acknowledgement names three teams and the projects they are responsible for, which is a useful signal about what is reused rather than built.

The Wan Team is credited for work on Wan, the OpenMOSS Team for MOVA, and the VideoX-Fun Team for VideoX-Fun. The Wan relationship is visible elsewhere too, since a Wan 2.2 checkpoint set is offered as an optional download for the video side.

Licensing is Apache 2.0, which is the permissive choice for a research framework whose weights are downloaded separately.

The release history is short and worth noting for expectations. The project directory was initialized on 2026-09-01 and the weights plus inference code for both stages were published on 2026-09-03. There are no tagged releases, and the last recorded push is 2026-09-03.

What is not here is as informative as what is. There is no top-level requirements file, no shared library between the two stages, no training code, and no reinforcement-learning component, so the repository is an inference package rather than a training package despite the framework framing.

## Conclusion

DreamX-Creator suits a research team evaluating native audio-video generation that has a single GPU to spare for the 7B stage and a second for the refiner, and that is willing to assemble checkpoints by hand before anything runs. It does not suit anyone needing low-latency generation today, since the distilled faster models are the one item still open on the release plan. Before you start, read the checkpoints README and confirm the layout matches your download, expect the default run to be a benchmark case rather than your own prompt, and note that the shipped generator weights are merged rather than adapter-based.

## FAQ

### How do I run DreamX-Creator joint audio-video generation?

Change into `audio_video_generation`, install its requirements, and run the inference script, which executes the default Verse-Bench case. Input overrides, CPU-offload options, output details, and multi-GPU sequence-parallel inference are documented in that subdirectory's README.

### Where do I download the DreamX-Creator model weights?

From HuggingFace or ModelScope under the GD-ML listing, then place them under `checkpoints/`. The weights are not in the git repository, and a README inside `checkpoints/` documents the expected layout and download instructions.

### Which checkpoints does DreamX-Creator need?

Three: the creator joint generator at 7B with LoRA merged, containing video and audio model folders plus a separate file of gated cross-modal attention weights; the CreatorDACVAE audio VAE; and the 2K refiner with its super-resolution transformer and a latent upsampler, the flash variant being the default. A Wan 2.2 checkpoint set is available but not required.

### How do I super-resolve a video to 2K with DreamX-Creator?

Change into `video_refiner`, install its requirements, and run the inference script with an input path pointing at a video file. The audio is left unchanged, and the subdirectory README lists the knobs, including KV cache and window attention for speed and quality trade-offs.

### What is still missing from DreamX-Creator 1.0?

Distilled, faster models with fewer sampling steps for reduced latency, which is the single unchecked item on the release plan. The released components are the 7B joint generator and the autoregressive 1-step 2K refiner, with the reinforcement learning and multimodal feedback stages described as part of the broader system rather than shipped.

## Sources

- [AMAP-ML/DreamX-Creator on GitHub](https://github.com/AMAP-ML/DreamX-Creator)
- [Issues](https://github.com/AMAP-ML/DreamX-Creator/issues)
- [License: Apache-2.0](https://github.com/AMAP-ML/DreamX-Creator/blob/main/LICENSE)
- [README](https://github.com/AMAP-ML/DreamX-Creator/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/amap-ml-dreamx-creator
