# Bernini: ByteDance's Unified Framework for Video Generation and Editing

> Bernini is a video generation and editing framework from ByteDance that pairs an MLLM-based semantic planner with a DiT-based renderer, covering text-to-video, image-to-video, video-to-video, and editing tasks through a single interface. It targets researchers and practitioners with H100-class GPU access who need strong instruction following in complex video editing workflows.

**bytedance/Bernini** — Bernini is a unified framework for video generation and editing that combines an MLLM-based semantic planner with a DiT-based renderer.

- Repository: https://github.com/bytedance/Bernini
- Website: https://bernini-ai.github.io/
- Stars: 1,322 · Forks: 103
- Language: Python
- License: Apache-2.0
- Published: 2026-09-18 · Updated: 2026-09-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/bytedance-bernini

## Two Model Families: Bernini and Bernini-R

The repository provides two distinct model families with different complexity tradeoffs. Bernini is the full pipeline: an MLLM-based semantic planner decomposes complex natural-language instructions into semantic changes, which a DiT-based renderer then executes. The README describes this decomposition as Bernini's core strength: it handles multi-step or ambiguous edit instructions by planning before rendering, producing stronger instruction following.

Bernini-R is a renderer-only model, fine-tuned from Wan2.1. It skips the semantic planning stage and passes instructions directly to the diffusion renderer. The tradeoff is a simpler setup with fewer moving parts. According to the benchmark table in the README, the 14B Bernini-R model performs close to the full Bernini pipeline on simpler tasks such as style transfer, subtitle removal, and local editing, but lags on complex scenarios like human generation.

A 1.3B Bernini-R variant is also available. The README notes it performs close to the 14B variant on simpler tasks, making it a lower-cost option for experiments that do not require the full model capacity.

Both families share the same task interface: t2i (text-to-image), i2i (image-to-image), t2v (text-to-video), v2v (video-to-video), rv2v (reference video-to-video), and r2v (reference-to-video). Checkpoints are hosted on Hugging Face: ByteDance/Bernini-Diffusers for the full pipeline and ByteDance/Bernini-R-Diffusers for the renderer.

## Semantic Planner and DiT Renderer Architecture

Bernini's architecture combines two components that communicate through a latent semantic representation rather than raw pixel data. The MLLM-based planner reads the source video and the edit instruction, then produces a semantic plan that describes what should change and where. The DiT-based renderer takes that plan and generates the output video frames through a diffusion process.

This two-stage design is what the paper, published on arXiv as 2605.22344, titles Latent Semantic Planning for Video Diffusion. The planning stage acts as an interpreter for complex instructions before any generation occurs.

Bernini-R removes the planner and routes the instruction directly to the DiT renderer. This is the same diffusion approach used in the Wan family of models, from which Bernini-R is fine-tuned. The README recommends Bernini-R for workflows where instruction clarity is high and the editing task is well-defined, citing simpler setup as the practical benefit.

A Gradio demo is included as gradio_demo.py. Training code for Bernini-R was released on 2026-07-13, with the full training guide in docs/bernini_r_train.md. The full Bernini inference code and model weights were released on 2026-06-11.

## Installation: Pinned Dependencies and the VeOmni Requirement

Bernini has strict environment requirements. The README specifies Python 3.11.2 (not any 3.11.x), CUDA 12.6, and a pinned PyTorch build of torch==2.7.1+cu126. An NVIDIA H100 is the reference GPU; other CUDA GPUs fall back from FlashAttention-3 to FlashAttention-2 or PyTorch SDPA.

The inference installation sequence starts with cloning and installing requirements:

```bash
git clone https://github.com/bytedance/Bernini.git bernini && cd bernini
pip install -r requirements.txt
pip install --no-deps git+https://github.com/ByteDance-Seed/VeOmni.git@v0.1.11
```

The VeOmni step uses --no-deps to prevent it from pulling in a different torch build that would override the pinned torch==2.7.1+cu126. The README states that Open-VeOmni is a required dependency for all inference paths, including single-GPU.

For training, the README recommends uv for environment management. The pyproject.toml declares all training dependencies and routes torch to the correct CUDA index automatically:

```bash
curl -LsSf https://astral.sh/uv/install.sh | sh
uv sync
uv sync --extra all
uv pip install --no-build-isolation "flash-attn==2.8.3"
```

FlashAttention-3 for Hopper GPUs (H100/H800/H200) is not on PyPI and must be built from source from the flash-attention repository's hopper/ directory at tag v2.8.3. This is an optional step; CUDA GPUs without Hopper architecture use FlashAttention-2 instead.

## Case Files and the Shared Task Interface

A run in Bernini is described by a case file, a small JSON document stored under assets/testcases/. Each case file specifies the task type (one of t2i, i2i, t2v, v2v, rv2v, r2v), the source media, the edit instruction, and any model-specific parameters. Both Bernini and Bernini-R use this same format, which means case files are portable across the two pipelines.

The inference entry points differ by GPU count. For single-GPU inference the repository provides infer_single_gpu.py, and for multi-GPU runs (using Ulysses sequence parallelism) there is infer_multi_gpu.py. The README links to model-specific guides in docs/bernini.md and docs/bernini_r.md for weight download and per-task inference commands.

The repository also notes support from vLLM-Omni as of 2026-06-03, with a recipe available at the vllm-omni repository. This adds a serving path that the repository itself does not provide directly.

## Benchmark Scope and What the Numbers Cover

The README includes a benchmark table covering four metrics: EditVerse, OpenVE, OpenS2V, and VBench, plus two Bernini-specific scores (Bernini-v2v and Bernini-rv2v). These figures come from the project's self-built arena platform, where human annotators blindly vote on paired edits and the votes are aggregated into Bradley-Terry scores.

The full Bernini pipeline at 7+14B scores 8.02 on EditVerse and 4.03 on OpenVE, while Bernini-R at 14B scores 7.99 and 3.78 on the same metrics. The 1.3B Bernini-R scores 7.74 and 3.65. The README does not claim these figures represent production performance on arbitrary prompts; they are the result of a specific evaluation protocol.

The README also states that the released checkpoint was trained with up to 128 A100 GPUs, with training conducted up to 768x768 image generation and 480p at 12 FPS video generation. Output quality is noted to vary across prompts, resolutions, duration, motion complexity, and editing scenarios, and the README explicitly describes the project as a research artifact rather than a polished product model.

## Limitations, Alternative Tools, and Licensing

The hardware requirement is the sharpest practical constraint. H100 or H800 or H200 GPUs are recommended for FlashAttention-3; the README notes other CUDA GPUs fall back to FlashAttention-2 or SDPA, which works but runs slower. The strict Python 3.11.2 and CUDA 12.6 requirements mean existing virtual environments are unlikely to be compatible without rebuilding.

There is no serving infrastructure included. The repository provides inference scripts and a Gradio demo but not a REST API, a Docker image for deployment, or a model server compatible with production traffic. Teams needing video editing as a service would need to build that layer themselves.

The closest comparable open project in the text-to-video space is CogVideoX from THUDM, which also targets video generation and editing but uses a different architecture (3D causal VAE with expert transformer) and does not include a semantic planning stage. Bernini's distinction is the MLLM-based planner for complex multi-step instructions.

The repository is licensed under Apache-2.0. The last push to the repository was on 2026-08-13. Contributions and bug reports are noted as welcome in the README.

## Conclusion

Bernini is worth evaluating for research teams working on video editing benchmarks or instruction-following pipelines who can supply H100-class hardware and the exact Python 3.11.2 and CUDA 12.6 environment. The 1.3B Bernini-R variant lowers the hardware bar and performs comparably on simpler tasks like style transfer and watermark removal. Teams seeking a production API or a polished deployment path should look elsewhere; the repository provides inference scripts and a Gradio demo, but no serving infrastructure. Before running inference, confirm that the VeOmni dependency installs cleanly with --no-deps alongside the pinned torch==2.7.1+cu126 build.

## FAQ

### What is Bernini AI from ByteDance?

Bernini is a unified framework for video generation and editing that combines an MLLM-based semantic planner with a DiT-based renderer. It supports tasks including text-to-video, video-to-video editing, and image-to-video generation, and is released under Apache-2.0 with weights on Hugging Face.

### What is Bernini-R and how does it differ from the full Bernini pipeline?

Bernini-R is a renderer-only model fine-tuned from the Wan2.1 diffusion renderer. It skips the MLLM-based semantic planner of the full Bernini pipeline, resulting in a simpler setup. The README notes Bernini-R performs comparably to the full pipeline on simpler tasks like style transfer and watermark removal, but lags on complex tasks such as human generation.

### Can Bernini run on GPUs other than the H100?

The README states that a Hopper GPU (H100/H800/H200) is recommended so FlashAttention-3 can be used. Other CUDA GPUs fall back to FlashAttention-2 or PyTorch SDPA. FlashAttention-3 is not on PyPI and must be built from source from the hopper/ directory of the flash-attention repository at tag v2.8.3.

## Sources

- [bytedance/Bernini on GitHub](https://github.com/bytedance/Bernini)
- [Issues](https://github.com/bytedance/Bernini/issues)
- [License: Apache-2.0](https://github.com/bytedance/Bernini/blob/main/LICENSE)
- [Project website](https://bernini-ai.github.io/)
- [README](https://github.com/bytedance/Bernini/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/bytedance-bernini
