Model or dataset
bytedance/Lance avatar
bytedance/Lance

Lance: ByteDance's 3B Unified Multimodal Model for Image and Video Tasks

A 3B-active-parameter native unified multimodal model for image and video understanding, generation, and editing.

1,349 stars93 forksPythonApache-2.0

At a glance

What is it?
Lance is a 3B active-parameter multimodal model from ByteDance that handles image understanding, image generation, image editing, video understanding, and video generation within a single architecture, trained from scratch on up to 128 A100 GPUs. It is a research artifact aimed at teams studying unified multimodal modeling under a constrained compute budget.
Who is it for?
Lance is a reasonable starting point for researchers studying unified image and video modeling at a relatively small scale, particularly those who want a single model covering understanding, generation, and editing tasks. The 40GB VRAM floor, the Python 3.10 plus CUDA 12.4 requirements, and the explicitly research-grade quality mean it is not a production model for user-facing applications.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 77 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Lance Is and the Problem It Studies

Most video and image AI models specialize in one task: a text-to-image model, a separate video editing model, a separate visual question-answering model. Each requires its own weights, its own fine-tuning pipeline, and its own serving infrastructure. Lance takes a different approach: a single 3B-parameter model trained to handle understanding, generation, and editing for both images and videos.

The README is direct about the motivation and its limits: the goal is to share a research artifact for studying unified image and video understanding, generation, and editing under a relatively small model and limited compute budget. ByteDance trained it from scratch using a staged multi-task recipe on up to 128 A100 GPUs, which is large by most standards but small compared to frontier production models.

The project targets researchers studying multi-task synergy, the hypothesis that training one model on multiple related tasks improves performance on each individual task relative to training separate specialized models. The README does not claim state-of-the-art output quality on all tasks; it notes output quality may vary across prompts, resolutions, duration, motion complexity, and editing scenarios.

Multi-Task Synergy: How Lance Is Trained

Lance is not a fine-tuned version of an existing model. The README states it is trained from scratch with a staged multi-task recipe. The training covered image generation up to 768x768 pixels and video generation at 480p and 12 FPS.

The staged approach means different task combinations are introduced at different phases of training, allowing the model to first develop core visual representations before being asked to handle the full range of generation and editing tasks simultaneously. The README titles this approach Multi-Task Synergy.

Fine-tuning code was released on 2026-06-17. The training guide is in train.md, with additional dataset documentation in train_dataset.md and train_dataset_zh.md. The train/ directory contains the training implementation. The README also lists a Hugging Face Space at huggingface.co/spaces/bytedance-research/Lance for trying the model without local installation.

Installation and Running Inference

The minimum hardware requirement is a GPU with at least 40GB VRAM. The README lists two tested dependency combinations on NVIDIA A100: PyTorch 2.8.0 with CUDA 12.6 and flash-attn 2.8.3, or PyTorch 2.5.1 with CUDA 12.4 and flash-attn 2.6.3.

The recommended setup uses the PyTorch 2.8.0 and CUDA 12.6 combination. The requirements.txt pins specific versions including transformers==4.49.0, diffusers==0.29.1, and accelerate==1.13.0:

bash
git clone https://github.com/bytedance/Lance.git
cd Lance
pip install -r requirements.txt

Inference runs through inference_lance.py or the shell wrapper inference_lance.sh. The README notes that for other GPU models, users should choose and validate a PyTorch build and a matching flash-attn version according to their driver, CUDA runtime, and Python version. This is not a one-command install; dependency compatibility requires careful validation.

A Gradio interface is available via lance_gradio.py and supports image generation, video generation, editing, and understanding tasks. The README notes the Gradio interface was updated on 2026-05-26.

Supported Tasks and the vLLM-Omni Integration

Lance handles six categories of input-output combinations. Text-to-image (t2i) generates images from text prompts. Image-to-image (i2i) transforms an input image according to instructions. Text-to-video (t2v) generates video from text. Image-to-video (i2v) animates a still image. Video editing transforms an existing video based on a natural-language instruction. Video understanding answers questions about video content.

As of 2026-06-03, Lance is supported in vLLM-Omni, which provides a serving path the repository itself does not include. The recipe for using Lance through vLLM-Omni is available in the vllm-omni repository under recipes/ByteDance/Lance.md. This is particularly relevant for teams that want to serve the model through an existing vLLM-Omni deployment.

The model weights are on Hugging Face at bytedance-research/Lance. The initial release of inference code and weights was on 2026-05-18, with subsequent updates adding image-to-video support on 2026-05-29.

Limitations: Hardware Floor, Research Quality, and No Serving Infrastructure

The 40GB VRAM requirement is a hard floor. Consumer GPUs with 24GB VRAM will not run Lance inference according to the README's stated requirement. This rules out a significant portion of accessible hardware and limits experimentation to data center or high-end workstation GPUs.

The README is explicit that Lance is a research project, not a polished product model. The team acknowledges opportunities to improve the post-training recipe and states that output quality varies across prompts, resolutions, duration, motion complexity, and editing scenarios. Using Lance as a backend for a user-facing application carries meaningful quality risk.

The repository provides no serving infrastructure. There is no Docker image, no API server, and no deployment guide beyond the inference scripts and Gradio demo. Teams that want Lance as a production service need to build that layer on top of vLLM-Omni or a custom serving stack.

The alternative for teams needing production-ready unified multimodal modeling is InternVL, which offers a comparable range of understanding and generation tasks with more documented deployment options. InternVL's approach differs in that it is built on a separate language and vision encoder architecture rather than a unified generative model trained from scratch.

Repository Layout and What the Directories Contain

The repository is organized to keep inference and training concerns separate. The modeling/ directory holds the model architecture code. The config/ directory contains training and inference configuration files. The inference entry points at the repository root (inference_lance.py and inference_lance.sh) provide ready-to-run scripts without requiring navigation into subdirectories.

The train/ directory contains the training implementation, with a separate top-level train.md as the guide and train_dataset.md for dataset format documentation. Both are also available in Simplified Chinese (train_zh.md and train_dataset_zh.md).

The benchmarks/ directory stores evaluation scripts and results. The common/ directory likely holds shared utilities used by both inference and training paths, though the README does not describe its contents directly.

The repository root contains both README.md and README_zh.md, meaning the team maintains documentation in both English and Simplified Chinese. A SECURITY.md is present. The .devin/ directory contains agent-specific configuration, consistent with the README's note that the project has been used with the Devin AI coding agent in the vLLM-Omni integration context.

The skills-lock.json file at the repository root is the same artifact produced by Claude Code skill management tooling, suggesting the project uses Claude Code for development.

License, Maintenance, and Contributing

Lance is released under Apache-2.0, which permits use, modification, and distribution including in commercial applications, provided attribution is maintained. The last push to the repository was on 2026-07-14.

The README invites contributions: it notes the team is actively updating and improving the repository and asks contributors to open an issue or submit a pull request for bugs and suggestions. A SECURITY.md is present in the repository. The README is also available in Simplified Chinese as README_zh.md, with Chinese-language training guides (train_zh.md, train_dataset_zh.md) indicating the primary development team writes documentation in both languages.

The common/ directory at the repository root contains shared utilities for both inference and training. The data/ directory provides data processing scripts. The scripts/ directory holds convenience scripts for dataset preparation and evaluation. The benchmarks/ directory contains evaluation code and benchmark configurations. Understanding this layout is useful when adapting the inference code to a custom serving setup, since the modeling/ directory contains the architecture definitions that any serving wrapper would need to import.

Editorial conclusion

Lance is a reasonable starting point for researchers studying unified image and video modeling at a relatively small scale, particularly those who want a single model covering understanding, generation, and editing tasks. The 40GB VRAM floor, the Python 3.10 plus CUDA 12.4 requirements, and the explicitly research-grade quality mean it is not a production model for user-facing applications. Teams needing reliable, polished output on arbitrary prompts should look at specialized production models. Before running inference, verify that your CUDA driver supports CUDA 12.4 or newer and that your GPU has at least 40GB VRAM; the README states this is required for inference.

Frequently asked questions

What GPU is required to run Lance inference?

The README states a GPU with at least 40GB VRAM is required. The reference hardware is the NVIDIA A100. The tested configurations use either PyTorch 2.8.0 with CUDA 12.6 and flash-attn 2.8.3, or PyTorch 2.5.1 with CUDA 12.4 and flash-attn 2.6.3.

Can Lance generate video from an image?

Yes. The README added image-to-video (i2v) support on 2026-05-29. Lance supports six task types in total: t2i, i2i, t2v, i2v, video editing, and video understanding, all within the same model and inference pipeline.

Is Lance fine-tunable for custom tasks?

Fine-tuning code was released on 2026-06-17. The training guide is in train.md in the repository, with dataset documentation in train_dataset.md. The README recommends reviewing the training guide before attempting fine-tuning, as the staged multi-task training recipe has specific dependencies and dataset format requirements.

Official sources

  1. bytedance/Lance on GitHub
  2. Issues
  3. License: Apache-2.0
  4. Project website
  5. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/bytedance-lance.svg)](https://hysenlabs.com/projects/bytedance-lance)