Stable Virtual Camera: a 1.3B diffusion model that invents the view you were about to fly through
Stable Virtual Camera: Generative View Synthesis with Diffusion Models
At a glance
- What is it?
- SEVA takes a handful of photos and target camera poses, then generates new views of the scene with 3D consistency. The repo ships inference code, two demos and a benchmark split, but no training script.
- Who is it for?
- SEVA is worth a look if you need multi-view video from a small number of input photos and can accept a diffusion model rather than a trained reconstruction. The 1.1 checkpoint is the one to use, the Gradio demo is the fastest path to a result, and the benchmark folder gives you the scenes to judge it on.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Activity is slowing. The repository last received commits 7 months ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 9, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What SEVA generates and the audience it was built for
Stable Virtual Camera, shortened to SEVA in the package name and the code, is a generalist diffusion model for novel view synthesis. You give it any number of input views of a scene plus a set of target camera poses, and it renders new images from those poses that hold together as a 3D consistent set. The paper frames the goal differently from most view synthesis work: instead of optimizing one model per scene, a single checkpoint generalizes across scenes, which is why the README calls it a generalist model rather than a per-scene method.
The practical shape of that matters. A per-scene method trains on the images you give it and produces a faithful reconstruction, because the training step is effectively fitting that specific scene. SEVA does no per-scene training. It ships a 1.3B parameter checkpoint that runs at 576P resolution and generates views from arbitrary camera trajectories, including trajectories that swing well outside the input views.
The authorship list explains the technical center of gravity: Stability AI with researchers from the University of Oxford visual geometry group and UC Berkeley BAIR. The associated paper is arXiv 2503.14489, and the README points at a project page, a Hugging Face model card and a hosted Gradio space, so the intended audience spans people who want to try a demo and people who want to run inference locally.
Cloning with submodules and installing the seva package
The installation path is three commands, and the first one matters more than it looks:
git clone --recursive https://github.com/Stability-AI/stable-virtual-camera
cd stable-virtual-camera
pip install -e .The `--recursive` flag is not decoration. The repository tree includes a `third_party/` directory, and the pyproject configuration points pyright at `third_party/dust3r`, so the depth-aware reconstruction code this work builds on is fetched as a submodule. Clone without the flag and the import fails at inference time rather than at install time, which is the least convenient moment to discover it.
Installing in editable mode pulls a dependency set that is worth reading before you commit a GPU to it. `pyproject.toml` names the distribution `seva` at version `0.0.0`, requires Python 3.10 or newer, and pins two versions exactly: `gradio==5.17.0` and `numpy==1.24.4`. The rest float, including `torch`, `diffusers`, `open-clip-torch`, `kornia`, `roma`, `viser`, `tyro`, `fire`, `einops`, `imageio[ffmpeg]`, `opencv-python` and `huggingface-hub`. A project that pins numpy to an old release while leaving torch unpinned is a specific kind of friction, and it is a common one in diffusion repositories that were developed against a fixed CUDA stack.
The README states two environment requirements: Python 3.10 or newer and torch 2.6.0 or newer. Windows users are told to use WSL, with the reason given plainly, flash attention does not yet have native Windows support in the relevant PyTorch issue. A `docs/INSTALL.md` file covers the extra dependencies for the demos and for development.
Authenticating with Hugging Face before the first run
The weights are gated, so authentication comes before anything else. The README is explicit that the code handles the credentials on the first run once you are logged in, and gives the command:
# This will prompt you to enter your Hugging Face credentials.
huggingface-cli loginAfter that you still have to visit the model card on Hugging Face and request access, which the README calls out as a separate step. Two gates, in other words: a local login and a hosted approval. Budget for the approval if you are evaluating this on a deadline, because a rejected or pending request produces a download error rather than a clear permission message.
The two checkpoints both live under the same Hugging Face repository as `model.safetensors` for version 1.0 and `modelv1.1.safetensors` for version 1.1. Keeping both in one place means switching versions does not mean juggling separate repositories or token scopes.
Two demos aimed at two different kinds of user
The repository ships a Gradio interface and a command line tool, and the README is clear about which one to pick. The Gradio demo is described as requiring no expert knowledge and suitable for general users:
python demo_gr.pyThe CLI demo is for power users and academic researchers, and it takes a data path plus whatever else you want to pass:
python demo.py --data_path <data_path> [additional arguments]That split tells you something about the design. The model is the same in both cases, so the GUI is not a simplified model, it is a restricted control surface. Anything you cannot express in the web form, you need the CLI for. Separate guides exist for each: `docs/GR_USAGE.md` and `docs/CLI_USAGE.md`.
Two scripts at the root, `demo.py` and `demo_gr.py`, sit next to a `seva/` package directory and a `benchmark/` directory. If you are evaluating the model rather than reading about it, the fastest path is the Gradio script against a handful of your own photos, then a move to the CLI once you know which parameters you actually want to vary.
Version 1.1 and the foreground detachment problem
The README carries a version table with one row of real engineering content. Version 1.1, released in June 2025, is described as fixing known issues from version 1.0 where foreground objects sometimes detach from the background. Version 1.0, from March 2025, is the initial release. Both are 1.3B parameters at 576P.
Foreground detachment is the failure mode you would expect from a model that hallucinates geometry: a person or a chair renders correctly against the original camera and then separates from the scene as the target viewpoint moves, because nothing in the architecture guarantees that the subject and its background occupy the same reconstructed space. A checkpoint fix reduces how often it happens. It does not promise it is gone, and the README does not quantify the improvement.
Selecting a version is a one-line change in your script. The README shows the pattern:
load_model(..., model_version=1.1)If you find detached subjects in your own scenes, check which checkpoint you loaded before you go looking for a problem in your input images or your camera poses. That single check resolves the most likely cause first.
The benchmark folder, and what the repo does not give you
The `benchmark/` directory and a release tagged `benchmark` cover the evaluation side. The release notes describe 17 splits organized from 10 different academic datasets, spanning single-view, sparse-view and semi-dense-view regimes. The README points researchers there for scenes, splits and the input and target views used in the paper. A second release tagged `assets_demo_cli` provides example scenes so you can drive `demo.py` before supplying your own data. Neither release is a version number, so there is no tagged release to diff against.
Two gaps are documented rather than hidden. The README's Q&A points at issue 27 and issue 42 for the training script and notes that a community contributor opened pull request 51 based on those discussions. In other words, the training code is not officially released. That makes this an inference repository, and anyone planning to fine-tune SEVA on a domain of their own is writing the training loop themselves.
The second gap is rights. The Q&A entry for output licensing points at issue 26 and states that the output follows the same non-commercial license. The repository's own license field reads NOASSERTION, which is what GitHub shows when it cannot classify a license, and the LICENSE file in the tree is what actually governs use. The model weights are additionally gated behind Hugging Face approval, which is a second set of terms applied on top of whatever the file says. If commercial use matters to you, read the LICENSE file and the model card rather than relying on either the metadata field or a summary.
Editorial conclusion
SEVA is worth a look if you need multi-view video from a small number of input photos and can accept a diffusion model rather than a trained reconstruction. The 1.1 checkpoint is the one to use, the Gradio demo is the fastest path to a result, and the benchmark folder gives you the scenes to judge it on. What you do not get is the training script, so this is an inference repository, not something you can fine-tune without writing that part yourself. Check the non-commercial output terms before you build anything commercial on top of it, and note that the last push landed on 2026-03-03.
Frequently asked questions
Do I need a specific GPU to run SEVA inference?
The README does not state a GPU requirement, which is worth noting given the model is 1.3B parameters at 576P resolution. It does state Python 3.10 or newer and torch 2.6.0 or newer, and Windows users must use WSL because flash attention has no native Windows support yet. Budget for a CUDA machine with enough memory for the checkpoint and check the torch version before you start.
Why did I get an error right after cloning the repository?
The install command starts with `git clone --recursive`, which pulls the submodules under `third_party/`. The pyproject configuration points pyright at `third_party/dust3r`, so a plain clone without submodules leaves the depth-aware reconstruction code missing. Re-clone with the recursive flag rather than trying to repair the existing checkout.
Can I train or fine-tune SEVA myself?
Not from this repository as it stands. The README's Q&A points at issues 27 and 42 for the training script and mentions a community pull request 51 built from those discussions, which means the training code is not officially released. You would be writing that part yourself, on top of an inference implementation.
Is SEVA 1.1 better than 1.0?
The README's version table says version 1.1 fixes known issues from version 1.0 where foreground objects sometimes detach from the background. Both checkpoints are 1.3B parameters at 576P. Load the newer one with `load_model(..., model_version=1.1)` unless you have a reason to reproduce the original paper numbers.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/stability-ai-stable-virtual-camera)