PersonaLive: streaming portrait animation, and what the repository still leaves open
[CVPR 2026] PersonaLive! : Expressive Portrait Image Animation for Live Streaming
At a glance
- What is it?
- A CVPR 2026 portrait animation framework built around streaming rather than fixed-length clips, with a research-only licence, a three-stage training script set, and a TensorRT path that trades a little output quality for roughly double the speed.
- Who is it for?
- PersonaLive is a good fit when the output has to run longer than a diffusion model naturally wants to, because streaming is designed into the loop rather than bolted on afterwards. What the repository settles is the runtime story: the conda setup, the exact weight directory layout, the inference flags, the xFormers workaround for Blackwell cards, and the TensorRT conversion path.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 40 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 20, 2026, and from our analysis. They are not legal advice.
Editorial analysis
A framework built for streaming rather than for fixed-length clips
PersonaLive is described as a real time and streamable diffusion framework capable of generating infinite length portrait animations. That combination of words is the whole design brief. Most talking head work produces a clip of a set length, because a diffusion model conditioned on a sequence has a natural span it was trained on, and everything past that span degrades. Building the loop around streaming means the model keeps a running state and consumes frames as they arrive.
The project is titled expressive portrait image animation for live streaming and was accepted by CVPR 2026. The arXiv identifier is 2512.11253. Five authors are credited: Zhiyuan Li, Chi-Man Pun, Chen Fang, Jue Wang and Xiaodong Cun, spread across the University of Macau, Dzine.ai and GVC Lab at Great Bay University. The repository sits under GVCLab with 3789 stars and 542 forks, and 35 open issues, which for a codebase of this kind is a modest number.
The framing as live streaming rather than as offline generation is why the flagship capability is described as infinite length output. It also explains why the streaming option lives inside the offline inference script rather than in a separate tool: the same loop, asked to keep going.
Two dates frame the maturity of the repository. The last push is dated 2026-08-28, which means the code is current rather than archived. But that same day the maintainers marked a todo item complete announcing EditaLive as an updated version of PersonaLive with full body animation and editing abilities. Anyone starting fresh has a reason to look at the successor first, and anyone who needs the specific streaming portrait behaviour still has a reason to stay here.
Two entry points, one offline script and one Node web interface
Installation is a standard conda environment, with Python pinned to 3.10:
# clone this repo
git clone https://github.com/GVCLab/PersonaLive
cd PersonaLive
# Create conda environment
conda create -n personalive python=3.10
conda activate personalive
# Install packages with pip
pip install -r requirements_base.txtThree requirements files sit at the root and separate the use cases cleanly: requirements_base.txt for inference, requirements_train.txt for training, and requirements_trt.txt for the TensorRT path. Base models are fetched with a helper script that pulls sd-image-variations-diffusers and sd-vae-ft-mse:
python tools/download_weights.pyAlternatively the weights can be downloaded by hand from Google Drive, Baidu Netdisk, ModelScope or Hugging Face and placed into the pretrained_weights folder, with the README noting that the Hugging Face repository is huaichang/PersonaLive.
Inference then splits into two modes. Offline inference is a single script:
python inference_offline.pyIts flags are few and well chosen. `-L` caps frame count and defaults to 100. `--use_xformers` toggles memory efficient attention and defaults to true. `--stream_gen` enables the streaming strategy and also defaults to true. `--reference_image` and `--driving_video` override whatever the config file specifies. Since the streaming strategy arrived in late December 2025 specifically to allow long videos within a 12GB VRAM budget, a machine with a consumer card is closer to viable here than the paper's own benchmarks would suggest.
Online inference runs a web interface instead, and it needs Node.js 18 or newer installed through nvm:
# install Node.js 18+
curl -o- https://raw.githubusercontent.com/nvm-sh/nvm/v0.39.1/install.sh | bash
nvm install 18
source web_start.shThe web interface was enhanced in December 2025 to support reference image replacement, which is what makes it usable as a tool rather than a demo. A webcam/ directory at the root, plus a data/ directory and a demo/ folder holding a driving video and a reference image, suggest the intended uses are a browser front end fed by a webcam and a batch path over recorded footage.
The weights directory describes the pipeline more clearly than the prose does
The README prints the expected directory layout, and it is worth reading closely because it is effectively an architecture summary. Under pretrained_weights there are four top level entries.
The personalive/ folder holds the six learned components as separate checkpoint files: denoising_unet.pth, motion_encoder.pth, motion_extractor.pth, pose_guider.pth, reference_unet.pth and temporal_module.pth. Those filenames separate the concerns of the system in a legible way. Motion is extracted from the driving video by one module and encoded by another. A pose guider conditions the denoising on body orientation. A second UNet handles the reference image rather than the frame being generated, which is what lets a still portrait drive an animated one. A temporal module exists specifically to keep successive frames coherent.
The remaining entries are borrowed foundation components: sd-vae-ft-mse for the VAE, and sd-image-variations-diffusers for the image encoder and unet, which is what allows a plain still image to be used as conditioning.
The onnx/ and tensorrt/ folders are the acceleration path rather than part of the model. Under onnx sit unet and unet_opt directories, with unet_opt.onnx and unet_opt.onnx.data as the exported optimised graph. Under tensorrt there is a single engine file, unet_work.engine, which is a compiled artefact rather than something a user regenerates by default.
Reading that layout against a config file in configs/ gives a reasonable picture of the system before running anything: a two UNet pipeline, a dedicated motion path, a temporal module, and a VAE borrowed from Stable Diffusion. The repository does not include a standalone architecture document beyond the framework diagram, so this directory listing does more explanatory work than it appears to.
TensorRT roughly doubles throughput at the cost of some output quality
Acceleration is optional and documented as a separate path rather than built into the base install. The claim is roughly a twofold speedup, with engine build time of about 20 minutes depending on the device, and an explicit warning that TensorRT optimisations may cause slight variations or a small drop in output quality.
That trade is stated plainly, which is refreshing. A framework that generates portraits is often used in contexts where visual artefacts matter, so a small quality drop is not automatically worth accepting for speed.
The conversion procedure has one non-obvious step. A positional embedding in the motion encoder has to be frozen before export, and the README identifies the file to edit directly:
# Install packages with pip
pip install -r requirements_trt.txt
# src/models/motion_encoder/FAN_temporal_feature_extractor.py
self.pos_embed.pos_embed.requires_grad = False
# Converting the model to TensorRT
python torch2trt.pytorch2trt.py at the repository root does the export. The requirements install can fail at the pycuda stage, and the README has a documented workaround for exactly that error: install pycuda through conda with a numpy pin rather than from source.
# Install PyCUDA manually using Conda (avoids compilation issues):
conda install -c conda-forge pycuda "numpy<2.0"The numpy pin below 2.0 is the part that matters, because pycuda has historically not built against newer numpy. The suggested fix is to install pycuda via conda, then comment out or remove the pinned pycuda line from requirements_trt.txt so pip does not try to reinstall it. That is a real dependency conflict rather than a missing package, and knowing it in advance saves an afternoon.
There is also a ComfyUI integration, a third party node project maintained by someone outside the core team, which appeared in December 2025. Its existence is a reasonable signal about where the community expects this kind of model to end up.
Blackwell cards need one flag changed before anything else
There is a specific hardware note that deserves more attention than it gets. xFormers is not yet fully compatible with the RTX 50 series architecture, and the README's instruction for those cards is to disable it:
python inference_offline.py --use_xformers FalseThis is the kind of detail that decides whether a first run works. Since `--use_xformers` defaults to true, anyone on a Blackwell card who follows the default invocation hits a crash before seeing a single frame. The failure is in the attention implementation rather than in the model, so the fix is disabling an optimisation, which means the card also loses the memory savings that made long videos possible in the first place.
That creates a real tension for a 12GB card. The streaming strategy was added so long videos would fit in that memory budget, and xFormers is part of how the base model stays inside it. Disable xFormers and you may need to reduce the frame count or rely more heavily on the streaming path to stay within memory. The README does not quantify that trade, which is a gap a reader has to resolve on their own hardware.
The workaround applies to offline inference. It is reasonable to assume the same attention path is used by the web interface, since both run the same model, but the README does not state this explicitly and the web UI has no documented flag for it. Anyone combining the web interface with a Blackwell card should expect to edit the configuration rather than pass a command line argument.
The broader pattern across this repository is that the base path is tested on datacenter hardware and the consumer paths are maintained reactively. The 20 minute TensorRT build, the frozen positional embedding, and the Blackwell note are all fixes layered onto a research release, each one documented once it became someone's problem.
Three training scripts, an academic licence, and no release history
Training code arrived on 2026-05-15, roughly five months after inference code, configs and pretrained weights shipped in December 2025. The repository root carries three separate entry points, train_stage1.py, train_stage2.py and train_stage3.py, which implies a staged procedure where each script depends on output from the previous one.
What is missing is the explanation. The README does not describe what each stage trains, what the data/ directory should contain, or what hyperparameters to use, so the staged structure is discoverable from the file names but not from documentation. Anyone hoping to retrain rather than fine tune will be reading the scripts themselves. This is the weakest documented axis of the project, and it is worth knowing before spending time on it.
The licensing position is unambiguous. The disclaimer states that the project is released for academic research only, that users must not use it to generate harmful, defamatory or illegal content, that the authors bear no responsibility for misuse, and that by using the code you accept sole responsibility for generated content. The licence file is Apache-2.0, which is permissive in the usual sense, but the disclaimer adds a use restriction on top of it. That combination deserves attention before any product work, since the permissive licence and the research-only statement are not obviously the same promise.
Finally, there are no published releases, so there is no version history to consult and no changelog beyond the dated todo list at the top of the README. That todo list is genuinely useful documentation: it dates the inference release, the paper, the ComfyUI integration, the streaming memory work, the WebUI improvement, CVPR 2026 acceptance, the training code, and the EditaLive successor. For a project with no tags, it functions as the changelog.
Taken together, the repository is strongest as a well documented inference path and weakest as a reproducible training path, and its licensing terms place it firmly in research use.
Editorial conclusion
PersonaLive is a good fit when the output has to run longer than a diffusion model naturally wants to, because streaming is designed into the loop rather than bolted on afterwards. What the repository settles is the runtime story: the conda setup, the exact weight directory layout, the inference flags, the xFormers workaround for Blackwell cards, and the TensorRT conversion path. What it leaves open is the training story, since the three stage scripts arrive without a walkthrough, and the licensing story, since the disclaimer restricts use to academic research and the successor project EditaLive already supersedes it for full body and editing work. Treat the code as a research artifact to read and reproduce rather than a component to ship in a product.
Frequently asked questions
What does PersonaLive generate, and how long can the output be?
It animates a reference portrait using motion from a driving video. The streaming strategy means output is not capped at a fixed length, and the changelog notes it was added specifically to generate long videos within a 12GB VRAM budget.
What do I need to run PersonaLive inference?
A conda environment with Python 3.10, the packages in requirements_base.txt, and the pretrained weights under pretrained_weights, which tools/download_weights.py can fetch. xFormers is on by default but must be turned off on RTX 50 series cards.
Can PersonaLive be used for commercial work?
The repository states it is released for academic research only, alongside an Apache-2.0 licence file and a disclaimer placing responsibility for generated content on the user. The research-only restriction is the more restrictive of the two and should be resolved before any product use.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/gvclab-personalive)