SparkVSR hands you the keyframes and leaves the artifact correction to you
[ECCV 2026] SparkVSR: Interactive Video Super-Resolution via Sparse Keyframe Propagation
At a glance
- What is it?
- An ECCV 2026 interactive video super-resolution model from Texas A&M and YouTube, Google: you super-resolve a sparse set of keyframes with any off-the-shelf image model, and SparkVSR propagates those priors across the sequence. The paper's gains are stated without a dataset, and the page's pipeline figures are empty.
- Who is it for?
- Read SparkVSR as a controllable pipeline rather than a single model, because that is what the design is. Any image super-resolution model produces the keyframes, and a reference-free guidance mechanism keeps the propagator working when those keyframes are missing or imperfect, which is the part that decides whether it is usable on your footage.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 69 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 11, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The keyframes are yours, produced by any image model
The starting point is a complaint about video super-resolution as a black box. Existing approaches run end to end at inference time, so an unexpected artifact cannot be corrected and the only option is to accept the output.
SparkVSR changes who does what. You take a sparse set of keyframes and super-resolve them yourself, with any off-the-shelf image super-resolution model. SparkVSR then propagates those keyframe priors across the rest of the sequence while staying grounded by the motion in the original low-resolution video. The image model is therefore a replaceable component rather than part of the checkpoint, and the correction happens before the propagator ever runs.
Selection is flexible in three named ways: manual specification, codec I-frame extraction, or random sampling. Codec extraction is the interesting one for real footage, since it takes the frames the encoder already marked as keyframes instead of asking a person to pick them. The training side is a keyframe-conditioned latent-pixel two-stage pipeline that fuses low-resolution video latents with sparsely encoded high-resolution keyframe latents, and at inference a reference-free guidance mechanism balances keyframe adherence against blind restoration so the model still runs when reference keyframes are absent or imperfect.
The reported margins are not attached to a dataset or a baseline
The results sentence reports that the method surpasses baselines by up to 24.6%, 21.8% and 5.6% on CLIP-IQA, DOVER and MUSIQ respectively, across multiple video super-resolution benchmarks, with improved temporal consistency. Those are three headline margins and no table. The visible text does not say which benchmark produced which number, which baseline was beaten, or whether the margins are absolute or relative.
The leaderboard links add a wrinkle. Two of them are filed under the video super-resolution task, RealVSR at a 4x restoration protocol with DOVE, and SPMCS at 4x with the RealBasicVSR degradation. The other two, UDM10 and YouHQ40, are filed under video restoration instead. So a metric name and a benchmark do not line up one to one, and a reader has to visit Papers with Code to see which protocol each margin belongs to.
That page reportedly carries 14 verified evaluations, curated by the Hugging Face open-source team, dated 2026.08.03 in the news list. It is the place to resolve the attribution question this repository does not answer.
The pipeline sections hold empty placeholders
Two of the most useful sections on the page are structurally there and visually empty. Inference Pipeline is a heading followed by a single centered paragraph tag with nothing in it. Training Pipeline is the same, one heading and one empty centered paragraph. The intended diagrams are not inline on the repository front page.
The demo block follows the same pattern. It points at `assets/demo.mp4` inside an anchor, so the video file is referenced and the player region is left blank, which means the front page gives you the filename rather than the footage.
The rest of the header block is intact: the project page at sparkvsr.github.io, a Hugging Face repository under JiongzeYu, an arXiv entry numbered 2603.16864, and the Papers with Code record for the same number. The affiliation block credits Texas A&M University and YouTube, Google, and lists seven authors. Anyone who needs the two pipeline diagrams has to leave the repository for the project page or the paper, since the section that would carry them is present but carries nothing.
Torch is installed by a pinned command and is absent from requirements.txt
The dependency section states Python 3.10 or newer, PyTorch 2.5.0 or newer, and Diffusers, with everything else in `requirements.txt`. The install recipe is more specific than the dependency list, because it pins exact versions against a single CUDA wheel index.
# Clone the github repo and go to the directory
git clone https://github.com/taco-group/SparkVSR
cd SparkVSR
# Create and activate conda environment
conda create -n sparkvsr python=3.10
conda activate sparkvsr
# Install all required dependencies
pip install torch==2.5.0 torchvision==0.20.0 torchaudio==2.5.0 --index-url https://download.pytorch.org/whl/cu124
pip install -r requirements.txtTwo details matter. The file itself never mentions torch, torchvision or torchaudio, so `pip install -r requirements.txt` on its own does not produce a working install, and the pinned line has to run first. And the page says outright that the command may need adjusting for your platform, CUDA version and desired PyTorch version, pointing at the official previous-versions page. The declared floor of 2.5.0 and the pinned 2.5.0 are not the same constraint.
The environment name is `sparkvsr` and the interpreter is fixed at 3.10 inside the create command, even though the dependency line allows 3.10 or newer.
Two dataset tables, two directory names, and a page that stops mid-sentence
Training uses the same data as DOVE: HQ-VSR and DIV2K-HR, both placed under `datasets/train/`. The table counts HQ-VSR as 2,055 videos and DIV2K-HR as 800 images, with HQ-VSR hosted on Google Drive and DIV2K-HR at its official source. The directory example underneath shows what happens with that naming.
datasets/
└── train/
├── HQ-VSR/
└── DIV2K_train_HR/The folder is DIV2K_train_HR, which is the archive's own name rather than the hyphenated DIV2K-HR used in the table and the prose. A script written against the table will miss the directory.
The test side is five sets with frame counts attached: UDM10 at 10 synthetic videos averaging 32 frames, SPMCS at 30 synthetic videos averaging 32, YouHQ40 at 40 synthetic averaging 32, RealVSR at 50 real-world videos averaging 50, and MovieLQ at 10 old-movie videos averaging 192. Four of the five downloads are Google Drive links. The section closes by telling you to make sure the `datasets/test/` path is correct, and the visible text ends in the middle of that sentence, so the instructions that would follow the dataset list are not on the page.
The requirements file pulls in a UI, a tracker and two video decoders
requirements.txt lists 26 packages, and the list is worth reading as a description of the project's shape rather than as a version manifest.
Three entries are user-facing or operational rather than modeling: `gradio>=5.5.0` for an interface, `wandb` for experiment tracking, and `openai>=1.54.0` as an API client. Video input is covered twice, by `decord` and by `av`, with `scikit-video>=1.1.11` and `imageio-ffmpeg>=0.5.1` alongside them. `torchdiffeq` sits next to `einops` and `safetensors`, which is consistent with an iterative propagation step rather than a single feed-forward pass.
Two pins stand out. `numpy==1.26.0` is the only exact pin in the file, so an environment with a newer NumPy will be downgraded or will fail to install. Everything else uses a floor, several of them well above their release versions at the time of writing.
The tree also carries a `finetune/` directory, and `peft` is in the requirements, so parameter-efficient finetuning is part of what the repository is built to do rather than something the README promises for later. Weights are not in the tree; the header links a Hugging Face repository instead.
Two delivery paths: a shell script and a ComfyUI node
Inference at the repository level is a pair of files: `sparkvsr_inference.sh` and `sparkvsr_inference_script.py`, alongside `run_eval_all.sh` for the benchmark sweep and a `test_input/` directory for sample media. The GitHub releases list is empty, so those scripts are how you run the model.
The second path is ComfyUI. A `ComfyUI-Spark/` directory sits at the root, and the news list dates the ComfyUI-SparkVSR release to 2026.05.11, two months before the repo itself was announced on 2026.03.17 in that same list. The TODO block on the page has all five items marked done: inference code, pre-trained models, training code, project page and ComfyUI.
A third path is community deployments rather than local installs. The 2026.06.20 entry points at RunningHub.ai and CNAPS.ai, described as community deployments, and the 2026.06.18 entry records the ECCV 2026 acceptance. Note what separates the three: the ComfyUI directory is in the tree, the hosted deployments are third-party sites, and the model files themselves are on Hugging Face under JiongzeYu. Nothing in the repository pins a version of the weights, because there is no release to pin.
The generality claim rests on one old-movie set and one sentence
The last abstract sentence makes a bigger claim than the rest: SparkVSR is presented as a generic interactive, keyframe-conditioned video processing framework that applies out of the box to unseen tasks such as old-film restoration and video style transfer.
The test table supports the first half of that and not obviously the second. MovieLQ is the only old-movie set in the list, 10 videos averaging 192 frames, and it is also the outlier in every other column: four times the frame count of the synthetic sets and nearly four times that of RealVSR. Old film is where sparse keyframes and blind restoration would matter most, so its inclusion is coherent, but ten videos is a small basis and no figure for it appears in the visible text.
Video style transfer has no dataset, no metric and no script in what is visible here. The claim is one clause long in the abstract, it is not in the TODO list, and nothing else on the page picks it up. Treat it as a direction the authors name, not as a shipped mode you can run with the files in this repository.
Editorial conclusion
Read SparkVSR as a controllable pipeline rather than a single model, because that is what the design is. Any image super-resolution model produces the keyframes, and a reference-free guidance mechanism keeps the propagator working when those keyframes are missing or imperfect, which is the part that decides whether it is usable on your footage. Three things to check before adopting it. The reported margins, up to 24.6% CLIP-IQA, 21.8% DOVER and 5.6% MUSIQ over baselines, are not attributed to a dataset or a baseline in the visible text, so treat the ordering as indicative rather than as a table you can plan around. The install recipe pins one PyTorch build from a single CUDA wheel index and leaves torch out of requirements.txt, so a different GPU means editing the command rather than editing a file. And the repository has no tagged releases; weights live on Hugging Face under the first author's account, with ComfyUI and two community deployment sites as the other distribution paths.
Frequently asked questions
What makes SparkVSR interactive rather than a black box?
You super-resolve a sparse set of keyframes yourself with any off-the-shelf image super-resolution model, and SparkVSR propagates those keyframe priors through the sequence while staying grounded by the original low-resolution motion. Keyframes can be specified manually, extracted from codec I-frames, or sampled randomly.
Which datasets does SparkVSR train on and evaluate against?
Training uses the same data as DOVE: HQ-VSR with 2,055 videos and DIV2K-HR with 800 images, placed under datasets/train/ in folders named HQ-VSR/ and DIV2K_train_HR/. Evaluation covers UDM10, SPMCS, YouHQ40, RealVSR and MovieLQ, with frame counts from 32 to 192 per video.
What does SparkVSR need installed before it will run?
Python 3.10 or newer plus PyTorch and Diffusers. The recipe pins torch==2.5.0, torchvision==0.20.0 and torchaudio==2.5.0 from the cu124 wheel index, and requirements.txt does not list torch at all. The page warns the command may need adjusting for your platform and CUDA version.
Can SparkVSR run when there are no reference keyframes?
Yes. A reference-free guidance mechanism continuously balances keyframe adherence against blind restoration, so the model keeps working when reference keyframes are absent or imperfect rather than requiring a complete set of user-supplied high-resolution frames.
Where are the SparkVSR model weights and does the project publish releases?
There are no GitHub releases. The header points at a Hugging Face repository under JiongzeYu for the weights, the TODO list marks pre-trained models as released, and the news list records Papers with Code with 14 verified evaluations on 2026.08.03. ComfyUI-Spark/ is in the repository and community deployments are hosted on RunningHub.ai and CNAPS.ai.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/taco-group-sparkvsr)