Open-source project
OpenVGLab/OmniLottie avatar
OpenVGLab/OmniLottie

OmniLottie: four entry points, two model formats, and a timing table measured in a different unit than the paper

[CVPR 2026🔥] 🧑‍🎨 OmniLottie, an open-sourced multi-modal instructed vector animation generator that produces Lottie JSONs.

799 stars40 forksPythonApache-2.0

At a glance

What is it?
Inference code and weights for a multimodal vector animation generator, plus a two million item dataset and a benchmark. The repository keeps both the original checkpoint format and the newer safetensors format side by side, names the same argument differently in each, and leaves training code as the one unticked item on its own release plan.
Who is it for?
Use OmniLottie if you need Lottie JSON rather than raster video and you have a CUDA machine with room for roughly fifteen gigabytes of memory. Five things to check before you plan around it.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Activity is slowing. The repository last received commits 6 months ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 10, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The timing table is linear and in a different unit than the paper

One row gives the hardware cost: 15.2 gigabytes of GPU memory, and a time column reading 8.34, 16.68, 33.38, 66.74 and 133.49 seconds for token budgets of 256, 512, 1024, 2048 and 4096. Each figure is exactly double the one before it, so the cost is a straight line in the budget and the largest setting is about two minutes and thirteen seconds. A red note under the table says the numbers here are measured per OmniLottie Lottie token, while the figures in the paper are measured per JSON code token so they can be compared fairly against baseline methods. Two different units sit in the same table as the project's own reporting, which is the reason to read the note before quoting either number.

Four entry points at the root, one argument with two names

The repository root holds four scripts that do the same work in pairs. The originals, inference.py and app.py, load a single checkpoint file, and the newer inference_hf.py and app_hf.py load safetensors shards with a config. The argument for the model is called one thing in the older pair and another in the newer pair, and the output changes shape too: the newer scripts take a single output file, the older ones take an output directory. The input switches differ as well, with a batch text file and a single image on the original path against separate text, image and video options on the newer one. Two parallel APIs for one model is the price of supporting two checkpoint formats, and the page does not map one onto the other.

Two model formats, and the switchover happened in March 2026

The older format is a single file named pytorch_model.bin, described as being for users who downloaded the model before HuggingFace format support arrived. The news list dates that change to 20 March 2026, and the same entry records that the safetensors shards with a config file became the recommended path for new users, with automatic downloading through the from_pretrained interface. The page states that both formats produce identical results and tells you to choose based on which one you have. The weights are a single 4B model at 8.46 gigabytes, dated 2 March 2026. So anyone who started in the first three weeks of March has a different download from everyone who started after, and the repository keeps both decoder files for the same reason. Fetching the new format means installing the hub client first and running one download command:

bash
huggingface-cli download OmniLottie/OmniLottie --local-dir /PATH/TO/OmniLottie

requirements.txt is a compatibility ledger with a dangling reference

The dependency file reads like a record of what broke and when. NumPy is capped below 2.3.0 with a comment naming the opencv release that needs 2.0 or later and the pandas release that needs the other side, so the range straddles the major version boundary on purpose. Pillow is held under 12.0 for the web interface, pydantic under 2.12, and fsspec is bounded at both ends for the dataset library, with the hub client given a floor that satisfies the web interface and the dataset library rather than the model. One comment cites a requirement from a training library, but no entry for that library appears anywhere in the file, which is consistent with training code being the one item left unopened.

The one unticked box is the training code

The open source plan has six items and five of them are ticked: the project page and technical report, the two million item dataset, the inference code and weights, a Gradio demo deployed on the hub, and the benchmark. The sixth, training code, is empty. So the release is inference and artefacts rather than a pipeline, which is the normal shape for a research code drop and the normal thing to check before planning a fine tune. The benchmark itself is in the repository as its own directory with a download script, and the dataset and the demo live on the model hub rather than here, along with a community plugin for a node-based image editor that is not part of this repository.

The top of the page still names a different project in a comment

The first thing in the file is a commented out block. Its heading reads Official repo for OmniSVG, in a title element wrapped in an HTML comment, and a second commented block near the badge row would have linked back to this repository on GitHub. Everything active below it is for OmniLottie. Small residue like that usually means the page was copied from a sibling project and edited rather than written, which matters here because the two decoder files, the two script pairs and the two weight formats all follow the same duplication pattern. The page is a paper artefact that has been kept in step with the code, not the other way round.

The environment is pinned to an older CUDA build

The installation sequence is conventional: clone, create a conda environment with Python 3.10, then install PyTorch from a specific index URL and the remaining dependencies from a requirements file. The PyTorch pin is a version with a CUDA 12.1 suffix, paired with a matching torchvision build, installed from a specific index URL:

bash
pip install torch==2.3.0+cu121 torchvision==0.18.0+cu121 --index-url https://download.pytorch.org/whl/cu121

The page says that combination is what the environment was tested against, with a link to the toolkit installer for anyone without it. Nothing on the page mentions a CPU path, a lower precision path, or a second CUDA version, and the 15.2 gigabyte memory figure is the only sizing information given. A machine without a CUDA device has no documented way to run the model from this repository.

Editorial conclusion

Use OmniLottie if you need Lottie JSON rather than raster video and you have a CUDA machine with room for roughly fifteen gigabytes of memory. Five things to check before you plan around it. The training pipeline is not open, so reproducing or fine tuning the model means writing your own loop on top of the released weights. The same weight download exists in two formats and the argument names differ between the old script and the new one, so pin the format first and read which entry point you are invoking. The timing table is linear in the token budget and is quoted per Lottie token rather than per JSON code token, so do not compare those numbers against figures in the paper. The environment is pinned to an older CUDA build, and the dependency file caps several libraries for compatibility reasons recorded in its own comments, one of which cites a package the file never installs. And with the last push in April 2026 and no releases published, the project page is a snapshot of a research release rather than a maintained tool.

Frequently asked questions

How much GPU memory does OmniLottie inference need?

The published table lists 15.2 gigabytes of GPU memory usage for the single 4B model, whose weights are 8.46 gigabytes. Times are given for token budgets from 256 to 4096.

What is the difference between the two OmniLottie model formats?

The original format is a single pytorch_model.bin used with inference.py and app.py, for downloads made before HuggingFace format support. The newer format uses safetensors shards with a config file, is used with inference_hf.py and app_hf.py, and is the recommended path for new users.

Is OmniLottie's training code available?

Not in this repository. The open source plan lists training code as the single unticked item, while the report, dataset, inference code, weights, demo and benchmark are all marked as released.

Which inputs does OmniLottie accept?

Text, images and video, according to the introduction and the quick start. The newer script exposes a text option, an image option and a video option, and the original script takes a batch text file or a single image.

What Python and CUDA versions does OmniLottie expect?

A conda environment with Python 3.10, and PyTorch installed from the CUDA 12.1 wheel index, which is the combination the page says the environment was tested against.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. OpenVGLab/OmniLottie on GitHub
  4. Project website
  5. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/openvglab-omnilottie.svg)](https://hysenlabs.com/projects/openvglab-omnilottie)