Self-hosted service
Netflix/vmaf avatar
Netflix/vmaf

VMAF: Netflix's Perceptual Video Quality Metric, and When It Misleads You

Perceptual video quality assessment based on multi-method fusion.

5,503 stars834 forksCNOASSERTION

At a glance

What is it?
VMAF scores compressed video against a reference using a fused model of several elementary metrics. Here is what the repository documents, how to build it, and where the metric stops being the right tool.
Who is it for?
Adopt VMAF if you encode video and need a reference-based quality number that tracks perception better than PSNR or SSIM, and if you can run the vmaf tool or the FFmpeg libvmaf filter over both the source and the encoded file.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 4 days ago.
What is it written in?
Mainly C, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What VMAF scores, and who needs a number like this

VMAF is a perceptual video quality assessment algorithm developed by Netflix. The problem it addresses is concrete: when you compress a video, you need to know how much visible damage the encoder did. Peak signal-to-noise ratio and structural similarity give you numbers too, but they do not track what viewers actually notice. VMAF computes a score by fusing several elementary quality features into one prediction, and the repository ships implementations of those underlying metrics alongside it: PSNR, PSNR-HVS, SSIM, MS-SSIM and CIEDE2000. The README describes the package as a stand-alone C library, libvmaf, plus a wrapping Python library, and the Python side is what you use to train and test a custom model.

The audience is narrower than the name suggests. VMAF is reference-based: you need the original video and the encoded video, and you compare them frame by frame. That makes it a tool for encoder developers, codec evaluation teams, and anyone tuning a rate-control or preset configuration who wants to know whether a change helped. It is not a tool for auditing a file you received from someone else, because you have nothing to compare against. The repository also notes that AOM has specified VMAF as the standard implementation metrics tool under the AOM common test conditions, which tells you where the project expects to be used: in codec comparison harnesses, not in one-off manual inspection.

How the fusion model and the libvmaf pipeline fit together

The architecture visible in the repository is a C library with a tool on top. libvmaf provides the interface to incorporate VMAF into your own code and, per the README, tools to integrate other feature extractors into the library. The command-line tool vmaf provides a complete algorithm implementation, so you can deploy it without writing C. The Python library sits above both and offers wrapper classes and scripts for software testing, model training and validation, dataset processing and data visualization.

The fusion idea is that no single elementary metric predicts perceived quality well on its own, so the model combines them. The score is therefore a model output, not a direct measurement, and that distinction matters when you interpret results. The repository keeps model documentation in two separate places: models_v0.md for the previous generation and models_v1.md for the v1 models released in 2026-06, with the README pointing to the v1 page for details and model selection guidance. If you are reading a blog post or a forum thread about VMAF scores, check which generation it refers to, because the two are documented as distinct and the project explicitly says it will continue to update the v1 models over time. A score produced by one generation is not automatically comparable to a score produced by the other.

There is also a mode worth knowing about before you trust any comparison. The README describes NEG, short for No Enhancement Gain, introduced in 2020 after Netflix published a memo on VMAF's behavior in the presence of image enhancement operations. The short version: sharpening and similar filters can raise a VMAF score without making the video better for a viewer, which distorts codec evaluation. NEG exists to remove that gain from the measurement. If you are comparing two encoders and one of them applies any enhancement, a plain VMAF run is measuring the wrong thing.

Building VMAF and running the vmaf tool

The repository offers several routes. The fastest to reason about is the Dockerfile, which builds from ubuntu:22.04, installs the build toolchain, retrieves nv-codec-headers at the tag pinned in the file, builds libvmaf with ENABLE_NVCC=true, installs the Python requirements from python/requirements.txt, sets PYTHONPATH to python, and sets the container entrypoint to vmaf. Building that image gives you a container whose default command is the VMAF tool itself.

bash
docker build -t vmaf .
docker run --rm -v "$PWD:/work" vmaf

The first command builds the image from the repository root. The second runs the container with your working directory mounted at /work, and because the entrypoint is vmaf the container prints the tool's usage output, which is the quickest way to confirm the binary is present.

If you prefer a native build, the top-level Makefile drives meson and ninja from a Python virtual environment. The target you want is build, which configures libvmaf with -Denable_float=true and -Denable_cuda=true, then runs ninja.

bash
make build

After that, the vmaf binary lands under libvmaf/build/tools, which is the path the Dockerfile also adds to PATH. The README does not reproduce a full argument list for the tool, so read libvmaf/tools/README.md before scripting anything. What matters operationally is that the tool consumes a reference and a distorted input and produces a report you can read per frame or as a pooled value. Keep the per-frame data: a single pooled number hides a bad scene. The README also does not document a rollback procedure for a model or a library upgrade, so if you are pinning VMAF inside a CI pipeline, pin the container tag or the git revision yourself rather than tracking the default branch.

FFmpeg users have a second route. The README states that VMAF is included as a filter in FFmpeg and is configured with --enable-libvmaf at build time, with a dedicated page at resource/doc/ffmpeg.md. That is convenient when your pipeline already decodes through FFmpeg, but it means your VMAF version is tied to your FFmpeg build.

Where VMAF is the wrong instrument

The reference requirement is the hard boundary. Without the original video you cannot compute a VMAF score at all, so any workflow that starts from a delivered file is out of scope. Teams sometimes expect VMAF to act as a general quality gate on incoming media, and it cannot do that.

The second limitation is the one the project itself flagged. Because VMAF is a learned fusion of features, an encoder that applies enhancement operations can move the score in ways that do not correspond to a better picture. Netflix's own memo on this property is the reason NEG mode exists. Turning NEG on is not free of consequences either: you are changing what the metric measures, so scores from a NEG run and a default run should not be mixed in the same table.

The third issue is model selection. The repository documents v0 and v1 models separately and points to a selection guide, which is an admission that the choice is not automatic. A model trained for one class of content or viewing condition applied to another produces a number that still looks authoritative. Nothing in the tool will warn you.

Finally, the codebase is C with SIMD paths and optional CUDA. The README notes that the v2.0.0 release introduced a fixed-point, x86 SIMD-optimized implementation, and the Makefile shows the CUDA path is gated behind ENABLE_NVCC and -Denable_cuda. Builds are therefore sensitive to your toolchain and to whether the CUDA dependencies are present; the Dockerfile installs nvidia-cuda-dev and nvidia-cuda-toolkit for exactly that reason. If you are not prepared to maintain that build, use the container.

VMAF against PSNR and SSIM, which ship in the same binary

The most direct alternative is not another project. It is PSNR or SSIM, both of which the README says are included in libvmaf alongside PSNR-HVS, MS-SSIM and CIEDE2000. The practical difference is that PSNR and SSIM are formulas, not trained models. You can compute them, explain them, and compare numbers across versions of the tool with reasonable confidence, because there is no model generation to track. VMAF trades that stability for correlation with human perception, and the trade is usually worth it when you are making encoding decisions.

A sensible setup uses both. Score with PSNR or SSIM as a cheap sanity check that nothing catastrophic happened, then use VMAF for the decision you actually care about. The vmaf tool supports this directly: the README describes the auxiliary features as part of the same command-line tool, so you are not maintaining two pipelines.

If your question is specifically about banding artifacts rather than overall quality, the repository also includes CAMBI, the Contrast Aware Multiscale Banding Index. The README documents two input parameters added to it: full_ref, which lets CAMBI run as a full-reference metric so it accounts for banding already present in the source, and max_log_contrast, which extends detection to higher contrasts than the default. The project notes that CAMBI was sped up substantially for 4K, and that the PCS 2021 paper describes an earlier version that no longer matches the code exactly. That last point is a good general warning about this repository: the documentation and the implementation drift, and the release notes are often more current than the papers.

Licence, releases and what an upgrade actually costs

The licence is not a standard identifier in the repository metadata, and the README explains why: VMAF was changed from Apache 2.0 to BSD+Patent in February 2020, described there as more permissive than Apache and including an express patent grant. That patent grant is the part worth reading if you are embedding libvmaf in a commercial product, and it is the reason the licence field does not show up as a plain SPDX string. Read the LICENSE file in the repository root rather than relying on a scanner's classification. Nothing here is legal advice.

Upgrade cost is dominated by model generation, not by the build. The project released libvmaf v3.0.0 in December 2023 with optimizations, bug fixes, and removal of APIs that had been deprecated in v2.0.0. That pattern tells you the maintainers are willing to break deprecated surface area on major versions. On the model side, the v1 models arrived in 2026-06 and the README states the project will continue to improve and update them over time. If your quality thresholds are calibrated to v0 scores, moving to v1 invalidates those thresholds until you recalibrate. Budget for that as a real task, not a version bump.

Maintenance status is easy to state from the facts: the repository is not archived, and the last push was on 2026-09-17, with v3.2.1 released on 2026-09-14. The project is being worked on. That does not mean the API is stable, as the v3.0.0 removals show.

Editorial conclusion

Adopt VMAF if you encode video and need a reference-based quality number that tracks perception better than PSNR or SSIM, and if you can run the vmaf tool or the FFmpeg libvmaf filter over both the source and the encoded file. Do not adopt it as a no-reference quality monitor for files whose source you no longer have, and do not treat a single VMAF number as a verdict on a codec without checking whether the encoder applies enhancement filters, which is exactly the case NEG mode was introduced to address. Before committing, verify three things: which model generation you are scoring with (v0 or v1, since the two are documented separately and the README says v1 was released in 2026-06), that your build actually enables the CUDA path if you expect GPU acceleration, and that your chosen model matches your content type, because the repository ships several models trained for different viewing conditions.

Frequently asked questions

What is a VMAF score?

It is the output of a perceptual video quality model that fuses several elementary metrics into a single prediction of how a compressed video compares to its reference. The repository ships the underlying metrics PSNR, PSNR-HVS, SSIM, MS-SSIM and CIEDE2000 alongside the fusion model.

How do I use Netflix VMAF?

You supply a reference video and a distorted video to the vmaf command-line tool, or to the libvmaf filter inside FFmpeg if your build was configured with --enable-libvmaf. The tool writes a report you can read per frame or as a pooled value.

Can I use VMAF without the original reference video?

No. VMAF is a full-reference metric, so it needs both the source and the encoded file to produce a score. The repository does not document a no-reference mode.

What is the difference between the v0 and v1 VMAF models?

They are two generations of models, documented separately in models_v0.md and models_v1.md. The README states that v1 was released in 2026-06 and points to the v1 page for model selection guidance, while v0 remains documented in its own file.

What is NEG mode in VMAF?

NEG stands for No Enhancement Gain. It was added because image enhancement operations can raise a VMAF score without improving perceived quality, which distorts codec evaluation. The README points to models_v0.md for how to disable enhancement gain.

Does VMAF also report PSNR and SSIM?

Yes. The README states that implementations of PSNR, PSNR-HVS, SSIM, MS-SSIM and CIEDE2000 are included in libvmaf, and that the vmaf tool provides these as auxiliary features.

Official sources

  1. Issues
  2. Netflix/vmaf on GitHub
  3. README
  4. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/netflix-vmaf.svg)](https://hysenlabs.com/projects/netflix-vmaf)