IQA-PyTorch: a noncommercial licence, a Python 3.6 floor the dependencies cannot honour, and a default test target that skips two checks
🔎 🖼️ 🔥PyTorch Toolbox for Image Quality Assessment, including PSNR, SSIM, LPIPS, FID, NIQE, NRQM(Ma), MUSIQ, TOPIQ, NIMA, DBCNN, BRISQUE, PI and more...
At a glance
- What is it?
- The package published as pyiqa is a toolbox of image quality metrics in Python and PyTorch, with implementations calibrated against the original reference scripts where those exist. Its calibration results are committed to the repository, which is unusual and useful. The packaging around it is looser: the declared interpreter floor is three major versions behind the dependencies it installs, and the two checks most likely to catch a wrong number are not in the default test target.
- Who is it for?
- Use it for research and evaluation work, and read the licence before you plan to put it in a product, because the manifest declares a noncommercial licence and ships a second licence file for lab-contributed code. Two things to verify before trusting a number.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 36 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 5, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Three names for one thing, and a licence that is not open source
Start with the naming, because a search for the project will not find the package.
The repository is named after the field and the framework. The distribution on the package index is a different, shorter string, and the readme says so in the first line with an install command. The importable module and the console command both use the short name again.
So there are three names, and the one you install is not the one you would guess from the repository URL. It is one command:
pip install pyiqaThree further routes follow it in the same block: an installer for a faster resolver, an install straight from the git repository that starts by uninstalling any existing copy so it cannot shadow it, and a clone followed by an editable install.
The licensing is the part to read carefully. The manifest declares a PolyForm noncommercial licence, which permits use that is not commercial and does not permit the opposite. The project also ships two licence files at the root: the main one, and a second one whose name ends with a lab suffix, listed alongside it as a licence file for the package.
Meanwhile the project's own metadata records no licence at all. So the authoritative statements are the manifest and the two files, and the practical consequence is that a company evaluating this toolbox for a product is looking at a noncommercial grant.
The declared interpreter floor cannot be satisfied by its own dependencies
The manifest declares a minimum Python version of three point six, and its classifier list stops at three point eight.
The dependency list, read from a separate file at build time, requires a deep learning framework at one point thirteen or newer, its vision counterpart at zero point thirteen, a training library at zero point eight, and a transformer library at five or newer. None of those versions install on three point six, and the transformer library in particular is several interpreter generations past it.
There are also no upper bounds anywhere in that list, so the floor is the only constraint, and it is a constraint the floor itself violates.
What this means in practice is that the declared floor describes history rather than capability: it records the oldest interpreter the code was ever written against. The classifier list compounds it, since a package index will show support for three point six through three point eight and nothing newer, which is the opposite of the truth.
The project classifies itself as alpha in its metadata, which is at least consistent with a manifest this far behind its requirements.
The default test target runs two of the five checks
The build file lists five test targets and makes two of them the default.
The default target runs a forward inference check and a calibration check. The other three, each a named target of its own, are a check that CPU and GPU produce the same answer, a check that gradients flow backwards, and a check that the bundled datasets load.
Two of those three omitted checks are the ones that would catch a wrong number. This toolbox's central claim is numerical agreement with the reference implementations, so an implementation that behaves differently on CPU and on GPU is exactly the bug a user would notice and exactly the bug the default target will not catch.
The calibration check is in the default set, and it is marked as its own category in the test runner, which is the right way to enforce the central claim. But a category marker on a test only helps if the category is run.
The long target runs all five, so the suite is complete. It simply is not what you get by typing the obvious word.
Every metric has its own direction, and the loss example inverts one by hand
The advanced usage section contains the line that every user of this toolbox has to read twice.
After creating a metric, the example prints the metric's own attribute that says whether lower or higher is better. It is not a convention; it is data on the object, and it differs between metrics.
The loss section then builds on that. It notes that gradient propagation is disabled by default and has to be switched on explicitly to use a metric as a loss, warns that not every metric supports backpropagation and points at a model card per metric, and says to be sure you use it in the direction that metric expects.
The example then does the arithmetic by hand for the one metric whose direction is the opposite way round: it computes the score and subtracts it from one, with a comment saying that metric is not lower-better.
So the correct usage of any metric as a loss depends on reading that metric's direction and, where needed, flipping it yourself. That is a real ergonomic cost in a library whose selling point is that you can compare dozens of metrics through one interface.
One metric sets the dependency floor for everything else
The newest entry in the changelog is the one that shaped the dependency list.
It adds a visual scorer built on a vision and language backbone, offered in three sizes with parameter counts of zero point eight billion, four billion, and nine billion. The changelog entry for it states a requirement on the transformer library at version five or newer.
The dependency file says the same thing and adds the detail that the rest of the toolbox is verified against the five series, that the new metric wants five point two for the native support, and that on five point zero and five point one it runs through a shim vendored inside the package.
So a two-dozen-metric toolbox now carries a dependency floor set by one scorer, plus a compatibility shim for the versions below what that scorer wants.
That is the ordinary cost of a fast-moving research dependency entering a stable library, and it is the kind of decision that is easier to accept when the dependency file explains itself, which this one does, in a comment above the line.
No upper bounds anywhere, in a package that promises agreement with reference implementations
The dependency file is twenty-five lines and almost none of them carry a version bound.
Four are constrained: a training library at zero point eight, the framework and its vision counterpart at one point thirteen, and the transformer library at five. Everything else is a bare name, including the numerical library, the scientific library, the image library, the array library, and a headless build of the computer vision library.
One carries a platform marker instead of a version, an accelerator library that is skipped on macOS and unconstrained everywhere else.
The tension is with the project's own headline claim: that its implementations are calibrated against the official reference scripts. Calibration is a numerical claim, and numerical claims are exactly what a major version bump in a numerical library can invalidate. A pin would turn that risk into a resolver error; leaving it open turns it into a different number.
Three of the unpinned entries are also development tools rather than runtime ones: a test runner, a linter, and a pre-commit driver. Since the manifest reads its dependency list from this file, installing the package installs those three as well.
A target called refresh ends in an upload to the package index
The build file has one target that is worth reading before you type it.
It is a single line chaining five other targets in order: clean, build, install, build the distribution, then release. The release step runs the upload tool against everything in the distribution directory.
So a target named for refreshing your working copy will, in one command, build a distribution and publish it to the public package index. There is no confirmation step and no dry-run flag in the chain.
The rest of the file is more careful. The clean target removes build artefacts, egg metadata and caches, and uninstalls the package with a failure tolerated so the target works on a machine where it was never installed. A lint target is present but entirely commented out, which suggests the linter moved into the pre-commit configuration instead, and the requirements file agrees by listing that same pre-commit driver as a runtime dependency.
Editorial conclusion
Use it for research and evaluation work, and read the licence before you plan to put it in a product, because the manifest declares a noncommercial licence and ships a second licence file for lab-contributed code. Two things to verify before trusting a number. The dependency list has no upper bounds at all, so a fresh install can move a metric's value without anything in the package changing. And the checks that would catch that, comparing CPU against GPU and confirming gradients flow, are only run by the long test target, not the default one.
Frequently asked questions
What does IQA stand for in the context of quality?
It stands for image quality assessment, and this repository is a toolbox of metrics for it: implementations of full-reference and no-reference metrics written in Python and PyTorch, with results calibrated against the original reference implementations where those exist. The calibration results themselves are committed to the repository.
What is pyiqa?
It is the published package name, which is shorter than the repository name. There is a command of the same name that lists the available metrics and scores an image or a directory, and an importable module whose main entry point creates one metric at a time from the same list.
Is PyTorch just Python?
This repository is Python built on that framework, and it uses the framework's device abstraction to select a GPU when one is available and a CPU otherwise. Its install requirement is a minimum framework version with no upper bound on any dependency in the list.
Is PyTorch used for CNN?
Nothing here answers that question directly. What the project does state is that its implementations run considerably faster on a GPU than the reference implementations they were calibrated against, which is the framework's accelerator being applied to a specific workload.
Is ChatGPT built on PyTorch?
Not a question this repository answers; it is a library of image quality metrics. What it does say about models is that its newest metric is a visual scorer built on a vision and language backbone, released in three sizes, the largest of them a nine-billion-parameter model.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/chaofengc-iqa-pytorch)