CLI tool
lynote-ai/ai-image-detector avatar
lynote-ai/ai-image-detector

ai-image-detector: the default backend scores one in three on the project's own three images

Open-source CLI, API, web UI, and reproducible benchmarks for probabilistic AI-generated image detection.

325 stars11 forksPythonMIT

At a glance

What is it?
A small open source detector with six backends, a CLI, an API, a web UI and benchmark runners, published at version 0.3.0. The most useful thing in the repository is its own smoke test, which scores the default model at 0.333 and shows the recommended ensemble making exactly the same decisions as a single model inside it.
Who is it for?
ai-image-detector is worth reading if you want a small, readable detection stack you can swap backends in, because the Python surface is four lines and the backend choice is a single flag. Two cautions come straight out of its own documentation.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 57 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The default model missed both AI images and produced no separation at all

The default backend is UnivFD, which is a CLIP ViT-L/14 image encoder with a tiny linear fake-or-real head on top. The project publishes a smoke test of three images and scores every backend on it.

The default's result:

| Image | `probability_ai` | Predicted | Expected | | --- | ---: | --- | --- | | `test_images/ai-generated.png` | 0.3098 | human | ai | | `test_images/ai_retouched.png` | 0.2832 | human | ai | | `test_images/human.jpeg` | 0.2801 | human | human |

Accuracy on this sample: `1 / 3 = 0.333`.

Two things stand out. The decision threshold defaults to a probability of at least 0.5 becoming the AI label, and all three scores sit between 0.28 and 0.31, so every image landed on the same side by a wide margin. The human image scored lowest of the three, which means the model's output on this sample carries no usable ordering.

The readme itself says borderline scores between 0.45 and 0.55 should be treated as weak evidence. Nothing in this table is borderline. It is uniformly uninformative.

The project is upfront about what this test is. Three images, labels taken from the filenames, and the sentence that it is not a publishable benchmark but only a quick sanity check. That framing is fair, and it is also the only published evidence for any backend.

The recommended ensemble makes the same three decisions as one of its four models

The backend list says `ultra` is the strongest current ensemble in the repository and the best first choice when you want the highest practical accuracy.

What goes into it is worth expanding. `ultra` adds Sentry on top of `hybrid-plus`. `hybrid-plus` is described as an ensemble of the project's internal hybrid with Nonescape. The hybrid is UnivFD blended with a lightweight Hugging Face classifier. So the strongest path loads four models: a CLIP encoder, a small image classifier, an adapted external detector, and a ConvNeXt.

On the smoke test the result is 0.667, and the note in the table reads: same decisions as Sentry on this three-image set.

That is the sharpest line in the repository. An ensemble of four models, recommended as the first choice, produces output identical to one of its members on the only evidence published. The ensemble may help on real data, but the repository contains nothing that shows it.

The other three alternatives also land at 0.667 and fail in three different ways. The Hugging Face backend is correct on the retouched image but false-positives on the human one. Nonescape Mini is correct on both AI-tagged images and also false-positives on the human one. Sentry ConvNeXt Small is correct on the plainly generated image and on the human one, so it is the one that misses the retouched image.

So on three images: one backend misses both AI images, one is fooled by the retouch, one is fooled by the human image, and the ensemble behaves identically to the third.

Two of the seven backends never appear in the recommendation list

The README names six backends in the model section: UnivFD as the default, then hybrid, nonescape-mini, sentry-convnext-small, hybrid-plus and ultra. The optional Hugging Face path adds a seventh under the name `hf`.

The section that answers which backend to use lists five of them. UnivFD as the simplest default, the ConvNeXt detector as the strong single external option, ultra as the strongest ensemble, nonescape-mini as useful extra signal and part of the stronger ensembles, and the generic Transformers path for standard checkpoints.

So `hybrid` and `hybrid-plus` are documented as selectable backends, with a documented flag for one of them, and neither appears in the list of what to pick.

The hybrid does have a usage example with a blend weight, which implies the two signals inside it can be traded against each other rather than simply averaged. `hybrid-plus` has no example at all, and appears only inside the description of ultra.

The naming convention is also worth noting. Three of the six carry a vendor or model family name, one carries a project's own name, and two are variations of another. A user choosing by reading the list has to guess which are separate models and which are combinations.

Three names for one package, one letter apart

The distribution is named `ai-image-detector`. The importable package is `aidetector`, with no separator and no dash. The console script is `aidetect`, one character shorter again.

So the three names a user types are `pip install ai-image-detector`, `from aidetector import ...`, and the `aidetect` command. The manifest declares the entry point as `aidetector.cli:app`, which is also why the console script is invoked as a bare command: the CLI is a Typer application rather than a set of argparse functions.

The gap between the module name and the command name is one letter, and both start with the same prefix, so a typo produces an import error rather than a command not found in a subtle way. It is a small friction that costs a newcomer about thirty seconds.

The repository name matches the distribution name, and the package directory at the root matches the import name. The other four top level entries are the packaging file, the benchmark runners, the test suite and the three smoke images, plus a security policy and the continuous integration configuration.

One more naming detail. The console script is installed unconditionally, which means the API and web extras do not add new commands to install; they add the dependencies the existing commands need.

The base install pulls a full CLIP stack and never mentions a model download

The pitch is that you install it, run one command, and get a probability. The install section is a virtual environment, an editable install, and five optional extras:

bash
python -m venv .venv
source .venv/bin/activate
pip install -e .
bash
pip install -e '.[eval]'      # Hugging Face dataset benchmarks
pip install -e '.[hf]'        # generic Hugging Face image-classification backend
pip install -e '.[api]'       # FastAPI server
pip install -e '.[web]'       # Gradio UI
pip install -e '.[dev]'       # tests and linting

The base dependency list is eleven entries and it is heavier than the description suggests. Two are Torch and Torchvision. Three more are the model-hugging stack: an open_clip implementation, the Hub client, and a safetensors reader. Two more are the vision model library and a progress bar. The remaining four are Pillow, the CLI framework, the console formatting library and a SOCKS proxy library.

That last one has no explanation anywhere in the file. It is a plausible dependency for fetching model weights through a proxy, but the readme never says so and never offers a way to configure one.

More importantly, none of these packages contains the detector weights. UnivFD is CLIP ViT-L/14 features plus a linear head, and CLIP ViT-L/14 is a model that has to be fetched. The install section says nothing about a first-run download, a cache location, an offline mode, or what happens when the network is unavailable. The result screen would block on that fetch.

The extras are each one dependency deep. The API extra adds three packages, the evaluation extra one, the Hugging Face backend one, the web interface one, and the development extra two. So the Transformers path is deliberately not in the base install even though the readme offers it as a first class backend option.

The operating point is called a more effective lever than the model

The benchmark commands accept post-hoc threshold calibration objectives, and two are named: balanced accuracy and F1. The readme's assessment is that this has been one of the most effective low-risk levers for improving held-out performance.

That is a different claim from choosing a better network. Threshold calibration moves the point at which a probability becomes a decision, which costs nothing to train and is exactly the kind of thing a calibration set exists for. If it is the highest-yield knob, then the interesting work is in the benchmark runner rather than in the model.

Two other pieces of guidance point the same way. The output section says that if a decision matters you should compare at least two backends, naming the ConvNeXt detector and the ensemble specifically, which implies the disagreement between two models is the signal rather than either model's score. And scores between 0.45 and 0.55 are called weak evidence rather than proof.

The hybrid's blend weight is the same idea at a different level. Passing a weight to the hybrid trades a CLIP-based signal against a small classifier signal, so the shipped defaults are operating points chosen by the author rather than by the data.

The recent research note in the same file says a different approach, combining CLIP semantics with low-level frequency and noise features, reports gains on two named benchmarks and would be a good target for a future backend. The current default is described as the simplest one that holds up. The smoke test in the repository says something different about that default, which is worth holding in mind when reading the preference.

An unauthenticated file upload endpoint bound to loopback

The API is three lines: install the API extra, run the command with a host and a port, then post a file with a multipart form request. Both values in the example are the loopback address and port 8000.

There is no authentication in the documented surface, no rate limit, no request size limit, no content type check, and no mention of either. For a loopback default that is a reasonable posture for a local tool.

It becomes a different question the moment somebody binds it to a public interface, which the host flag permits without further comment. The endpoint accepts an arbitrary image, runs a model over it, and returns a probability. There is no stated concurrency limit either, and the base install already carries Torch, so a handful of concurrent uploads on a small machine is a memory question nobody has written down.

The web interface is the same shape with a different dependency. Installing the web extra adds one package and gives you a `serve` command.

The Python surface is the smallest of the four. Four lines create a detector with an automatic device selection, predict on a path, and print the result as a dictionary. There is no documented batch method, no streaming, and no way to pass a preloaded tensor.

One release, cut at the same second as the last commit

There is a single release, version 0.3.0, and its timestamp is 2026-08-05 at 08:56:41. The last push carries the identical timestamp. So the release was cut from the final commit in the same second it was made, and nothing has been committed since, about two months ago.

That is not stale by the usual measure, and the repository is not archived, but it is a project at version 0.3.0 with a single tag, no changelog file, and no statement anywhere about a roadmap beyond the research note about a future backend.

The packaging is otherwise tidy. The build backend is declared with a floor, the package list is explicit rather than discovered, the console script is declared, and the linter is configured with a line length of a hundred characters and nothing else: no rule selection, no per-file ignores, and no type checker in the manifest at all despite the codebase being typed Python.

There is also no test configuration. The development extra installs pytest and the repository has a test directory, but the packaging file declares no test paths, no markers configuration and no coverage settings. The tests run on whatever default discovery finds.

The readme's own formatting gives a hint about how it is maintained. The opening disclaimer, the line about treating output as one signal rather than proof, ends with a raw HTML break tag, and the block of sibling project links that follows uses the same tag with bare unformatted URLs.

Editorial conclusion

ai-image-detector is worth reading if you want a small, readable detection stack you can swap backends in, because the Python surface is four lines and the backend choice is a single flag. Two cautions come straight out of its own documentation. The default backend performed worst of the five tested on the repository's smoke images, so do not treat the default as the safe choice, and the project itself calls its own calibration objectives a more effective lever than the model, which means you should expect to tune a threshold rather than pick a network. And treat every score as one signal, which is what the file asks you to do in its first paragraph.

Frequently asked questions

Are AI image detectors trustworthy?

The project's own position is that detection is probabilistic and the output should be treated as one signal rather than proof. It publishes a three-image smoke test in which the default backend scores 0.333 and the recommended ensemble scores 0.667, and it advises comparing at least two backends when a decision matters.

What is an ai image detector?

A tool that returns a probability that an image was machine-generated. In this project the default backend is a CLIP ViT-L/14 encoder with a tiny linear fake-or-real head, and the CLI applies a default threshold where a probability of 0.5 or above becomes the AI label.

Which backend should I use with ai-image-detector?

The readme calls `ultra` the strongest current ensemble and the best first choice, and offers `sentry-convnext-small` as the simpler strong single model. Note that on the repository's own three-image test the ensemble makes the same decisions as the ConvNeXt model, and two of the seven backends, `hybrid` and `hybrid-plus`, are absent from the recommendation list.

How does ai-image-detector stay offline and reproducible?

By default it makes no network calls for detection once the weights are present, and the base dependencies include a SOCKS proxy library, which is presumably for fetching weights through a proxy. The readme does not document where weights are cached, how to pre-download them, or what happens with no network, and it does not mention a first-run download at all.

Can I tune how ai-image-detector decides?

Yes, and the project says it is the highest-yield lever. The benchmark commands accept post-hoc threshold calibration objectives such as balanced accuracy and F1. The hybrid backend also takes a blend weight, and the output section calls scores between 0.45 and 0.55 weak evidence rather than proof.

Official sources

  1. License: MIT
  2. lynote-ai/ai-image-detector on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/lynote-ai-ai-image-detector.svg)](https://hysenlabs.com/projects/lynote-ai-ai-image-detector)