# WatermarkRemover-AI: Florence-2 finds the mark, LaMA fills the hole

> A Python desktop application that detects watermarks with a Microsoft vision model, inpaints the region with LaMA, and ships both a PyWebview GUI and a command line entry point for batch image and video work.

**D-Ogi/WatermarkRemover-AI** — AI-Powered Watermark Remover using Florence-2 and LaMA: Remove watermarks from images and videos, including AI-generated content from Sora, Runway, and others. Features a modern PyWebview GUI.

- Repository: https://github.com/D-Ogi/WatermarkRemover-AI
- Stars: 1,958 · Forks: 383
- Language: Python
- License: MIT
- Published: 2026-10-06 · Updated: 2026-10-06 · Language: en
- Canonical page: https://hysenlabs.com/projects/d-ogi-watermarkremover-ai

## Two models doing two different jobs

The architecture is the easiest thing to understand about this project, and the README is direct about it: Florence-2 from Microsoft handles watermark identification, and LaMA performs the inpainting that fills in the removed regions. Those are genuinely different problems. Detection is a question about where something is, and inpainting is a question about what should occupy that space once it is gone, which is why using a single model for both is unusual.

The tech stack section names the rest: PyWebview as the cross platform window, Alpine.js for the interface markup, and PyTorch as the deep learning backend. The repository topics agree with that reading and add one term that hints at a secondary use, since alongside florence-2, inpainting and lama-cleaner there is a dataset-creation topic.

The project states its target plainly, naming AI-generated video from Sora, Sora 2 and Runway as the motivating case. That is a narrower goal than the name suggests. Watermarks from those tools tend to sit in predictable places, semi-transparent, often moving, which is a friendlier detection problem than a photographer's copyright mark embedded across a whole image.

## Setup scripts per platform, and a Windows path with no Python

Installation is handled by three scripts in the tree, `setup.bat` alongside `setup.ps1` for Windows and `setup.sh` for Linux and macOS. The Windows route in the README is three lines:

```powershell
git clone https://github.com/D-Ogi/WatermarkRemover-AI.git
cd WatermarkRemover-AI
.\setup.ps1
```

and the shell equivalent adds the executable bit before running:

```bash
git clone https://github.com/D-Ogi/WatermarkRemover-AI.git
cd WatermarkRemover-AI
chmod +x setup.sh
./setup.sh
```

The macOS and Linux path requires Python 3.10 or newer on the system, while Windows needs none, because the setup script pulls down a portable Python environment by itself. Afterwards you launch with `run.bat` on Windows or `./run.sh` elsewhere.

The versioned Windows release goes further and removes the setup step entirely. Release 0.68.0, published on 2026-09-16, offers a CPU plus NVIDIA package at roughly 1.68 GiB using 7z compression to stay inside GitHub's per asset limit, and a smaller CPU only zip at about 470 MiB. Both bundle the EXE launcher and their own Python and Qt runtime, and the application picks CUDA when it is available and falls back to CPU otherwise.

## The CLI is where the real options live

The GUI covers the common path: pick a language and theme, choose single file or batch, set input and output paths, then start processing, with settings saved between launches. The command line exposes considerably more, and it is where you should look first if you want to understand what the application can actually do.

```bash
# Basic usage
python remwm.py input.png output_folder/

# With options
python remwm.py ./images ./output --overwrite --max-bbox-percent=15 --force-format=PNG
```

`remwm.py` is the entry point in the tree, with `remwmgui.py` for the graphical shell. The options table in the README is worth reading closely because it shows the assumptions built in. `--max-bbox-percent` defaults to 10, meaning a detection larger than a tenth of the image is rejected, which is a sensible guard against a model that decides the whole frame is a watermark. `--detection-prompt` defaults to the word watermark, and changing it is how you would target a specific logo or a fixed corner badge instead.

`--preview` is the option that makes the rest usable. It runs detection and draws the boxes without processing anything, so you can see whether the model found the thing you meant before committing to inpainting.

## Two pass detection and fade handling for video

Video gets its own set of controls, and the reasoning behind them is the most interesting design decision in the project. Detection runs on a sampling interval rather than every frame, which the README calls two pass mode and says is faster with `--detection-skip` greater than one. A watermark in a generated clip tends to stay in the same place across frames, so sampling every third frame and reusing the box is a large saving for modest risk.

```bash
# Process video with two-pass detection
python remwm.py video.mp4 ./output --detection-skip=3 --fade-in=0.5 --fade-out=0.5
```

The fade parameters address the failure mode that sampling creates. A watermark that fades in at the start of a clip will not be visible during the frames that got sampled, so `--fade-in` extends the mask backwards by a number of seconds and `--fade-out` extends it forwards. Same idea, opposite directions, and both expressed in seconds rather than frames, which matches how someone watching the clip would describe the problem.

Audio preservation needs FFmpeg installed separately, which the README handles per platform with `sudo apt install ffmpeg` and `brew install ffmpeg`. Without it the video still processes but loses its audio track. Supported containers are listed as MP4, AVI, MOV, MKV, FLV, WMV and WEBM, and the output format can be forced through `--force-format`.

## LaMA runs from TorchScript weights without IOPaint

The most consequential line in the README is a short paragraph near the end, and it is about migration rather than features. LaMA inpainting runs directly from verified TorchScript weights, and IOPaint is no longer required. Anyone who installed an earlier version and follows a stale tutorial will otherwise waste time assembling a dependency chain the application no longer uses.

The guidance that follows is practical. Use a fresh application environment, check it with `python -m pip check`, and consult the runtime and cache notes in `docs/lama-runtime.md` when upgrading. The three requirements files in the tree explain the structure: `requirements.txt` includes `requirements-core.txt` and adds the desktop UI dependencies, with per platform PyWebview extras so Linux gets concrete Qt wheels instead of a GTK source build.

```text
-r requirements-core.txt

psutil
pyyaml
```

The UI is built with a Tailwind pipeline declared in `package.json`, and the repository carries both `tailwind.config.cjs` and a lock file, so the styles are compiled rather than hand written. The Python tooling targets version 3.12 under Ruff, configured in `pyproject.toml` alongside the pytest settings.

## A multilingual GUI and an honest limit on the method

The interface is offered in English, French, Chinese, Japanese and Portuguese, selectable from the top right alongside the theme, and the README includes a deliberately silly sixth option labelled Brainrot. Small things like that are usually a sign the author is enjoying the project, and the theme support suggests someone who expects to use this for more than an afternoon.

The limit worth naming is the one no documentation can fix. Inpainting is a guess. LaMA reconstructs a plausible region given what surrounds it, which is why the results are convincing on a flat gradient and obviously synthetic on a face or a piece of text. On the AI-generated video this project targets, the regions are usually small, usually low contrast, and often over texture that repeats, which is the friendly case.

The second limit is detection. Florence-2 is given a text prompt, and the default is the single word watermark. A mark that reads as a logo, a piece of body text, or an overlay with no consistent appearance may simply not be found, which is exactly what `--preview` exists to reveal before you spend any processing time.

## Conclusion

WatermarkRemover-AI is a competent implementation of a specific two stage pipeline rather than a general purpose image editor, and understanding that boundary tells you whether it fits your task. Florence-2 is asked where the watermark is, LaMA is asked what should be there instead, and everything else in the application exists to get files in and out of those two calls. For still images and short clips the project covers the ground well: batch folders, a preview mode that draws detections without touching pixels, per platform setup scripts, and a portable Windows build that needs no system Python. The parts that need judgement are the parts the documentation leaves open, namely how well LaMA reconstructs a specific region and whether your watermark is something the detection prompt will find at all. Start with the CLI and the preview flag on a handful of frames, since that combination shows you what the model sees before you spend GPU time filling anything in.

## FAQ

### Is there any AI that can remove watermarks?

There are open source projects that do it with machine learning, and WatermarkRemover-AI is one of them. It runs Florence-2 from Microsoft to detect the watermark and LaMA to inpaint the region, with a PyWebview GUI and a CLI, under an MIT licence.

### What is a watermark remover?

It is a tool that identifies a watermark and reconstructs the pixels underneath it. In this project the identification step is a vision model given a text prompt, and the reconstruction step is an inpainting model that fills the detected box with what the surrounding image suggests belongs there.

### How do I check what the detector found before processing a whole video?

Use the preview mode, which runs detection and draws the boxes without writing anything. Combine it with a small detection skip on a short clip so you can see what Florence-2 considers a watermark and adjust the detection prompt before committing to a full pass.

## Sources

- [D-Ogi/WatermarkRemover-AI on GitHub](https://github.com/D-Ogi/WatermarkRemover-AI)
- [Issues](https://github.com/D-Ogi/WatermarkRemover-AI/issues)
- [License: MIT](https://github.com/D-Ogi/WatermarkRemover-AI/blob/main/LICENSE)
- [README](https://github.com/D-Ogi/WatermarkRemover-AI/blob/main/README.md)
- [Releases](https://github.com/D-Ogi/WatermarkRemover-AI/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/d-ogi-watermarkremover-ai
