watermarks-remover: an agent skill that strips AI provenance marks from files you own
A privacy-first app that strips AI watermarks from content you own.
At a glance
- What is it?
- The project pairs deterministic Unicode cleanup with an optional model rewrite and a C2PA/EXIF scrubber for 20-plus file formats. It is a hygiene tool for content you own, not a way to launder someone else's media.
- Who is it for?
- Adopt it if you write or generate files with AI assistance and want the invisible Unicode, C2PA manifests and EXIF doc props gone before those files reach a repository, a client or a public upload. Skip it if your goal is pulling a stock-photo watermark off an image you did not produce: nothing in this repository is built for that, and the README frames the whole tool as hygiene for content you own.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 4 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The marks this project removes, and the ones it cannot
AI-generated text and media carry provenance in several unrelated places, and the README splits them into three layers rather than pretending one pass handles everything.
Layer A is deterministic: invisible Unicode, exotic space characters, bidirectional control characters and tag characters. These are bytes, so a Python script can find and remove them without a model in the loop. Layer B is statistical: token-sampling watermarks, where the signal lives in the choice of words rather than in any character. No regex touches that, which is why the project routes Layer B through an agent rewrite plus an optional rewrite_text.py hook. The third layer is file metadata: C2PA manifests, EXIF, XMP and document properties across PNG, JPEG, WebP, AVIF, HEIC, BMP, GIF, TIFF, SVG, PDF, DOCX, XLSX, PPTX, EPUB, ODT, HTML, Markdown, MP4/MOV/M4A/M4V, WAV, MP3 and FLAC.
The README names the vendor classes it targets at a class level: Claude, Gemini and SynthID-Text, OpenAI provenance surfaces, and open-LLM schemes such as Kirchenbauer-style green-list marks and keyed-Gumbel or EXP marks in the Aaronson family. That is a useful honesty marker. Class-level targeting means the tool is built against the published scheme, not against a specific model checkpoint.
The audience is narrow and the README says so: privacy and hygiene on content you own. If you are trying to lift a stock photo's visible watermark, this is the wrong repository.
How the skill, the service and the hook divide the work
The skill ships no code. It is markdown that calls the service over HTTP, so the agent host needs no Python installed. That split is the central design decision: the agent supplies judgement for the statistical layer, and the service supplies the deterministic file work.
The service is stdlib Python and listens on 127.0.0.1:8765 by default. The skill talks to it at WATERMARKS_SERVICE_URL when that variable is set to something other than the loopback address.
The interesting part is the hook. The README makes an argument worth repeating: a skill is an instruction, and the model decides whether to invoke it, while the model is also the thing producing the marks. A hook is executed by the harness on every matching tool call, cooperation not required. So the plugin registers a PostToolUse hook on Write|Edit|MultiEdit|NotebookEdit and runs service/scripts/hook_written_file.py against whatever file the agent just wrote. Two modes exist, following the pre-commit convention of checking by default: check reports provenance marks and leaves the file alone, sending findings to the model with exit 2 so it can offer to clean them; clean strips the marks in place and tells the model the file on disk changed.
Detection reuses audit_lib's scan_file and is_actionable, so the hook, the pre-commit gate and the CI SARIF export agree on what counts as actionable. Cleaning shells out to clean_file.py rather than duplicating logic. In clean mode the hook writes to a sibling temp file and swaps only on a real difference, so already-clean files keep their mtime and do not retrigger file watchers.
Installing the skill and running a first clean
The installer is stdlib Python and needs 3.10 or later with no dependencies. It validates the skill against the Agent Skills packaging rules that claude.ai uploads and the Skills API enforce before it writes anything: spec-only frontmatter, a lowercase hyphenated name of at most 64 characters matching the directory, and a non-empty description of at most 1024 characters.
The README gives this command for the full, service-backed skill:
python3 install_skill.py --skill remove-ai-marks --target claude-codeThat lands the skill in ~/.claude/skills/remove-ai-marks and honors CLAUDE_CONFIG_DIR. On Windows the README says to use py install_skill.py instead, and install-skill.sh is the wrapper for macOS and Linux shells. Two skills ship: remove-ai-marks, which is service-backed, and clean-user-facing-text, which is text only and self-contained. Passing --list prints them. Existing installations survive unless you pass --force, and replacement is staged first with the previous install kept as a uniquely named backup. --link symlinks the checkout instead of copying, so edits are picked up live.
For a project-scoped install, the target changes and takes a directory:
python3 install_skill.py --skill remove-ai-marks --target claude-project --project-dir .That writes into .claude/skills/ inside the given path. The other two targets are cowork, which produces dist/remove-ai-marks.zip for upload under Customize then Skills, and cursor, which is the default target and writes to ~/.cursor/skills/.
Next, bring up the service. The compose file maps the port to loopback only, so the service is not reachable from the network by default:
docker compose up --build -dAfter that, the skill and any web app talk to wr-core at http://127.0.0.1:8765. If you run the service somewhere else, set WATERMARKS_SERVICE_URL to match. With no configuration at all, the core service works: WATERMARKS_SERVER_API_KEY is empty by default, and setting it makes every request carry an Authorization: Bearer header.
The rewrite backend is a real dependency, not a detail
Layer B is where the project stops being a file utility. A token-sampling watermark is a statistical property of the text, so removing it means producing different text, and that means a model. The Makefile makes the default explicit and the reasoning is worth reading closely: the rewrite backend defaults to DeepSeek as a cross-model, non-origin choice.
REWRITE_BACKEND ?= openai-compatible
REWRITE_MODEL ?= deepseek-v4-flash
REWRITE_BASE_URL ?= https://api.deepseek.com
REWRITE_ALLOW_REMOTE ?= --rewrite-allow-remoteThose are Makefile variables you can override on the command line, for example pointing REWRITE_BASE_URL at a local server on port 8000 and dropping REWRITE_ALLOW_REMOTE. The default sends your text to a remote API. If your reason for using this tool is privacy, that default is in tension with the goal, and the override is the fix rather than a footnote. Running a local OpenAI-compatible server and clearing the remote flag keeps the rewrite on your machine.
The non-origin point matters more than it looks. Rewriting Claude output with Claude risks re-applying a mark from the same scheme. The Makefile's default is chosen to avoid exactly that, and if you swap in a model from the same family that produced the text, you have undone the design.
What a hook cannot reach, and where the tool stops
The README is unusually direct about the boundary. No hook can rewrite the assistant's chat message before you read it. Claude Code's Stop hook receives last_assistant_message read-only, and there is no pre-send filter for final responses. The project documents the same limit for Cursor rules.
The practical consequence: the deterministic guarantee covers files the agent writes, plus the pre-commit gate for anything heading into git. Text that only ever exists in the chat transcript still depends on the skill workflow, which is the cooperative half. If your threat model is a transcript you paste somewhere later, the hook does not help you and the skill only helps if the model chooses to run it.
Two more constraints sit in the compose file. The ctrlregen and synthid images bake in upstream code that is not publicly redistributable, under an all-rights-reserved or non-commercial Research License, so they build from source locally and are never pushed to GHCR. Those two live behind the heavy profile and are one-shot CLIs: up starts them with --help to confirm the image, and real jobs run through docker compose run. The harness profile adds markllm and markdiffusion. The README also notes that the MarkLLM harness is same-config-only detection and not a vendor oracle, which is a meaningful caveat if you were hoping for a general detector.
One more: the .env.example records that Google removed SynthID text watermarking from the Generative Language API, and the gemini-synthid-text detector was removed as a result. Vendor detection surfaces move, and a detector can disappear underneath you.
How it compares with a general-purpose media watermark tool
The obvious alternative is a general watermark remover aimed at visible marks on images and video, which is what most search traffic for this topic is actually looking for. The difference in approach is not a matter of quality. Those tools inpaint or clone pixels over a logo; this project does not touch pixels for that purpose at all. It reads and strips provenance metadata and invisible characters, and it rewrites statistical text marks.
So the two do different jobs. If you have a video with a platform overlay burned into the frame, a C2PA scrubber is irrelevant, and the README's own framing, hygiene on content you own, tells you this is not the tool for it. Conversely, if you have a DOCX with document properties pointing at an AI service, or a PNG with a C2PA manifest, a pixel-level remover will not touch either.
Within its own category, the notable choice is the two-part architecture: a thin skill plus a service, with the deterministic work in the service and the statistical work delegated to an agent. That is what lets the agent host run without Python while the file formats stay handled by code that does not depend on model cooperation.
Licence, distribution and what to check before you depend on it
The repository is MIT licensed, which covers the code in this repository. It does not cover everything the compose file can build: the ctrlregen and synthid images incorporate upstream code under an all-rights-reserved or non-commercial Research License, and the compose file states plainly that those images are never pushed to GHCR. If you enable the heavy profile, you are building and running third-party research code under its own terms, and the MIT licence on this repository says nothing about that. The harness profile pulls gated models that need HF_TOKEN, and the .env.example notes the token is read from the environment only, never from argv.
The last push to the repository was on 2026-09-13, and v0.7.0 was released on 2026-09-03 with a /clean Layer B rewrite, a watermark-stealing module, audio and video watermark removal, and broader benchmark tooling. The release cadence visible in the repository is roughly one minor version every two to three weeks, with v0.6.0 on 2026-08-26 and v0.5.0 on 2026-08-14. That pace is the upgrade cost you are signing up for: the detector surface tracks vendor behaviour, and the .env.example already records one vendor detector being retired when Google changed its API.
A migration note worth catching: the skill was formerly named remove-claude-marks, and the slash alias /remove-claude-marks is still documented. If you have older instructions or scripts referencing that name, they may still work, but the current skill name is remove-ai-marks.
Editorial conclusion
Adopt it if you write or generate files with AI assistance and want the invisible Unicode, C2PA manifests and EXIF doc props gone before those files reach a repository, a client or a public upload. Skip it if your goal is pulling a stock-photo watermark off an image you did not produce: nothing in this repository is built for that, and the README frames the whole tool as hygiene for content you own. Before you commit, confirm two things for yourself: that your target host is one of the four installer targets (claude-code, claude-project, cowork, cursor), and that your rewrite backend is a model other than the one that produced the text, since the Makefile defaults to DeepSeek precisely to avoid re-marking with the origin model.
Frequently asked questions
What is the best way to remove a watermark?
It depends on what kind of watermark you mean, and this project only covers AI provenance marks. The README splits the work into invisible Unicode and exotic characters handled by deterministic Python scripts, statistical token-sampling text marks handled by an agent rewrite plus an optional rewrite_text.py hook, and file metadata such as C2PA, EXIF, XMP and document properties.
What is the best watermark remover?
watermarks-remover is built for privacy and hygiene on content you own, not for stripping visible overlays from media you did not produce. It covers text and file metadata across formats including PNG, JPEG, PDF, DOCX, MP4/MOV and MP3/FLAC, and it targets Claude, Gemini and SynthID-Text, OpenAI provenance surfaces and open-LLM green-list and keyed-Gumbel schemes at a class level.
How can I make a homemade watermark remover?
You can start from the deterministic half, which is plain Python with no dependencies. The project ships install_skill.py, which needs Python 3.10 or later, and a stdlib service you can start with docker compose up --build -d on 127.0.0.1:8765. The statistical layer is harder to build yourself because it requires a rewrite model, which is why the Makefile defaults to a cross-model, non-origin backend.
What is the best high-quality watermark remover?
The README does not make quality claims or publish comparison results, so there is no ranking to quote. What it does describe is coverage: invisible Unicode, bidi and tag characters in Layer A, token-sampling marks in Layer B, and C2PA/EXIF/XMP plus document properties across more than twenty file formats. The release notes for v0.7.0 mention benchmark and tooling breadth.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/guillaumemeyer-watermarks-remover)