Model or dataset
oxbshw/watch-skill avatar
oxbshw/watch-skill

Watch Skill: timestamped video and audio evidence for AI agents, with contracts that verify the work

Give AI agents eyes, ears, and verifiable results. Watch Skill turns video, audio and screen activity into searchable, timestamped evidence and proves work with deterministic contracts, not model opinion. DeepWatch is the agent workspace built on DeepSeek Harness. Python + npm, MCP, CLI, REST, Web.

403 stars58 forksPythonMIT

At a glance

What is it?
Watch Skill indexes video, audio and screen recordings into searchable, timestamped evidence, and checks work against frozen contracts instead of asking a model whether it looks right. DeepWatch is the workspace built on top of it.
Who is it for?
Adopt Watch Skill if you already run an MCP client and need an agent to cite a moment in a recording instead of paraphrasing it, or if you want a verification verdict that a separate process produces. Skip it if you only need a one-off transcript and never intend to re-query the source: the indexing step is the product, and without it you are paying for storage you do not use.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 15 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem Watch Skill solves, and who it is for

An agent that can read text has no way to answer a question about a recording. You can hand it a transcript, but a transcript loses the screen. You can hand it frames, but frames lose the audio. Watch Skill's answer is to index a source once and keep both, with every extracted artifact carrying an absolute timestamp, so a later question returns a citation the reader can open rather than a summary they have to trust.

The second half is verification, and it is the part that is easy to overlook. The README describes a frozen contract covering file digests, JSON values, SQL results, HTTP responses and DOM state, evaluated by a separate process. The verdict is one of VERIFIED, FAILED, UNVERIFIED or INCONCLUSIVE. The README's own framing for four answers instead of two is that three of them are not the same as "no", which is a fair point: a check that never ran and a check that ran and failed are different facts, and collapsing them into a boolean hides the difference.

The intended audience is narrow and specific. You already have an agent (the README names Claude Code, Cursor, Codex and any MCP client) and you want to give it video and audio input without building the extraction and indexing pipeline yourself. The second audience is anyone who needs an agent's work checked by something other than the agent that produced it. If neither describes you, the project is more machinery than the task needs.

How indexing and verification actually fit together

The README splits the system into two capabilities that work apart. Perception turns video, audio and screen activity into frames, transcripts and OCR text, each with a timestamp, and indexes the source so it can be queried for as long as you keep it. Verification takes a contract and evaluates it in a separate process.

The repository layout supports the split. There is a src/ tree for the Python engine, a workspace/ tree for the DeepWatch agent workspace, a schemas/ directory, and a top-level Dockerfile that installs the full dependency tier. The README describes four interfaces over the same engine: MCP, a CLI, a REST API, and Python. The MCP server is described as exposing 39 tools, which is a lot of surface area for a client to present, and worth checking in your own client before you commit to it.

DeepWatch is the second layer. The README says it is built on the official DeepSeek Harness with Watch Skill composed in, installed by one command, and that every tool call leaves a receipt naming what it touched. It also states that every path a tool declares is checked against one workspace boundary, so a tool cannot quietly write outside it, and that results carry a Core verdict and persist in a Library across a restart. Compare, in the browser UI, puts two runs of the same contract side by side and shows where the verdicts diverged.

Installing Watch Skill and asking a video your first question

The README's first path is for people who already have an MCP client. The install line carries an extra, and the README is explicit that this matters: a bare pip install gives you the CLI, the verifier and the Bridge, but cannot extract a frame, and watch stops at perceive.missing_dependency on the first video.

bash
pip install 'watch-skill[standard]'   # frames, retrieval and the MCP server
watch-skill doctor                    # checks, and repairs what it can
watch-skill watch <video-url-or-file>
watch-skill ask <id> "what changed at 3:12?"

Run doctor before anything else. The README says it names the exact command for whatever is missing, which is the fastest way to find out whether your environment can extract frames at all. The extras are additive: [standard] covers frames, retrieval and MCP; [ocr] reads on-screen text; [whisper] handles local transcription for sources with no captions; [loop] adds the browser; [all] takes everything.

To expose the engine to an agent, start the MCP server. The README describes it as a stdio server with 39 tools.

bash
watch-skill serve              # stdio MCP server, 39 tools

If you would rather install the skills into many agents at once, the README gives this command:

bash
npx skills add oxbshw/watch-skill -g

There is also a container path for people who do not want the dependency tier on their machine. The Dockerfile comments put the resolved install at roughly 600 MB of wheels, and the run commands mount a named volume. The comment is blunt about why: the persistent index is the product, and without the volume every run re-downloads and re-transcribes.

bash
docker run --rm -v watch-skill-data:/data ghcr.io/oxbshw/watch-skill --help
docker run --rm -i -v watch-skill-data:/data ghcr.io/oxbshw/watch-skill serve

The image installs ffmpeg but not the Playwright browser, because Chromium plus system libraries roughly doubles the image. The Dockerfile says everything except THE LOOP works as-is and that doctor reports the gap. Capturing browser sessions in the container means extending the image and running playwright install --with-deps chromium.

Where Watch Skill is the wrong tool

The extras are the first real constraint, and the README does not soften it. A bare install produces a CLI that fails at perceive.missing_dependency the moment you point it at a video. That is a deliberate design choice, keeping the default install small, but it means the failure arrives at first use rather than at install time. Read the extra list before you run pip.

The second constraint is storage and time. The Dockerfile comment states that the persistent index is the product, and that without the volume every run re-downloads and re-transcribes. That is a direct statement about cost: if you do not keep the index, you pay the extraction cost again on every run. The flip side is that keeping it means keeping frames, transcripts and lessons on disk, and the project gives no retention policy in the README.

The third is the browser path. The container ships the Playwright package but not its browser, so THE LOOP needs an extended image. If your work is entirely browser capture inside a container, you are maintaining a derived Dockerfile before you capture anything.

Finally, the verification verdict is not a general correctness proof. It reports on a contract you wrote and froze. If the contract is wrong, VERIFIED is a true statement about a bad check. The four-value verdict is more honest than a boolean, but it does not tell you whether the contract was worth writing.

How this differs from pointing a multimodal model at a file

The obvious alternative is to upload the recording to a model that accepts video and ask your question directly. The difference is what you get back. A model answer is prose you have to trust; Watch Skill's answer cites a timestamp you can open, because the source was indexed into timestamped artifacts first. For a one-off question about a short clip, the direct route is cheaper and simpler, and nothing in the README suggests otherwise.

The second alternative is a conventional transcription pipeline: speech-to-text, then a text index, then retrieval. That is a reasonable stack, and it is what most teams already have. It drops the visual channel entirely. Watch Skill's OCR extra exists precisely because on-screen text is not in the audio, and the README's own example question, "what changed at 3:12?", is a question a transcript-only index usually cannot answer.

The third alternative is agent observability built on logs and traces. Those record what the agent did. Watch Skill records what the agent saw, and the verification contract records whether the result held up. That is a different claim about the same run, and the two are complementary rather than competing.

Licence, maintenance and the cost of staying current

The licence is MIT, stated in pyproject.toml as license = { text = "MIT" } and listed as the classifier "License :: OSI Approved :: MIT License". That is permissive and carries no copyleft obligation on your own code. It also means no warranty, so the verification verdicts are your responsibility to interpret, not the maintainer's to stand behind. This is a description of the licence text, not legal advice; read LICENSE in the repository if the distinction matters to your organisation.

Maintenance is active as of the repository state. The last push to the default branch was on 2026-09-08, and the same date carries three releases: core-v1.4.3, deepwatch-v0.1.3 and deepwatch-v0.1.4. The version in pyproject.toml is 1.4.3, matching the core release. The repository is not archived.

Upgrade cost splits by package. The Python engine carries a versioned release and a lockfile (uv.lock) is present, so the dependency graph is pinned and reproducible. DeepWatch is at 0.1.x, which the project's own release numbering suggests is early; expect the workspace layer to move faster than the core. The Dockerfile resolves dependencies from the lockfile in a separate layer specifically so that source edits do not invalidate a roughly 600 MB install, which tells you the maintainers expect rebuilds to be expensive and planned for it. Running watch-skill doctor after an upgrade is the cheapest check the README offers.

Editorial conclusion

Adopt Watch Skill if you already run an MCP client and need an agent to cite a moment in a recording instead of paraphrasing it, or if you want a verification verdict that a separate process produces. Skip it if you only need a one-off transcript and never intend to re-query the source: the indexing step is the product, and without it you are paying for storage you do not use. Before wiring it into anything, run watch-skill doctor, confirm which extras you actually installed, and check that your MCP client can see the 39 tools after watch-skill serve.

Frequently asked questions

What does Watch Skill do?

It turns video, audio and screen activity into frames, transcripts and OCR text, each carrying an absolute timestamp, so an agent can answer questions with citations. Separately, it evaluates frozen contracts covering file digests, JSON values, SQL results, HTTP responses and DOM state, and returns VERIFIED, FAILED, UNVERIFIED or INCONCLUSIVE.

What is the Claude Code skill for watching videos?

The README lists Claude Code among the MCP clients that can use Watch Skill. You install the package with the standard extra and start the stdio MCP server with watch-skill serve, which the README describes as exposing 39 tools.

Can you give me some examples of agent skills?

Watch Skill is one: the README also gives npx skills add oxbshw/watch-skill -g to install the skills into more than 25 agents at once. The repository ships twenty numbered examples, from examples/01-watch-and-ask/ through examples/20-observer-loop/.

Why does watch-skill watch fail with perceive.missing_dependency?

The README states that a bare pip install watch-skill gives you the CLI, the verifier and the Bridge, but cannot extract a frame, so watch stops at perceive.missing_dependency on the first video. Installing with an extra such as watch-skill[standard] adds frames, retrieval and the MCP server, and watch-skill doctor names the exact command for whatever is still missing.

Can I run Watch Skill without installing its dependencies on my machine?

Yes. The Dockerfile supports running the image with a mounted volume, for example docker run --rm -v watch-skill-data:/data ghcr.io/oxbshw/watch-skill --help. The Dockerfile comments note that the volume matters because the persistent index is the product, and without it every run re-downloads and re-transcribes.

Official sources

  1. License: MIT
  2. oxbshw/watch-skill on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/oxbshw-watch-skill.svg)](https://hysenlabs.com/projects/oxbshw-watch-skill)