CLI tool
simonw/shot-scraper avatar
simonw/shot-scraper

shot-scraper: a Python CLI for screenshots, video demos and JavaScript scraping

A CLI utility for taking screenshots of websites, recording video demos and scraping sites using JavaScript

2,581 stars128 forksPythonApache-2.0

At a glance

What is it?
shot-scraper wraps Playwright in a command-line tool that captures web pages, records video demos and extracts data with JavaScript. It suits documentation pipelines and scripted captures, not interactive browsing.
Who is it for?
Adopt shot-scraper when you need repeatable, scriptable page captures in a Python environment or a CI job, and you are willing to run shot-scraper install to fetch the browser it drives. Skip it if you want an interactive browser tool or a visual editor, since the CLI has no GUI and the README does not describe one.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 17 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 24, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What shot-scraper solves, and who reaches for it

Documentation screenshots go stale the moment a UI changes, and doing them by hand does not scale past a handful of pages. shot-scraper turns that chore into a command. The README describes it as "a CLI utility for taking screenshots of websites, recording video demos and scraping sites using JavaScript", and the examples it lists are exactly that kind of work: the Datasette documentation generates its screenshots through a shot-scraper GitHub repository, and a Twitter bot built by Ben Welsh uses it with GitHub Actions to capture news homepages. The tool is aimed at people who already think in shell commands and Python packaging: technical writers wiring screenshots into a docs build, maintainers who want a captioned image regenerated on every release, and data people who need a number out of a page that only renders correctly in a real browser. It is not aimed at someone who wants to click around and crop an image by hand.

Playwright underneath, click and YAML on top

The dependency list in pyproject.toml tells you most of the architecture. shot-scraper depends on playwright>=1.62.0, which supplies the actual browser automation, plus click and click-default-group for the command surface, PyYAML for configuration files, and pydantic for validating what goes into them. The package itself lives in the shot_scraper/ directory and registers one entry point, shot-scraper, mapped to shot_scraper.cli:cli. So the data flow is: you type a command, click parses it, shot-scraper translates it into Playwright calls, a real browser loads the page, and the result is written to disk as an image or printed as text. That indirection is the whole value. Playwright is a large Python API with its own concepts of contexts, pages and selectors; shot-scraper reduces the common cases to flags and a YAML file. The cost is that anything shot-scraper does not expose still requires you to drop into Playwright directly, and the README does not describe an escape hatch for that.

Installing shot-scraper and taking a first screenshot

Installation is two commands, and the second one matters. The first installs the Python package from PyPI; the second downloads the browser binary that Playwright drives. Skipping the second step leaves you with a CLI that cannot launch anything.

bash
pip install shot-scraper
# Now install the browser it needs:
shot-scraper install

Once that finishes, point the tool at a URL. The README uses the Datasette homepage as its example, and notes that the output filename is derived from the URL with dots replaced by hyphens.

bash
shot-scraper https://datasette.io/

According to the README, that command "will create a screenshot in a file called datasette-io.png". If you see that file appear in your working directory, the install is complete. The README points to a Taking a screenshot page in the full documentation for the many options beyond this default, so treat the bare command as a smoke test rather than the full feature set. For teams that would rather not install anything locally, the README also describes shot-scraper-template, a GitHub template repository that takes screenshots of a page using shot-scraper.

Where shot-scraper is the wrong tool

The browser dependency is the first real constraint. playwright>=1.62.0 is not a small package, and shot-scraper install fetches a browser build on top of it. In a locked-down CI image or a container where you cannot download browser binaries, the tool will not run, and the README does not document an offline or bring-your-own-browser path. The second constraint is scope. This is a capture and extraction tool, not a testing framework: it does not assert that a page rendered correctly, it does not diff two screenshots, and it does not fail a build when a selector goes missing. If your actual problem is regression testing a UI, you want Playwright's own test runner or a visual-diff service, and shot-scraper will only get you the raw images. The third case is interactive work. There is no GUI, no element picker and no live preview described in the README, so if you need to point at a region and crop it visually, the command line is the wrong interface. Finally, the README does not document rollback or a dry-run mode, so a misconfigured batch of captures is something you discover after the files are written.

shot-scraper against raw Playwright

The obvious alternative is Playwright itself, in Python or in its Node.js form. The difference is one of level, not capability. Playwright gives you a browser object, a context, pages, selectors, network interception and a test runner; shot-scraper gives you a command that takes a URL and writes a PNG. If your job is a single screenshot in a shell script, the raw Playwright equivalent is a dozen lines of Python plus your own filename logic, and shot-scraper removes that boilerplate. If your job is a suite of assertions across a logged-in session, shot-scraper is the wrong layer and you will end up rebuilding Playwright on top of it. The YAML configuration is the other dividing line: shot-scraper's PyYAML dependency exists so that a list of pages and capture settings can live in a file next to your repository, which is a documentation-workflow concern rather than a browser-automation one. Pick shot-scraper when the deliverable is an image or a scraped value; pick Playwright when the deliverable is a verdict about whether the page behaved.

Licence, releases and what maintenance costs you

shot-scraper is Apache-2.0, declared both in the README badge and in the licence field of pyproject.toml, with the full text in the LICENSE file at the repository root. Apache-2.0 is a permissive licence that includes an explicit patent grant, which is the usual reason organisations prefer it over MIT for tooling they may redistribute; this is a description of the licence terms, not legal advice, and you should read LICENSE and your own policy rather than take my summary. The project requires Python 3.10 or newer, so an older interpreter is a hard blocker before anything else. Maintenance looks steady rather than dormant: the last push was on 2026-09-13, and release 1.12 carries the same timestamp, following 1.11 on 2026-07-12 and 1.10 on 2026-06-30. The upgrade cost is concentrated in the Playwright pin. Because pyproject.toml requires playwright>=1.62.0, a pip install can pull a newer Playwright than you tested with, and since shot-scraper install downloads the matching browser, a version bump can change both the library and the browser binary at once. Pin your dependencies if you care about reproducible captures.

Editorial conclusion

Adopt shot-scraper when you need repeatable, scriptable page captures in a Python environment or a CI job, and you are willing to run shot-scraper install to fetch the browser it drives. Skip it if you want an interactive browser tool or a visual editor, since the CLI has no GUI and the README does not describe one. Before committing, verify that playwright>=1.62.0 installs cleanly in your environment and that a plain shot-scraper https://datasette.io/ run produces datasette-io.png as the README states.

Frequently asked questions

How do I take a screenshot with shot-scraper?

Install the tool with pip install shot-scraper, run shot-scraper install to fetch the browser, then pass a URL such as shot-scraper https://datasette.io/. The README states this creates a file named datasette-io.png.

Can shot-scraper take screenshots automatically?

Yes. The README describes shot-scraper-template, a GitHub template repository that creates a repository taking screenshots of a page using shot-scraper, and the Datasette documentation generates its screenshots through a shot-scraper GitHub repository.

What Python version does shot-scraper need?

pyproject.toml sets requires-python to >=3.10, so Python 3.10 or newer is required. It also depends on playwright>=1.62.0, click, PyYAML and pydantic.

What licence is shot-scraper released under?

Apache-2.0. The README shows the Apache 2.0 badge and pyproject.toml declares license = "Apache-2.0", with the full text in the LICENSE file at the repository root.

Does shot-scraper only take screenshots?

No. Its description covers taking screenshots of websites, recording video demos and scraping sites using JavaScript, and the README links a scrape-hacker-news-by-domain project that uses shot-scraper javascript to scrape a page.

Official sources

  1. License: Apache-2.0
  2. Project website
  3. README
  4. Releases
  5. simonw/shot-scraper on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/simonw-shot-scraper.svg)](https://hysenlabs.com/projects/simonw-shot-scraper)