CLI tool
brijr/iris avatar
brijr/iris

iris and the Chrome it borrows instead of shipping

Screenshots of live websites. Minimal interface, powerful engine.

476 stars24 forksRustMIT

At a glance

What is it?
A screenshot tool for coding agents that renders with the Chrome you already have, driven over the DevTools Protocol, and serves the same capture engine to agents as a single MCP tool returning pixels inline. The benchmark section is the most careful part of the documentation.
Who is it for?
It fits an agent workflow where the model needs to see a page rather than be told what is on it, and where the same capture has to be reproducible from a terminal, from a script and from an MCP call. It does not fit browser automation, since interaction scripting, diffing and review workflows are explicitly out of scope, and it does not fit a machine without a Chrome-family browser.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 13 days ago.
What is it written in?
Mainly Rust, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.

Editorial analysis

One command, one image, rendered by your own Chrome

The stated pitch is a camera for coding agents, where one fast command or one MCP tool call produces one trustworthy image. What makes that cheap is that iris does not ship a browser. It renders with your real installed Chrome, driven over the DevTools Protocol, and the only runtime dependency is a Chrome-family browser you already have: Chrome, Chromium, Edge or Brave. Waiting is the other half of the trustworthiness claim. Iris waits for fonts, for image loads and for entrance animations before capturing, and on a full page capture it also scroll-triggers lazy-loaded content so that sections below the fold exist before the shutter. The everyday invocations are one line each, with the flags doing the framing work:

code
iris example.com
iris --full --dark tailwindcss.com
iris --size iphone stripe.com
iris --selector '#hero' --padding 24 app.dev

The default capture is a 1440 by 900 viewport at twice scale, written next to your prompt as a PNG named after the host. The rest of the surface covers the cases that come up in scripts rather than in demos: a batch of URLs captured concurrently into a directory, a list piped in on standard input where comment lines are tolerated, an explicit output extension to choose the format, a wait for a selector to appear before shooting, and a machine-readable mode that emits one JSON object per capture.

code
iris -o shots/ a.com b.com c.com
cat urls.txt | iris - -o shots/
iris -o hero.jpg --wait-for 'h1' app.dev
iris --selector 'h1' --json app.dev

Element capture takes one selector and the first match

Element capture is CSS-selector based on purpose, and the limits are stated rather than discovered later. Iris captures the first element matching the selector in document order, scrolling it into view and letting newly visible content settle before it frames the shot. Two flag relationships follow from that design. Padding only makes sense around a selected element, so it requires a selector, and a selector cannot be combined with a full page capture since those are contradictory requests. Cross-origin iframe contents are not supported, and capturing every match is not offered, so a page with repeated components gives you the first one. Each capture reports what it actually did, in a single JSON object per image:

json
{"status":"ok","url":"https://example.com/","output":"/absolute/example.com.png","mode":"element","selector":"h1","padding":24,"css_width":180,"css_height":72,"scale":2.0,"format":"png","bytes":14231}

The CSS dimensions, the scale, the format and the byte count are all in that line, which is what makes an agent's claim about a screenshot checkable rather than decorative.

Viewport presets, and an automatic retreat from the 16k limit

Viewports are handled by presets so that a capture is comparable across runs. The size flag accepts explicit dimensions or one of three named presets: desktop at 1440 by 900 at twice scale, iPhone at 390 by 844 at three times scale, and iPad at 1024 by 1366 at twice scale. The iPhone preset also switches to a mobile user agent, which is the part that actually changes what the server returns rather than just how the image is scaled. Output is retina by default, and there is one automatic exception worth knowing: a full page taller than Chrome's roughly sixteen-thousand-pixel render limit falls back to single scale instead of failing, and the report tells you which scale you got rather than silently resizing. Format follows the output file extension when it is a recognised one, otherwise the format flag decides between PNG as the default, jpg and webp, with JPEG and WebP encoded at quality 90. Two flags adjust timing and appearance rather than geometry: the dark flag emulates a dark colour scheme preference so you can capture the dark variant of a site without switching your own system theme, and the wait flag adds an extra settle delay on top of the smart waiting that already covers fonts, images and animations.

The MCP server returns pixels instead of a file path

The same binary serves an MCP server over stdio, which is the part aimed at agents. Registration for Codex is one command:

code
codex mcp add iris -- iris mcp

Other clients use the equivalent command and arguments configuration, and the server exposes exactly one tool, capture, which returns the image inline with structured metadata. Two behaviours matter for how an agent should be prompted. It writes nothing by default, so a file only appears if you pass an output path, which means an agent can look at a page without littering the working directory. And the address handling is opinionated in a useful way: bare localhost, .localhost and loopback addresses are treated as HTTP automatically, while other bare hosts are assumed to be HTTPS, so a developer server works without ceremony and a public host does not silently downgrade. A tool request carries the same options as the flags, from url and selector through padding, size, dark, format and a timeout in seconds, with output added when a file is wanted. If the server picks the wrong browser, its help subcommand is where a Chrome binary gets selected. And the registration order matters in practice: some MCP clients need a fresh agent task before the newly added capture tool becomes visible.

Batch mode will not die because one URL is broken

Batch behaviour is defined by what happens when something fails. Pointing the output flag at a directory switches to batch mode, multiple URLs on one line are captured concurrently, and a list can be piped in on standard input with comment lines tolerated. One browser process serves concurrent tabs rather than one process per page, and a URL that fails prints a failure marker without taking the rest of the batch down; the exit code is 1 if anything failed, so a script can still tell the difference between a clean run and a degraded one. Output filenames are derived from the URL, so a capture of a pricing section on a known host becomes a predictable name, and collisions get numeric suffixes rather than overwriting each other. Machine-readable mode writes one JSON object per completed capture to standard output in completion order rather than input order, and a capture failure is also emitted as JSON, so a consumer never has to parse a human-facing error line.

The benchmark section separates CLI cost from MCP cost

Most projects publish one number. This one explains why there are two. The CLI starts a Chrome process for every invocation, while the MCP server keeps a single Chrome process alive for its whole lifetime, so the two paths cost different amounts and have to be measured separately. The methodology is spelled out: compare the same URL, capture mode, viewport, scale, Chrome version and hardware; prefer a deterministic local page when comparing releases because public URLs add network and server variance; and for the MCP path, initialize once, record the first capture, then report the median and p95 of at least ten identical subsequent calls, measured from JSON-RPC request to complete tool response with model time excluded. The repeatable command uses hyperfine against a local fixture:

sh
cargo build --release
hyperfine --warmup 1 --runs 10 \
  './target/release/iris file://$PWD/tests/fixtures/precise-capture.html \
    --selector .capture-target --padding 24 --scale 1 \
    -o /tmp/iris-benchmark.png'

The published reference figures, for version 0.4.1 on an Apple M2 Max with a specific Chrome build, are a one-shot CLI median of 1.00 second, a first MCP capture of 965 milliseconds, and then 366 milliseconds median with 383 milliseconds p95 for the next ten. They are labelled reference results rather than a performance guarantee, and the documentation asks you to record the iris version, the Chrome version and the exact command whenever you publish a number of your own.

What Iris deliberately leaves out

The scope statement is short and worth reading before you build on it: the MCP subcommand serves a single-image capture tool over stdio, and batch capture, browser interaction scripting, diffs and review workflows remain outside what the tool does. That is a narrower product than a browser automation suite, and the choice shows up in the dependency list too, where the browser driver and the MCP server library are present but no scripting or comparison framework is. The setup prompt the project hands to your coding agent encodes the same boundaries: its goal is to install iris, connect the local MCP server, and prove that both the CLI and MCP paths work, with an explicit instruction not to modify the application repository. The repository itself is small, with the source, a tests directory holding the fixture the benchmark uses, an install script, a committed lock file, and a release profile that strips symbols and enables link time optimisation. Building needs Rust 1.88 or newer, and the crate is published under a different name than the command it installs. The dependency list matches that narrow scope closely, with a browser automation library driving Chrome, an MCP server library on standard input transport, an argument parser, a base64 encoder for inlining images, and an async runtime, and nothing that suggests a comparison or scripting layer. Versions move in small steps: three tags landed within about half an hour of each other in mid-August, and the last push to the main branch was on 2026-09-18.

Editorial conclusion

It fits an agent workflow where the model needs to see a page rather than be told what is on it, and where the same capture has to be reproducible from a terminal, from a script and from an MCP call. It does not fit browser automation, since interaction scripting, diffing and review workflows are explicitly out of scope, and it does not fit a machine without a Chrome-family browser. Before you trust any timing from it, read the benchmarking rules and record the version, Chrome build and exact command alongside your own numbers, because the published figures are labelled reference results rather than a guarantee.

Frequently asked questions

What does iris need installed before it will work?

A Chrome-family browser, meaning Chrome, Chromium, Edge or Brave, and nothing else at runtime. Building from source needs Rust 1.88 or newer, and the package on crates.io is named iris-screenshot while the installed command is iris.

How do I connect iris to a coding agent?

Run iris mcp, which serves a single capture tool over stdio, and register it with your client. For Codex that is codex mcp add iris -- iris mcp, and other clients use the equivalent command and arguments configuration.

Does the iris MCP tool write a file when an agent uses it?

Not by default. The capture tool returns the image inline with structured metadata and writes nothing unless you pass an output path.

What does iris not support for element capture?

It captures the first selector match in document order, so capturing every match is not supported, and cross-origin iframe contents are not supported either. Padding requires a selector, and a selector conflicts with full page capture.

How should I benchmark iris?

Compare the same URL, capture mode, viewport, scale, Chrome version and hardware, and prefer a deterministic local page over a public URL. The CLI starts Chrome per invocation while MCP keeps one Chrome alive, so measure them separately.

Official sources

  1. brijr/iris on GitHub
  2. Issues
  3. License: MIT
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/brijr-iris.svg)](https://hysenlabs.com/projects/brijr-iris)