# OCR It: a pinned region, a hotkey, and a point instead of a selector

> OCR It is a Chrome and Firefox extension for documents trapped in a viewer: pin a screen region once, then every hotkey press screenshots that rectangle and OCRs it offline with a bundled Tesseract build. The interesting parts are the permissions model, the point-based page turning, and the one viewer it cannot drive.

**thiagotigaz/ocr-it** — Chrome extension: pin a screen region once, then hotkey your way through a paginated document. OCR runs 100% offline via bundled Tesseract.

- Repository: https://github.com/thiagotigaz/ocr-it
- Website: https://thiagotigaz.github.io/ocr-it/
- Stars: 410 · Forks: 28
- Language: JavaScript
- License: MIT
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/thiagotigaz-ocr-it

## The repository root is already a Chrome extension, Firefox needs a build

One source tree, two targets. `npm run build` writes a loadable directory per browser into `build/`, and for Chrome the build step is optional: the repository root is a valid Chrome extension as checked in, so Load unpacked can point straight at the clone.

```sh
git clone https://github.com/thiagotigaz/ocr-it.git
cd ocr-it
npm run build          # -> build/chrome, build/firefox
```

Firefox has no equivalent shortcut. You go to `about:debugging#/runtime/this-firefox` and pick `build/firefox/manifest.json`, so the build is what produces the file you load. That route installs a temporary add-on, and a temporary add-on is unloaded when Firefox quits; only Mozilla-signed add-ons install permanently, which is what the store listing is for. With Mozilla's own tooling, `npm run start:firefox` opens a scratch profile with the extension already installed and reloads it on every edit.

So Chrome is one click from a store listing or three clicks from a clone, while Firefox needs a build and a profile that resets on every restart.

## Chrome can silently leave your hotkey blank

Three hotkeys do the work: ⌥⇧S captures the region once, ⌥⇧A starts or stops an automatic run, and ⌥⇧R draws or redraws the region. None of them is guaranteed to arrive.

After installing on Chrome you are told to check `chrome://extensions/shortcuts`, because Chrome silently leaves a hotkey blank when something else already claims it. Nothing inside the extension can report that; the shortcut simply does not exist until you assign one. Firefox has no navigable shortcut editor at all, so the popup's hotkey buttons point you at about:addons, the gear icon, then Manage Extension Shortcuts.

The region is adjustable before you commit it. Draw it with ⌥⇧R, drag it around, pull the handles, or nudge it a pixel at a time with the arrow keys, holding ⇧ to resize, and ↵ keeps it. One instruction matters more than the controls: draw a little inside the text margins, because everything inside the rectangle gets read, page numbers and running headers included.

## No site access at install, two grants when you need them

The extension asks for no site access when it installs. Single captures ride on `activeTab`, the permission the browser hands over at the moment you press the hotkey or open the popup, so the common case needs nothing else.

Two features need a durable grant, and the project names both: an automatic run that outlives a page load, and turning pages inside a cross-origin iframe. When either matters, the popup offers an Allow button for the site you are on. Firefox exposes the same grant under about:addons, then OCR It, then Permissions.

That boundary explains a behaviour that otherwise looks arbitrary. A manual capture works on any page without a standing grant, but an unattended run that has to survive a navigation, or a run driving a reader embedded in another origin, cannot work on a transient one. It is the difference between the hotkey and the loop.

The offline claim rests on the same arrangement. OCR runs locally against a bundled Tesseract build, the extension makes no outbound requests at all, and PRIVACY.md sits at the top of the repository next to LICENSE and PUBLISHING.md.

## A stored point, not a selector, is what turns an embedded reader

Page turning has two modes. Pick control asks you to click the viewer's next-page button and stores a point, not a CSS selector. The key mode dispatches a keyboard event, ArrowRight by default, into whichever frame owns the middle of your capture region, so the reader receives it rather than the host page.

The point over a selector is a deliberate choice with two concrete payoffs. A stored point survives the DOM re-renders that routinely invalidate a selector, and it reaches two places a selector cannot: cross-origin iframes, where nothing the top frame can express addresses an element inside one, and shadow DOM, which `document.querySelector` cannot see into.

The mechanics get more involved from there. At advance time the point is offered to every frame and the one that owns it acts. A frame computes its position inside the top-level viewport by walking up its same-origin ancestors, and across an origin boundary the parent hands the offset down by `postMessage`, because `window.screenX` reports the browser window rather than the frame. The owning frame then resolves the point through any shadow roots, walks up to the nearest real control, and emits the whole pointerdown, mousedown, pointerup, mouseup and click sequence, so viewers that page on pointerdown behave like those waiting for click.

## Chrome's own PDF viewer is declared unreachable

Every turn attempt records a verdict, shown in the popup and as an on-page toast, and one of those verdicts is a refusal rather than a retry.

`no next-page control picked yet` means auto-advance is switched on and nothing has been picked. `an embedded viewer owns that point` means the control you marked belongs to Chrome's PDF viewer or to a plugin, and the documentation states plainly that this is unreachable by any extension. That is the hard ceiling of the whole approach: the extension can drive a page it is allowed to script, and Chrome's built-in viewer is not one of them. A scanned book inside a web reader works; the same book opened in the browser's own viewer does not.

Test now exists for exactly this reason. It fires an advance immediately, without capturing a page, and reports what happened, so you learn in one click whether the viewer will turn before committing to a long unattended run. Once it works, ⌥⇧A loops capture, turn, capture, turn until the document ends, and Esc on the page stops it.

## The only automatic quality flag is DUPLICATE

Recognition is not scored. What you get instead is evidence: every page is listed with a thumbnail of exactly what was cropped, so a region that drifted out of alignment is obvious at a glance rather than eighty pages later. Text is editable in place, and a single bad read can be re-run on its own without recapturing the page.

The one automatic signal is a marker rather than a score. A page flagged `DUPLICATE` had text identical to the one before it, which is nearly always a sign that the document never actually turned. That is a useful failure detector precisely because it fires on the most common way an unattended run goes wrong, and it costs nothing to compute. It is not a confidence measure, so a page that reads badly without repeating itself still needs a human to notice it in the thumbnail.

Export is deliberately plain. Copy all and Download .txt emit the pages in order with `--- page N ---` separators between them, which is the format you would paste into a summariser.

## npm install is for the tests, not for the extension

Everything the extension needs at runtime is committed, including the vendored Tesseract build under `vendor/`, which is what makes the offline promise work without a package install. `npm install` is only for the tests, the Firefox linter, or re-vendoring Tesseract.

The scripts show how much tooling sits around a small extension. `test` and `test:headed` run `node tools/e2e.mjs` with or without a visible browser, `fixture` generates test material through `node tools/fixture.mjs`, `build` runs `node tools/build.mjs`, and `lint:firefox` calls web-ext against `build/firefox`. There are separate entries for icons, screenshots, packaging and vendoring, and one for benchmarks. package.json is marked private and typed as an ES module, so none of it is published to a registry.

The top level matches that split: manifest.json and package-lock.json for what ships, src/ and vendor/ for the code and its bundled engine, tools/ for the build and test harness, and docs/ for the rest.

## The benchmark script exists, the figures do not

One script is called bench and points at `tools/bench/run.mjs`, and the header of the README links a benchmark page hosted at thiagotigaz.github.io/ocr-it/bench/. The README quotes no accuracy figure, no timing figure and no comparison of its own.

That absence matters more here than in most projects, because the workflow ends with handing a transcript to a model. If the OCR drops characters, merges two columns, or reads a running header as body text, nothing downstream will tell you where the damage started; the transcript simply looks authoritative. The thumbnail strip and the editable text are the mitigation, and both are manual.

Release v0.3.0, tagged with Firefox support, is the newest release, dated 2026-08-25, and package.json carries the same 0.3.0. The last commit to the default branch is dated 2026-08-28, so the Firefox target is one release old and the tree has moved a little further since.

## Conclusion

OCR It earns a place if your document is a scanned book or a slide deck inside a viewer you control, because the pinned region, the offline Tesseract build and the point-based page turning solve exactly that job. It will not help with Chrome's own PDF viewer, which the project itself calls unreachable by any extension, and the region has to be drawn inside the text margins or you will be reading running headers all afternoon. Before a long run, use Test now and check chrome://extensions/shortcuts, because Chrome can leave a claimed hotkey blank and every failed advance reports a verdict rather than failing quietly.

## FAQ

### Does OCR It send my document pages to a server?

No. OCR runs locally against a bundled Tesseract build, the extension makes no outbound requests at all, and no API key is needed. A PRIVACY.md sits at the top of the repository next to LICENSE.

### Why does OCR It not turn pages in Chrome's PDF viewer?

Because that viewer is unreachable by any extension. The verdict table names the case: an embedded viewer owns that point means Chrome's PDF viewer or a plugin owns the control you marked, so the advance reports that verdict instead of turning the page.

### What does OCR It store when I pick a next-page control?

A point, not a CSS selector. A point survives DOM re-renders and reaches cross-origin iframes and shadow roots, and at advance time it is offered to every frame while the frame that owns it acts.

### How do I install OCR It in Firefox for development?

Run npm run build so build/firefox exists, then open about:debugging#/runtime/this-firefox and load build/firefox/manifest.json as a temporary add-on, which is unloaded when Firefox quits. npm run start:firefox does the same in a scratch profile that reloads on every edit.

### Which browsers and versions does OCR It support?

Chrome 116 and newer from the Chrome Web Store, and Firefox 140 and newer from Firefox Add-ons. The newest release is v0.3.0, tagged with Firefox support, and package.json carries the same 0.3.0 version.

## Sources

- [License: MIT](https://github.com/thiagotigaz/ocr-it/blob/main/LICENSE)
- [Project website](https://thiagotigaz.github.io/ocr-it/)
- [README](https://github.com/thiagotigaz/ocr-it/blob/main/README.md)
- [Releases](https://github.com/thiagotigaz/ocr-it/releases)
- [thiagotigaz/ocr-it on GitHub](https://github.com/thiagotigaz/ocr-it)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/thiagotigaz-ocr-it
