CLI tool
TheJoeFin/Text-Grab avatar
TheJoeFin/Text-Grab

Text Grab: OCR for Windows That Never Leaves Your Machine

Use OCR in Windows quickly and easily with Text Grab. With optional background process and notifications.

5,032 stars332 forksC#MIT

At a glance

What is it?
Text Grab is a Windows OCR utility from TheJoeFin that captures text from screens, images and PDFs using the built-in Windows OCR API, with four capture modes and a background process for global hotkeys. It is local-only, MIT-licensed, and shipped through the Microsoft Store, GitHub Releases, scoop and choco.
Who is it for?
Adopt Text Grab if you work on Windows 10 or later and regularly need text out of screenshots, PDFs or apps that block selection; the local OCR path means no upload step and no per-use cost, and the Edit Text Window saves a round trip to another editor. Skip it if you are not on Windows, since every feature depends on the Windows OCR API, or if you need a scriptable OCR pipeline rather than an interactive capture tool.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 8 days ago.
What is it written in?
Mainly C#, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem Text Grab solves, and who hits it

Text lives in places Windows will not let you select it. A screenshot pasted into a chat, a scanned PDF, a video frame, a licence key rendered inside an installer, a chart label in a dashboard. The README frames the whole project around that gap: text trapped inside images, videos, PDFs and parts of apps where you cannot select it. Text Grab takes a screenshot or opens a supported file, runs it through OCR, and sends the result to the clipboard or into an editor.

The audience is narrow and specific. This is a Windows desktop tool, not a cross-platform library. The README lists Windows 10 or later as a requirement for all features that use the Windows OCR API, and Windows 11 on a Copilot+ PC with the Microsoft Store install for the Windows AI features that use the on-device Neural Processing Unit. If you are on macOS or Linux, nothing here applies to you.

The second audience is people who already copy text from images occasionally but resent the round trip. Cloud OCR services charge per use and require uploading the image. Text Grab runs entirely on the device, which the README states plainly: no internet connection, no cloud service, no per-use cost. For anyone handling contracts, medical screenshots or internal dashboards, that removes a review step before you can even paste the text.

Four capture modes built on one OCR engine

The architecture is a WPF application (the repository topics list wpf alongside dotnet and msix) wrapping the Windows OCR API. That single dependency explains most of the behaviour.

Full-Screen Mode lets you select any region of the screen and copies the recognized text straight to the clipboard. Because the Windows OCR API draws a bounding box around each recognized word, a single click can target one word. If you click or select an area with no text, the Text Grab window stays active so you can retry rather than dismissing and restarting. Escape, right-click then Cancel, or Alt+F4 exits. The README notes this mode is the basis of the PowerToys Text Extractor, which is a useful signal about the recognition approach even though the two are separate products.

Grab Frame Mode is a movable OCR window you park over part of the screen. You grab text by searching for it, clicking a word border, or clicking the Grab button. It uses the same engine as Full-Screen Mode, so accuracy characteristics are identical; the README suggests adjusting the size and position of the frame as the practical accuracy lever. That is honest about the constraint: you are tuning the input, not the model.

The Edit Text Window is where the tool stops being a clipboard helper. It offers plain text, spreadsheet-style and markdown modes, plus cleanup operations: collapse to a single line, toggle UPPERCASE/lowercase/Titlecase, trim spaces and empty lines, remove duplicate lines, replace reserved characters, and extract by pattern (phone numbers, emails, custom patterns). Search and extract covers find and replace, regular expressions, and launching URLs. Structure tools convert stacked data to table format and transpose captured tables when OCR returns the right data in the wrong orientation. A Calc Pane evaluates expressions line by line, supporting arithmetic, functions such as sin, cos, sqrt, abs and log, variables and constants like pi and e, and unit conversions written as 5 miles to km.

The fourth surface is the background process. The README describes enabling it so global hotkeys work anywhere in Windows, and notes you can also open specific modes from the command line. That is the difference between a utility you launch and one that is always one keystroke away.

Installing Text Grab and running a first grab

The official channels are the Microsoft Store and GitHub Releases. Community packages exist for scoop and choco. Pick one; they are alternative distribution paths, not steps in a sequence.

With scoop, the README gives this exact command:

bash
scoop install text-grab

With Chocolatey, the equivalent is:

bash
choco install text-grab

After installing, launch Text Grab from the taskbar. For a first real use, open Full-Screen Mode, drag a rectangle over any text on screen, and release. The recognized text lands on the clipboard. If the region contains no text, the window stays open so you can drag again; press Escape to leave.

If you want the tool available without launching it each time, enable the background process so the global hotkeys respond anywhere in Windows. The README does not document which hotkeys are bound by default, so check the application settings after enabling it rather than assuming.

Building from source is a separate track and only makes sense if you intend to modify the project. The README requires the .NET 10 SDK and notes the repository pins SDK 10.0.100 in global.json. The commands are:

bash
dotnet restore Text-Grab.sln
dotnet build Text-Grab\Text-Grab.csproj
dotnet test Tests\Tests.csproj

For Visual Studio, the README specifies the Universal Windows Platform development, .NET desktop development and .NET cross-platform development workloads, plus Windows 10 SDK 10.0.22621.0, with Text-Grab-Package as the startup project and the CPU target set to x64 or ARM64.

Where Text Grab fails, and when it is the wrong tool

OCR accuracy is the first limit, and the README says so directly: OCR is not perfect, and the suggested remedy for Grab Frame is to adjust the frame's size and position. There is no accuracy setting, no model swap, no confidence threshold exposed in what is documented. You are working with the Windows OCR API's output and cleaning it afterwards.

The second limit is the hard platform dependency. Windows 10 or later is required for all features using the Windows OCR API. Windows AI features additionally require Windows 11 on a Copilot+ PC and the Microsoft Store install, which means the higher-accuracy NPU path is unavailable to a large part of the Windows install base and to anyone using the GitHub release on a machine that would otherwise qualify. That is a real fork in the distribution, not a footnote.

Third, this is an interactive capture tool, not an automation library. The README mentions opening specific modes from the command line, but the documented workflow is human-driven: select a region, park a frame, clean up results in a window. If you need to batch-process ten thousand scanned pages on a server, Text Grab is the wrong shape entirely. The same goes for anyone who needs OCR inside a CI pipeline or a Linux container.

Finally, the clipboard is a shared resource. Watch clipboard for changes is listed as a workflow helper, which is useful, but it also means the tool observes clipboard content. The README does not describe what happens to that data or whether it is persisted, so treat that as unverified rather than assuming either way.

Text Grab against PowerToys Text Extractor and cloud OCR

The most direct comparison is PowerToys Text Extractor, and the relationship is unusual: the README states that Full-Screen Grab mode is the basis of PowerToys Text Extractor. So the recognition core is shared. The difference is scope. Text Extractor is a single-purpose region grab inside the PowerToys suite. Text Grab adds Grab Frame, the Edit Text Window with spreadsheet and markdown modes, pattern extraction, table transposition, the Calc Pane, and a background process for global hotkeys. If you already run PowerToys and only ever want the region grab, Text Extractor covers it and you avoid a second utility. If you want the cleanup and structuring layer, that is what Text Grab adds.

The other comparison is cloud OCR. Services in that category accept an image over the network and return text, which usually brings better recognition on difficult material and works from any platform. The trade is that the image leaves your machine and usage is metered. Text Grab inverts both: the README states all OCR runs entirely on the device, no internet connection and no per-use cost, and the price is that you are limited to what the Windows OCR API can do, on Windows only. Neither is strictly better. The choice is whether your material can leave the device and whether you need accuracy beyond what the local engine provides.

Maintenance, licensing and upgrade cost

The repository is not archived and the last push was on 2026-09-22, one day before the date used for this assessment. Recent releases show a steady cadence: v4.14.2 on 2026-06-20, v4.15.0 on 2026-08-01 with Smart Patterns, Description, Text to Speech, HDR and settings changes, and v4.16.0-beta1 on 2026-09-19. The beta tag matters: v4.16.0 is not yet a stable release, so if you want the newest features you are opting into pre-release software.

Upgrade cost is low for the Store and package-manager paths, since those handle versioning. It is higher if you build from source: the repository pins SDK 10.0.100 in global.json, so a machine with a different .NET 10 patch level may need adjustment, and the Visual Studio path requires a specific Windows 10 SDK version and three workloads. Budget that setup time if you plan to modify the code.

Licensing is MIT, per the repository. That is permissive and places few obligations on reuse beyond preserving the licence notice. ThirdPartyNotices/ exists at the top level, which is where dependency attributions live; read it before redistributing a build. This is a description of the licence, not legal advice, and the MIT text in LICENSE is the authoritative document.

Editorial conclusion

Adopt Text Grab if you work on Windows 10 or later and regularly need text out of screenshots, PDFs or apps that block selection; the local OCR path means no upload step and no per-use cost, and the Edit Text Window saves a round trip to another editor. Skip it if you are not on Windows, since every feature depends on the Windows OCR API, or if you need a scriptable OCR pipeline rather than an interactive capture tool. Before committing, verify that the Windows OCR language pack for your target language is installed, and check whether you want the Store build (which the README ties to Windows AI features on Copilot+ PCs) or the GitHub release.

Frequently asked questions

What is Text Grab?

Text Grab is a Windows application that captures text with OCR, lets you clean it up, and moves it into your workflow. It runs OCR entirely on your device using the Windows OCR API, with no cloud service and no per-use cost.

How can I use OCR in Windows with Text Grab?

Install Text Grab from the Microsoft Store, GitHub Releases, scoop or choco, then launch it and use Full-Screen Mode to select a region of the screen. The recognized text goes straight to your clipboard. Windows 10 or later is required for all features that use the Windows OCR API.

How do I grab text from the screen with Text Grab?

Use Full-Screen Mode to select any region and copy the recognized text to the clipboard, or click once to target a single word, since the Windows OCR API draws a bounding box around each recognized word. Grab Frame Mode instead gives you a movable OCR window you position over the text.

Can I grab text from an image with Text Grab?

Yes. The README describes taking a screenshot or opening a supported file, running it through the OCR engine, and sending the result to the clipboard or into an editor. The Edit Text Window also lists copying text from every image in a folder.

Is there a Text Grab alternative for Windows?

PowerToys Text Extractor is the closest one, and the README states that Text Grab's Full-Screen Grab mode is its basis, so the recognition core is shared. Text Grab adds Grab Frame, the Edit Text Window with spreadsheet and markdown modes, pattern extraction, the Calc Pane and a background process for global hotkeys.

How do I use Text Grab?

Launch it and pick a mode: Full-Screen Mode to select a region of the screen, or Grab Frame Mode for a movable OCR window you position over the text. Results go to the clipboard or into the Edit Text Window, where you can clean and restructure them.

Official sources

  1. License: MIT
  2. Project website
  3. README
  4. Releases
  5. TheJoeFin/Text-Grab on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/thejoefin-text-grab.svg)](https://hysenlabs.com/projects/thejoefin-text-grab)