Self-hosted service
arcships/light-ocr avatar
arcships/light-ocr

arcships/light-ocr: offline OCR for Node.js and C++ with PP-OCRv6

Fast, offline OCR for Node.js & C++. PP-OCRv6 with Core ML / WebGPU hardware acceleration, recognize text in images with confidence scores & coordinates. npm: @arcships/light-ocr.

509 stars42 forksC++Apache-2.0

At a glance

What is it?
light-ocr bundles PP-OCRv6, PDFium and a Chinese fallback font into one npm install, returning text, confidence and quadrilateral boxes. It is a good fit for local OCR inside Node.js apps and native C++ integrations, and a poor fit if you need Python or a hosted service.
Who is it for?
Adopt light-ocr if you are building a Node.js CLI, desktop app or service that must read text from images or PDFs without sending files off the machine, and you are on Node.js 22 or 24. Do not adopt it if you need a Python binding, a hosted API, or Japanese support in the smallest model tier, since the Tiny tier is documented as covering 49 languages without Japanese.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 5 days ago.
What is it written in?
Mainly C++, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

The problem light-ocr solves for Node.js and C++ teams

Running OCR locally usually means assembling several moving parts: a model file, an inference runtime, a PDF rasterizer if you care about documents, and a font that can render CJK glyphs before recognition. Each of those is a separate download, and each is a place where an install can fail on a machine with no network access or a locked-down CI runner.

light-ocr collapses that into one package. The npm package carries PP-OCRv6 Small, the OCR runtime, PDFium, and a checksum-pinned Noto Sans SC fallback font through its platform dependency. The README states that installation and runtime need no postinstall fetch, compiler, model download, PDF engine download, or font download. That is the specific claim worth testing in your own environment, because it is the whole reason to pick this over wiring the components together yourself.

The audience is narrow and stated: local OCR in Node.js apps, CLIs, desktop software, and native C++ integrations. If you are writing Python, this is not your library.

How the engine, execution modes and result schema fit together

The entry point is createEngine(), which picks an execution mode for the current platform. On macOS with Apple Silicon the README says Auto tries Core ML on macOS 15 or later and falls back to CPU. On Linux x64 with glibc it tries WebGPU through Vulkan, and on Windows x64 WebGPU through D3D12, both falling back to CPU. macOS on Intel, Linux arm64 and Windows arm64 go straight to CPU. Applications that want to override this can pass auto, cpu, apple or webgpu through the execution option.

Results are line-oriented rather than a single text blob. Each line carries the recognized text, a confidence score, and a quadrilateral box giving its position in the original image, and the README says lines come back in reading order. Pages add metadata and timing. For images you can call recognizeEncoded with file bytes, or recognize() with decoded pixels in GRAY8, RGB8, BGR8 or RGBA8 if your application already decodes images. EXIF orientation is corrected automatically.

Documents take a different path. recognizeDocument is an async iterator over pages, accepting a PDF path or an array of image buffers, with a dpi option. The CLI mirrors this: a .pdf argument routes directly to document OCR, while the document command handles explicit multi-source jobs. Output follows a versioned schemaVersion: 1 contract, which matters if you parse CLI output in a pipeline.

Recognition runs off the JavaScript main thread, and the README lists queues, cancellation and explicit cleanup as supported. The try/finally around engine.close() in the quick start is not decoration; the engine holds native resources.

Installing light-ocr and running a first recognition

Node.js 22 and 24 are supported. Install the package from npm:

bash
npm install @arcships/light-ocr

The README gives this TypeScript example. createEngine() resolves the platform build, recognizeEncoded takes the raw file bytes, and each line exposes text, confidence and box. The finally block closes the engine.

ts
import { createEngine } from "@arcships/light-ocr";
import { readFile } from "node:fs/promises";

const engine = await createEngine();

try {
  const result = await engine.recognizeEncoded(
    await readFile("image.jpg"),
  );

  for (const line of result.lines) {
    console.log(line.text, line.confidence, line.box);
  }
} finally {
  await engine.close();
}

CommonJS reaches the same exports through require("@arcships/light-ocr").

The CLI installs alongside the library, so no extra setup step is needed. This command prints text with coordinates as JSON:

bash
light-ocr image.png --format json

For a PDF, pass the file directly. The README notes the renderer is already included, and the default is 150 DPI:

bash
light-ocr report.pdf --pages 1-10 --format jsonl

Before trusting any of this on a new machine, run the diagnostics command. It reports hardware and providers, which is the fastest way to confirm whether you are on Core ML, WebGPU or CPU:

bash
light-ocr doctor --json

A region of interest can be cropped with --region 100,80,640,320, and detect returns boxes without recognition.

Where light-ocr stops being the right tool

The support matrix is the first constraint. Node.js 22 and 24 are the documented runtime versions, and there is no Python binding. A team with an existing Python OCR pipeline gets nothing here.

Hardware acceleration is uneven. Only macOS on Apple Silicon reaches Core ML, and only on macOS 15 or later. Linux arm64 and Windows arm64 fall back to CPU, as does macOS on Intel. If your deployment target is an ARM Linux container, you are running the CPU path regardless of what the acceleration table suggests at a glance.

Model tiers are split across packages. Small is the stable default at roughly 30 MB. Tiny and Medium are preview packages published under the next tag, so installing them means opting into a dist-tag rather than a stable release. Tiny is documented as covering 49 languages with no Japanese, which rules it out for Japanese text even though it is the smallest install. Medium is roughly 139 MB and the README excerpt cuts off before describing its coverage.

The README also does not document rollback for a failed upgrade, nor does it describe what happens when a platform build is missing for your architecture. Those are gaps to resolve by reading packages/light-ocr/README.md rather than assuming.

One more trade-off worth naming: the no-download design means the model ships inside the platform package. That is good for air-gapped installs and bad if you wanted to swap in your own model weights. The tier table shows each install contains only its selected model, so the choice is made at install time, not at runtime.

light-ocr compared with a hosted OCR API

The obvious alternative is a cloud OCR service. The difference is not accuracy claims, which this project does not make, but where the data goes and what the install looks like. A hosted API needs credentials, a network path, and a per-request cost model; light-ocr keeps images, PDFs and results on the machine and installs through npm.

That trade reverses under specific conditions. A hosted service can update its model without you redeploying, and it can serve languages and scripts that a 30 MB local model does not cover. light-ocr pins you to PP-OCRv6 in the tier you installed. If your document mix drifts toward scripts outside the model's coverage, the local package does not adapt on its own.

For C++ integrations the comparison is different again: there is a CMakeLists.txt, CMakePresets.json, include/ and src/ at the repository root, so the project is buildable natively rather than only through the Node binding. That is the path for embedding OCR in a C++ application without dragging in a JavaScript runtime, and it is the reason the project lists native C++ integrations as a target audience rather than treating Node.js as the only consumer.

Licence, maintenance and upgrade cost

The repository is licensed Apache-2.0, with a NOTICE file alongside the LICENSE. Apache-2.0 permits commercial use and modification and includes an explicit patent grant, but it also carries notice and attribution obligations: if you redistribute the package or a derivative, the licence and NOTICE terms apply. The bundled components are worth checking individually, since PDFium and the Noto Sans SC font are separate works with their own terms. That is a question for your legal team, not something this article can settle.

The repository is not archived, and the last push was on 2026-08-05. Recent releases follow a steady cadence: v0.5.5 shipped built-in PDF support on 2026-07-27, v0.5.6 restored Chinese PDF rendering on 2026-07-31, and v0.5.7 added support for re-signed macOS artifacts on 2026-08-05. The 0.5.x line is still moving, and the patch releases address rendering and signing problems rather than adding surface area.

Upgrade cost is mostly the platform package. Since the model, PDFium and font travel inside it, a version bump can change binary payload size and native behavior, not just JavaScript. The macOS re-signing rule is the sharpest edge: downstream packagers may re-sign the native binaries with their own Developer ID or ad-hoc identity, and on macOS the loader accepts a re-signed Mach-O when its signature verifies and its signing identity matches the host application, either the same TeamIdentifier or both ad-hoc. Other platforms and unsigned mutations keep the strict size plus SHA-256 gate. If you notarize a distributed app, that rule determines whether your build runs at all.

Editorial conclusion

Adopt light-ocr if you are building a Node.js CLI, desktop app or service that must read text from images or PDFs without sending files off the machine, and you are on Node.js 22 or 24. Do not adopt it if you need a Python binding, a hosted API, or Japanese support in the smallest model tier, since the Tiny tier is documented as covering 49 languages without Japanese. Before committing, verify that your target platform appears in the six-build acceleration table, run light-ocr doctor --json on the machine that will do the work, and check the macOS re-signing rule if you notarize your own app bundle.

Frequently asked questions

What is light-ocr?

It is a fast, offline OCR library for Node.js and C++, built on PP-OCRv6 with Core ML and WebGPU acceleration. It recognizes text in PDF, JPEG, PNG and raw image data and returns lines with confidence scores and quadrilateral coordinates.

How do I install and run light-ocr for the first time?

Install it with npm install @arcships/light-ocr on Node.js 22 or 24. The README example calls createEngine(), passes file bytes to recognizeEncoded, iterates result.lines for text, confidence and box, and closes the engine in a finally block.

Does light-ocr need to download a model at install time?

No. The README states the model, OCR runtime, PDF renderer and Chinese fallback font are included through the npm package's platform dependency, with no postinstall fetch, compiler, model download, PDF engine download or font download.

Which platforms does light-ocr accelerate with hardware?

Auto mode uses Core ML on macOS 15 or later with Apple Silicon, WebGPU through Vulkan on Linux x64 with glibc, and WebGPU through D3D12 on Windows x64. macOS on Intel, Linux arm64 and Windows arm64 fall back to CPU.

What is OCR?

OCR stands for optical character recognition: extracting text from images. light-ocr performs this locally, returning each recognized line with a confidence score and a quadrilateral box giving its position in the original image.

Official sources

  1. Official README
  2. Project repository
  3. Release notes
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/arcships-light-ocr.svg)](https://hysenlabs.com/projects/arcships-light-ocr)
Community notes

Community notes