# Tesseract.js: the CDN snippet in the README still installs v5

> Tesseract.js wraps a WebAssembly build of the Tesseract engine for the browser and Node.js, Apache-2.0 and installable from npm or a CDN. Its quick start, its worker lifecycle and its output defaults all carry sharp edges: the documented script tag pins v5, one worker must outlive a whole batch, and every output format except text is switched off unless you ask.

**naptha/tesseract.js** — GitHub describes it as Pure Javascript OCR for more than 100 Languages 📖🎉🖥. The repository metadata lists JavaScript as its primary language. The metadata lists the Apache-2.0 license. This article stays within the project description and details documented in the GitHub repository README.

- Repository: https://github.com/naptha/tesseract.js
- Website: http://tesseract.projectnaptha.com/
- Stars: 38,752 · Forks: 2,396
- Language: JavaScript
- License: Apache-2.0
- Published: 2026-08-13 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/naptha-tesseract-js

## One worker per batch, not one worker per image

The whole quick start is four calls inside an async function:

```javascript
import { createWorker } from 'tesseract.js';

(async () => {
  const worker = await createWorker('eng');
  const ret = await worker.recognize('https://tesseract.projectnaptha.com/img/eng_bw.png');
  console.log(ret.data.text);
  await worker.terminate();
})();
```

The language code is the first argument to createWorker, so 'eng' here is what selects the English model, and the recognizer returns a data object whose text property holds the transcription. What matters for a batch job is the sentence immediately after that snippet. Users are told to create a worker once, run worker.recognize for each image, and run worker.terminate once at the end, rather than running the snippet for every image. Each createWorker call stands up the WebAssembly engine again, so a loop that wraps the whole snippet in a forEach pays that setup on every iteration and leaves nothing running between calls.

If you are porting code written against an older release, the language argument is not where it used to be. The v5 notes record that createWorker arguments changed, that setting a non-default language and OEM now happens in createWorker with a call shaped like createWorker("chi_sim", 1), and that worker.initialize and worker.loadLanguage functions should be deleted from your code. Code from v3 and v4 that still calls those two functions is calling an API the current version removed.

## The documented CDN tag pins v5 while the package sits at 7.0.0

The installation section offers three routes: a script tag through a local copy or a CDN, webpack through npm in the browser, and npm or yarn on Node.js. The copy-and-paste browser route is this:

```html
<!-- v5 -->
<script src='https://cdn.jsdelivr.net/npm/tesseract.js@5/dist/tesseract.min.js'></script>
```

The version is in the path, and it is v5. The ESM alternative points at https://cdn.jsdelivr.net/npm/tesseract.js@5/dist/tesseract.esm.min.js, also v5. Copy either and you deploy a version two majors behind the current release, because package.json declares 7.0.0. After including the script the Tesseract variable is globally available and a worker is created with Tesseract.createWorker, which is the same lifecycle as the import form with a different entry point.

The Node side is where the version gap bites differently, since the runtime floor moved:

```shell
# For latest version
npm install tesseract.js
yarn add tesseract.js

# For old versions
npm install tesseract.js@3.0.3
yarn add tesseract.js@3.0.3
```

Tesseract.js v7 requires Node.js v16 or newer, where v6 required v14 or newer. The two instructions are the whole reason to read the version rather than assume it: a v5 script tag on a modern project, a v7 install on a v14 runtime, or a pinned 3.0.3 left in a manifest from years ago will each fail differently, and none of the failures will name the version skew.

## Since v6 every output format except text is off unless you ask

Version 6 is where the library got quieter by default. The notes credit it with fixing a memory leak in previous versions and reducing runtime and memory usage, then list breaking changes, and the first of them is that all output formats other than text are disabled by default. The result you get is a transcription and nothing else, even if your code was written to read bounding boxes or hocr markup.

Turning one back on is a third argument to recognize:

```javascript
worker.recognize(image, {}, { hocr: true })
```

The empty object is the options slot and the third is the output selection. If you upgrade from v5 and your pipeline read structured output, nothing throws. You get text, and the downstream code that expected word level geometry finds undefined. The v6 notes also mention minor changes to the structure of the blocks object, so even after re-enabling an output you may be reading a different shape than the one you tested against.

One inconsistency is worth flagging before you rely on any of this. The v4 notes record that the getPDF function was replaced by a pdf recognize option, while the current project scope section says the library does not support PDF files. The README does not settle which description applies to v7, so treat PDF as out of scope and check the API documentation before designing around that option.

## No PDF support, and the project names Scribe.js as the way out

The scope paragraph is unusually blunt about the boundary. Tesseract.js aims to bring the Tesseract engine, described as a separate project, to the browser and to Node.js, and it does that by wrapping a WebAssembly port of Tesseract. It does not modify core Tesseract features. Most notably it does not support PDF files, and it does not modify the Tesseract recognition model to improve accuracy.

That is a design decision with a named escape hatch. For feature requests outside that scope, the README points at Scribe.js, described as an alternative library created to accommodate common feature requests outside the scope of this repository. The difference in approach is specific: Scribe.js includes improvements to the Tesseract recognition model and supports extracting text from PDF documents, where Tesseract.js deliberately does neither. A comparison document is linked for the two of them.

Community work fills some of the gap in the middle. The community list includes an example for converting PDF to text, an example using the blocks output for word and symbol level data, plus Electron and TypeScript examples, and the README says official examples live in the examples directory, with browser and node subdirectories. The stated bar for the list is that examples be well documented enough for a new user to run and that projects be functional and actively maintained, which is the standard to hold the PDF example to before you build on it, since the project itself declines the job.

## package.json swaps the worker implementation with a browser field

The engine exists in two shapes, and which one you get is decided by a single mapping in the package manifest:

```json
"browser": {
  "./src/worker/node/index.js": "./src/worker/browser/index.js"
}
```

Any bundler that honours the browser field resolves the node worker path to the browser worker instead. That is why one import line works in both environments without a conditional, and it is also why the worker code is split into a node directory and a browser directory under src/worker. The entry points say the same quiet thing: main is src/index.js, type is commonjs, and types is src/index.d.ts, so the published package is source with a declaration file, not a compiled bundle. The dist/ directory exists for the CDN, and package.json points unpkg and jsdelivr at dist/tesseract.min.js.

There is no module field and no exports map in the manifest, which leaves ESM resolution to whatever your bundler does by default. The build produces two artifacts to serve that gap, and the build script names both:

```json
"build": "rimraf dist && webpack --config scripts/webpack.config.prod.js && rollup -c scripts/rollup.esm.mjs"
```

A webpack production bundle for the script tag, then a rollup ESM build. prepublishOnly runs that same build, so what is on the registry is whatever these two steps produced.

## The install hook asks for money and then ignores its own failure

There is one postinstall script in the manifest, and it is not part of OCR:

```json
"postinstall": "opencollective-postinstall || true"
```

It runs on every install, including transitive ones in a CI image build, and the trailing || true means the shell discards whatever happens. The script exists to print the project's funding message, which the README also links as a badge. Two things follow. The first is that your installs reach the network for a donation prompt even when the rest of your build is offline and pinned, which is the kind of thing that surprises people auditing a locked dependency tree. The second is that the escape hatch is deliberate: a failed funding prompt cannot break a consumer's install, at the cost of a failure that is never surfaced anywhere.

None of that is a defect so much as a choice, and it is a small one to reason about. It matters more because the rest of the package is otherwise conventional. Linting is eslint over src with a fix variant, the airbnb-base config is in the dev dependencies, and the repository ships .eslintrc alongside babel.config.json and karma.conf.js, so the tooling around the library is ordinary Node work.

## Running the tests means launching a real browser

The test script is two stages, and the second one is a browser suite:

```json
"test": "npm-run-all -p -r start test:all",
"test:all": "npm-run-all wait test:browser test:node:all",
"test:browser": "karma start karma.conf.js",
"test:node": "nyc mocha --exit --bail --require ./scripts/test-helper.mjs"
```

The shape is start a local server, wait for the built file to appear at http://localhost:3000/dist/tesseract.min.js, then run karma and mocha. Karma is configured through karma.conf.js at the root with karma-chrome-launcher in the dev dependencies and a firefox launcher beside it, so the browser half of the suite drives a real installed browser rather than a DOM shim. Contributors on a headless machine need one available before the suite is meaningful.

For users, the cost of that rigour shows up as the upgrade path instead. Three of the last four majors carry breaking changes: v4 made createWorker async and replaced getPDF with a pdf recognize option, v5 moved language and OEM into createWorker and removed worker.initialize and worker.loadLanguage, and v6 disabled non-text outputs by default. The v5 notes point at a separate issue for the full list, and there is a dedicated guide for moving from v2 to v5, so the migration work is documented but it is three migrations deep for anyone on an old release.

## Accuracy is bounded by the upstream model, and preprocessing is the only lever

The clearest statement of what this library will not do is the line about not modifying the Tesseract recognition model to improve accuracy. Transcripts come out as good as the engine and the trained data allow, and no setting in this wrapper pushes past that. The alternative named for accuracy work is Scribe.js, whose stated difference includes improvements to the recognition model.

Inside its own scope, Tesseract.js did add an accuracy feature, and it is a preprocessing one. The v4 notes record rotation preprocessing options, including auto-rotate, added for significantly better accuracy, and the same release made processed images retrievable, so the rotated, grayscale and binary versions the engine actually worked with can be fetched back. That combination is the practical lever available: let the library fix orientation before recognition, and inspect the result when a document comes back wrong, because a page photographed at an angle is a different problem from a page with bad print.

No accuracy figure appears anywhere in the documentation, and benchmarks/ exists in the repository without any results quoted, so a claim about how well it reads your documents is not something this project makes. Test it against your own pages, in the language you need, before you commit a pipeline to it. The language list itself is a link to a separate document rather than an inline table, and the project description claims more than 100 languages with English quality as the only one singled out.

## Conclusion

Adopt Tesseract.js when you need OCR in a browser bundle or a Node service, the input is an image rather than a PDF, and text output is enough. Do not adopt it expecting better accuracy than the upstream model, because the project states it does not modify the Tesseract recognition model, and do not adopt it for PDF extraction, which the project scope explicitly excludes and hands to Scribe.js. Before you ship, run the exact version you will deploy rather than the one the README's CDN tag pulls: that tag points at tesseract.js@5 while the current release is v7.0.0, and v6 is where every output format other than text was disabled by default.

## FAQ

### What is Tesseract.js used for?

It gets words out of images in almost any language, running in the browser or on Node.js. The library wraps a WebAssembly port of the Tesseract OCR engine, which is a separate project, and it does not modify core Tesseract features.

### Is Tesseract.js available on NPM?

Yes, as tesseract.js, installed with npm install tesseract.js or yarn add tesseract.js, and the current release is v7.0.0. There is also a CDN route through jsdelivr and unpkg, and the manifest points those fields at dist/tesseract.min.js.

### Is Tesseract OCR free?

The library is released under the Apache-2.0 licence, declared in package.json and shipped as LICENSE.md. It depends on Tesseract itself, which is a separate project, and on a WebAssembly port of that engine kept in its own repository.

### How accurate is Tesseract.js?

The project does not publish accuracy figures and states that it does not modify the Tesseract recognition model to improve accuracy. The accuracy feature inside its scope is preprocessing: v4 added rotation options including auto-rotate, and made the rotated, grayscale and binary processed images retrievable.

### Is there a better OCR than Tesseract?

The README makes no accuracy comparison and names one alternative: Scribe.js, which includes improvements to the Tesseract recognition model and supports extracting text from PDF documents, features the Tesseract.js project scope excludes. A comparison document between the two is linked from the README.

### How do I install tesseract js?

Three routes are given: a script tag via a local copy or a CDN, npm through webpack in the browser, and npm or yarn on Node.js. Tesseract.js v7 requires Node.js v16 or newer, where v6 required v14 or newer, and the documented CDN script tag points at tesseract.js@5 rather than the current release.

## Sources

- [Official documentation](http://tesseract.projectnaptha.com/)
- [Official README](https://github.com/naptha/tesseract.js#readme)
- [Project repository](https://github.com/naptha/tesseract.js)
- [Release notes](https://github.com/naptha/tesseract.js/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/naptha-tesseract-js
