# pdf2json: pulling text and AcroForm fields out of PDFs with a vendored pdf.js

> pdf2json is a Node.js parser that turns PDF binaries into JSON, plain text and form field data, built on a fork of pdf.js compiled into the package itself. It is event driven, has no npm dependencies, and ships a command line tool as well as a library.

**modesty/pdf2json** — converts binary PDF to JSON and text, for server-side PDF processing and command-line use. Zero dependency.

- Repository: https://github.com/modesty/pdf2json
- Website: https://github.com/modesty/pdf2json
- Stars: 2,213 · Forks: 394
- Language: Java
- License: NOASSERTION
- Published: 2026-10-07 · Updated: 2026-10-07 · Language: en
- Canonical page: https://hysenlabs.com/projects/modesty-pdf2json

## What the library actually parses out of a PDF

Three outputs, and the third is the one that matters. The feature list on the README names PDF text extraction into structured JSON, interactive form element handling for flexible data capture, and use either as a web service or as a standalone command line tool.

The form support is the part that distinguishes this from a generic text extractor. The examples show a call to `pdfParser.getAllFieldsTypes()` written straight out to a `F1040EZ.fields.json` file, and the topics on the repository include `pdf-form` and `pdf2form`. That is a different question from "what does this document say": it is "what values are in the boxes of this form, keyed by field name". For an intake pipeline, a claims form or a tax document, that is the output you actually want.

Text output is available separately through `getRawTextContent()`, which the README shows being written to a `.content.txt` file containing only the textual content of the PDF. So you can produce both from one parse rather than paying for two passes.

The package manifest describes the mechanism in a line: a PDF file parser that converts PDF binaries to JSON and text, powered by porting a fork of PDF.JS to Node.js. That single phrase explains most of the design decisions below.

## Installing it, and the claim of zero dependencies

The install is the standard npm line:

```bash
npm i pdf2json
```

Or globally for the command line utility:

```bash
npm i pdf2json -g
```

And to update an existing global install:

```bash
npm update pdf2json -g
```

The feature list claims zero dependencies, dependency-free since v3.1.6, and only pure JavaScript code. The repository tree supports the second half of that claim and complicates the first: there is a `lib/` directory and the coverage script explicitly excludes `lib/pdfjs-code.js`, which is a compiled pdf.js sitting inside the repository rather than an npm dependency. There is also a `package-lock.json` and no runtime `dependencies` block in the manifest.

So the accurate reading is that pdf2json has no third-party npm packages at runtime because the third-party code is vendored in and built into `dist`. That is a real advantage for deployment, because there is no transitive dependency tree to audit or to break on a major release. It is also a maintenance trade: the fork of pdf.js has to be updated by hand, and you are trusting a single maintainer's port rather than a stream of upstream fixes. The `base/` and `lib/` directories plus the rollup configuration are where that port lives.

## An event driven API, not a promise

The parsing interface is built on Node events, and the basic shape is short. You construct a parser, subscribe to a data event, and tell it to load a file:

```javascript
import fs from "fs";
import PDFParser from "pdf2json";

const pdfParser = new PDFParser();

pdfParser.on("pdfParser_dataError", (errData) =>
 console.error(errData.parserError)
);
pdfParser.on("pdfParser_dataReady", (pdfData) => {
 fs.writeFile(
  "./pdf2json/test/F1040EZ.json",
  JSON.stringify(pdfData),
  (data) => console.log(data)
 );
});

pdfParser.loadPDF("./pdf2json/test/pdf/fd/form/F1040EZ.pdf");
```

Two events, one for failure and one for success, both mandatory. There is no promise to await and no thrown error on a malformed file, which is the first thing to know before you write production code on top of it: a parse that fails delivers `pdfParser_dataError` with a `parserError` property, and if you do not subscribe you find out nothing.

For finer control there are page-level events, available since v2.0.0:

```javascript
pdfParser.on("readable", (meta) => console.log("PDF Metadata", meta));
pdfParser.on("data", (page) =>
 console.log(page ? "One page paged" : "All pages parsed", page)
);
pdfParser.on("error", (err) => console.error("Parser Error", err));
```

A buffer works too, through `parseBuffer` after reading the file, which is the path to use when the PDF arrives over HTTP rather than from disk. And the parser is a stream, so the README shows input and output piped through it, which means a large PDF does not have to be fully materialised before processing starts.

## The test corpus is the project's real specification

The README documents the testing setup in unusual detail, and reading it tells you more about what the library handles than any feature list could.

`npm test` runs the unit suites, which it describes as 7 suites and more than 74 tests built with the Node.js built-in `node:test` runner. The suites live in `test/_test_*.cjs` in CommonJS, and the run also covers `parse-r` and `parse-fd` with ES modules via the command line. Because the pretest step builds bundles and source maps for both ES Module and CommonJS into `./dist`, the tests exercise the shipped artefacts rather than the source tree.

The broader suites are more revealing. `npm run test:forms` scans and parses 260 PDF AcroForm files under `./test/pdf`, running with the `-s -t -c -m` command line options, and generates a primary output JSON, additional text content JSON, form fields JSON and a merged text file for each one. The README reports it usually takes about 20 seconds on a MacBook Pro, and that an update from 2024-04-27 measured 7 to 8 seconds on an M2 Mac. `npm run test:misc` covers 15 PDF files where three are expected to throw, and the README names the exact failures: a bad XRef entry, an unsupported encryption algorithm and an invalid XRef stream header.

`npm run parse-r` then scans 165 PDFs under `./test/pdf/fd/form/` through the Stream API. A library that ships several hundred real-world PDFs as fixtures and asserts which ones fail is doing something quite different from a project with a mock-file unit test, and it is the closest thing here to a compatibility claim you can actually check.

## Command line flags and the logging you may want to silence

The command line tool lives at `bin/pdf2json.js`, invoked as `node ./bin/pdf2json.js` with an input flag and an output flag, and the npm scripts show the shape of it:

```bash
node ./bin/pdf2json.js -f ./test/pdf/fd/form/F1040.pdf -o ./test/target/fd/form
```

The four-letter shorthand `-s -t -c -m` recurs across the test scripts and is the flag vocabulary to learn. The README explains the logging behaviour in a section aimed squarely at CI, which is worth reading before you wire this into a pipeline.

There are two logging systems in the code. One consumes the standard `console.log` and `console.warn` APIs. The other consumes the project's own shared log function in `base/shared/util.js`. Mocking the console calls handles the first. For the second you have three documented options: set the environment variable `PDF2JSON_DISABLE_LOGS` to `1`, pass the `-s` silent flag on the command line, or pass `VERBOSITY_LEVEL` as 0 when invoking `PDFParser.loadPDF`.

That is a small thing, and it is the kind of small thing that costs an afternoon if you do not find it. Note also that the command line tool and the library expose the same underlying parser, so a flag that silences the CLI does not silence an in-process parse.

## Licensing, versioning and the honest limits

Two facts sit oddly together in this repository and both are worth checking before you adopt it.

The first is the licence. GitHub reports no recognised licence type for the repository, and the default branch carries a `license.txt` file whose terms the metadata does not classify. A `LICENSE` file whose contents the platform cannot identify is a signal to read it rather than assume. For a package that inlines a fork of Mozilla's pdf.js, whose upstream licence is Apache 2.0, the terms of the vendored code matter as much as the terms of the wrapper, and the file is the place where that should be settled.

The second is versioning. The manifest reads 4.1.1 while the most recent tagged release is v4.1.0, published on 2026-09-11, with v4.0.3 on 2026-04-16 and v4.0.2 on 2026-01-17. The repository is not archived and the last push was on 2026-09-11, so the default branch is one patch ahead of the last tag.

The real functional limits are clearer than either. This is a structured extractor, not a layout renderer: it will give you text content and form fields, not positioned glyphs with coordinates and font information, so reconstructing a visual page is out of scope. It is also a Node library, so the Python question has a short answer: there is no Python package here, and a Python service would shell out to the CLI or call a Node process. And it is not an OCR engine, so a scanned PDF with no text layer yields nothing to extract.

## Conclusion

pdf2json earns its place on server-side jobs where AcroForm field extraction matters as much as text: it pulls `getAllFieldsTypes()` out of a form and gives you structured data rather than a rendering, it runs in a plain Node process with nothing installed alongside it, and it has been accumulating a decade of PDFs as test fixtures. It is the wrong choice for reading a PDF layout faithfully, because it produces structured text and field data rather than positioned text with font metrics, and it is the wrong choice for a browser application, because pdf.js itself is the tool designed to run client side. Decide on one point before installing: if your documents are ordinary prose, use pdf.js directly or a text extraction service, and if your documents are forms, this package has the field-level access the others leave you assembling. v4.1.1 is current in the manifest and v4.1.0 shipped on 2026-09-11, so check which one your lockfile resolves before reading older examples against it.

## FAQ

### What is a PDF JSON file?

It is the structured output this library produces from a PDF: the page contents and text extracted into a JSON document rather than rendered. In pdf2json the parsed result arrives in the `pdfParser_dataReady` event as a JavaScript object, and the examples write it out with `JSON.stringify(pdfData)`.

### Is it possible to convert PDF to JSON?

Yes, that is the package's main job. You construct a `PDFParser`, subscribe to `pdfParser_dataReady` and to `pdfParser_dataError`, then call `loadPDF(path)` or `parseBuffer(buffer)`. The result is a JavaScript object you can write to a file, or you can pipe the parser into an output stream.

### Can I convert a PDF to JSON for use in an e-invoice?

Form-heavy documents are where this library has something the general extractors do not. Calling `pdfParser.getAllFieldsTypes()` returns the interactive form field data as a structure you can serialise, which is how the field values of an AcroForm become JSON. There is no e-invoice-specific output format in the library, so mapping those fields onto an invoice schema is your work.

## Sources

- [Issues](https://github.com/modesty/pdf2json/issues)
- [modesty/pdf2json on GitHub](https://github.com/modesty/pdf2json)
- [Project website](https://github.com/modesty/pdf2json)
- [README](https://github.com/modesty/pdf2json/blob/master/README.md)
- [Releases](https://github.com/modesty/pdf2json/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/modesty-pdf2json
