# invoice2data: turning PDF invoices into CSV with templates you have to write

> A Python library and command line tool that extracts issuer, amount, date and invoice number from PDF documents by matching regular expressions against text pulled out by a pluggable backend, and a 1.0 release that moved a lot of the internals around.

**invoice-x/invoice2data** — Extract structured data from PDF invoices

- Repository: https://github.com/invoice-x/invoice2data
- Website: https://invoice2data.readthedocs.io/
- Stars: 2,221 · Forks: 557
- Language: Python
- License: MIT
- Published: 2026-10-06 · Updated: 2026-10-06 · Language: en
- Canonical page: https://hysenlabs.com/projects/invoice-x-invoice2data

## Three steps, and none of them involve understanding an invoice

The README describes the pipeline in three numbered steps, and the order matters when you are deciding whether this tool fits your documents.

First, text is extracted from PDF files through a pluggable, cascading backend. The default is `pdfium`, which the README notes requires no system dependencies, and the alternatives include `pdftotext`, `text`, `pdfminer` and `pdfplumber`, plus OCR paths through `tesseract`, `ocrmypdf`, `docTR`, `paddleocr` and Google Cloud Vision. Cascading is the important word: the backend is a preference order rather than a single hard dependency.

Second, regular expressions are searched for in the result using a YAML or JSON based template system, with an optional AI fallback documented separately. Third, the results are saved as CSV, JSON or XML, or the PDF files themselves are renamed to match their content.

So the tool does not parse an invoice's semantics. It finds the string that looks like an invoice number because you told it the pattern that matches an invoice number. That is the whole bargain, and it is a good one when the documents are text based and the layouts are stable.

## Installing needs nothing but pip

The quickstart is three lines, and the first one is the whole install:

```bash
pip install invoice2data
invoice2data invoice.pdf                          # extract -> CSV
invoice2data --output-format json invoice.pdf     # or JSON / XML
```

The README states directly that no system libraries are required by default, because the `pdfium` backend bundles its own engine. That is a change in kind from the older poppler-dependent approach, where installing the package was only half the task.

Optional backends and extras are covered in the installation guide in the documentation, and the release notes show those extras being loosened deliberately, with a change to relax hard `==` pins in optional extras rather than leaving users pinned to a stale version.

Used as a library, the entry point is a single function:

```python
from invoice2data import extract_data

result = extract_data("invoice.pdf")
```

The project metadata in `pyproject.toml` declares Python 3.10 and above as the floor, lists classifiers for 3.10 through 3.13, names Manuel Riel as the author, and records MIT licensing. It also declares a development status of Production/Stable, which is a claim the 1.0 tag is presumably meant to support.

## Templates are the actual product

The template system is where the effort goes, and the README lists what the template language lets you express. You can precisely match content in PDF files, use plugins to match line items and tables, define static fields that are identical on every invoice, define custom fields needed by your organisation or process, attach multiple regular expressions to a single field when the layout or the wording changes, and define the currency.

Line items get a dedicated plugin, developed by Holger Brunn, called `lines`. That is the part most invoice projects do badly and the part this one has chosen to support explicitly, and the documentation has a separate page for recommended fields describing the canonical output schema.

Two features added later show the template language widening. A `replace:` key on each field handles per-field data cleanup, so you can normalise a value at extraction time rather than downstream. A per-template `required_fields:` override lets an individual template opt out of the invoice-shaped defaults, which matters for the non-invoice documents the project still wants to support.

The output shape is a Python dictionary of records, with keys like issuer, amount, date, invoice_number, currency, desc and template_name, where the last one tells you which template matched. Seeing the template name in the output is a useful diagnostic, because when a field comes out wrong you can immediately tell whether the wrong template matched or the right template's regex is stale.

## Generating a first template without writing one by hand

Writing regex for an unknown layout is the tedious part of this workflow, and the README says so indirectly by shipping a flag that does it for you. Running `invoice2data --new-template SAMPLE` guesses parameters for a new invoice format, and the README emphasises that this is deterministic and needs no AI. Adding `--ai` makes it ML-assisted, and a separate runtime flag, `--ai-fallback`, lets the model step in when a template fails to match.

Table parsing is opt-in rather than automatic. An advanced table plugin uses Camelot, installed as an extra with `pip install invoice2data[camelot]`, and the documentation also describes an Excalibur to template converter. The choice to keep tables behind a flag is reasonable, since table extraction is where a clean text layer usually does not exist.

For containers, the `Dockerfile` in the repository root is where the full dependency set appears, starting from a slim Python image and installing tesseract, poppler utilities, ImageMagick and Ghostscript before the package itself:

```bash
RUN pip install --no-cache-dir -U pip invoice2data ocrmypdf

RUN groupadd -r invoice2data && useradd -m -r -g invoice2data invoice2data
```

The image also drops to an unprivileged user and sets `invoice2data` as its entrypoint, which is the shape you would expect from a maintained image.

## Version 1.0 was a cleanup, not a rewrite

Version 1.0.0 landed on 2026-06-23 and its release notes read mostly like dependency bumps and a few CI improvements, which suggests the version bump reflected accumulated internal work rather than a change in interface. Version 1.0.1 followed on 2026-08-01 with a more telling list: splitting the library API out of `__main__` and replacing user-facing asserts, preserving extracted currency by switching a clobbering assignment to a setdefault, skipping empty template files instead of crashing, raising a `TemplateSyntaxError` for missing area-block keys, and emitting a single CSV header row across all invoices.

Every one of those is a bug report from a real pipeline. The currency one matters most for accounting use, since a template-defined currency being overwritten by the parser is exactly the kind of error that produces a plausible looking but wrong spreadsheet. The single CSV header fix matters too, because a header repeated per invoice breaks naive concatenation.

The same release added a `windows_strict` gate and a template-rot schema validator, support for Windows with a Windows-specific prefix command for ImageMagick, and a bump of torch from 2.10.0 to 2.13.0 in response to a Dependabot alert. Version 0.5.0, published 2026-05-23, was a single change, modernising type hints to PEP 585 and 604 syntax. The last push to the repository is 2026-09-28.

## Optional compilation, and what is still deferred

One packaging detail is worth knowing if you care about speed. A `setup.py` exists alongside `pyproject.toml` for a single purpose: optionally compiling hot-path modules into C extensions with mypyc. It only activates when the `INVOICE2DATA_COMPILE_MYPYC` environment variable is set to 1, which the docstring says the cibuildwheel release job does. Otherwise a pure Python package is built, so the default build from source distribution stays pure Python.

The list of compiled modules is selective and the reasoning is documented in comments. `extract/utils.py`, `extract/_regex.py`, the regex and lines parsers and the tables plugin are compiled, while `__main__.py` is excluded because Click sets attributes on the decorated function object, and `invoice_template.py` is excluded because its OrderedDict subclass and dynamic attributes do not compile cleanly.

The README's open task list is short and specific. Canonical schemas for non-invoice document types such as waybills, delivery notes and purchase orders are still deferred, even though the per-template override makes them workable in practice. A visual template builder is being pursued through the OCA/edi community project, an `account_invoice_import_invoice2data_db_templates` module currently in review, which would put a click-to-suggest interface on top of the disk templates. Tighter identifier validation through python-stdnum is deferred pending alignment with the stdnum version in Odoo core.

## Conclusion

invoice2data's design assumption is worth stating plainly before anything else: the tool does not read invoices, it matches regular expressions against extracted text using a template you supply. That is why it is fast, deterministic and easy to reason about, and why a new supplier's layout costs you a template rather than a fine tuning session. The 1.0 line makes the packaging considerably more comfortable, with Python 3.10 through 3.13 declared, a declared stable development status, and optional mypyc compilation for the hot path modules. The honest limit is coverage of document types: waybills, delivery notes and purchase orders are listed as still deferred, with only a per-template override available today. Start with the `--new-template` flag on one sample PDF to get a draft template, correct it by hand, then reach for `--ai-fallback` only on the layouts where regex has genuinely stopped generalising.

## FAQ

### How do I extract structured data from an invoice PDF?

With invoice2data you extract text from the PDF using a pluggable backend, then match it against a template of regular expressions that you supply, then write the result out as CSV, JSON or XML. The default backend is pdfium, which needs no system libraries, and a first template can be generated from a sample with the --new-template flag rather than written from scratch.

### Is invoice2data a Python library or a command line tool?

Both. It installs with pip install invoice2data and runs as invoice2data invoice.pdf from a shell, and it is importable, where the whole library entry point is extract_data from the invoice2data package. It requires Python 3.10 or newer and the project metadata lists support through 3.13.

### What is a template in invoice2data and what can it define?

A template is a YAML or JSON file describing how to match content in a PDF. It can define static fields repeated on every invoice, custom fields for your own process, several alternative regular expressions per field when wording changes, the currency, and plugins for line items and tables. The name of the template that matched is included in the output record, which makes diagnosing a bad extraction straightforward.

### Can invoice2data read scanned or image-only invoices?

Yes, through OCR backends rather than the default path. The README lists tesseract, ocrmypdf, docTR, PaddleOCR and Google Cloud Vision as OCR options, and the Dockerfile installs tesseract, poppler utilities, ImageMagick and Ghostscript to support them. Text-based PDFs should still go through pdfium first, since OCR is slower and less reliable on clean text layers.

### What is the newest version of invoice2data?

Version 1.0.1, published on 2026-08-01, following 1.0.0 on 2026-06-23 and 0.5.0 on 2026-05-23. The 1.0.1 changes include preserving extracted currency instead of overwriting it, skipping empty template files rather than crashing, and writing a single CSV header row across all invoices.

### Does invoice2data work on documents that are not invoices?

Partially. A per-template required_fields override lets a template opt out of the invoice-shaped defaults today, so non-invoice documents are workable, but the README lists canonical schemas for waybills, delivery notes and purchase orders as still deferred. Line item and table handling is provided by plugins rather than a separate schema.

## Sources

- [invoice-x/invoice2data on GitHub](https://github.com/invoice-x/invoice2data)
- [License: MIT](https://github.com/invoice-x/invoice2data/blob/master/LICENSE)
- [Project website](https://invoice2data.readthedocs.io/)
- [README](https://github.com/invoice-x/invoice2data/blob/master/README.md)
- [Releases](https://github.com/invoice-x/invoice2data/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/invoice-x-invoice2data
