# Camelot's neural parser is documented as one that never invents a value

> An MIT-licensed table extraction library with five parsers, a bundled PDF backend and no system dependencies, whose optional neural backend supplies only table structure while cell text always comes from the document's own text layer. Its distribution is named differently from its repository, the optional extras carry version floors split by Python version, and the section arguing why to use it ends mid-sentence.

**camelot-dev/camelot** — A Python library to extract tabular data from PDFs

- Repository: https://github.com/camelot-dev/camelot
- Website: https://camelot-py.readthedocs.io
- Stars: 3,839 · Forks: 548
- Language: Python
- License: MIT
- Published: 2026-09-23 · Updated: 2026-09-23 · Language: en
- Canonical page: https://hysenlabs.com/projects/camelot-dev-camelot

## The model supplies structure, never text

Two paragraphs are worth reading together, and together they are the strongest claim the readme makes. One explains that for scanned or image-only files you install the neural backend with recognition and read with the model flavour: the model reads the structure from the page image and recognition supplies the text. The other then states the guarantee plainly: the model only supplies the table structure, while cell text comes from the document's own text layer, or from recognition for scans, so it never invents or alters a value. For financial and regulatory extraction that is the difference between a tool you can audit and one you cannot, and it is achieved by an architectural choice rather than by a warning label: the model is never given the job of reading values at all. A benchmark on a public table-recognition dataset is cited for the structure recovery, which is the half the model is responsible for.

## The parser choice is a table, and it names the benchmark

Five parsers are offered and the readme gives a decision table rather than a feature list. Ruled tables go to the lattice parser by default, because detecting a grid from ruled lines is deterministic, with an option that unions the document's own vector lines with image-detected ones so faintly ruled tables are still found. Borderless tables go to two text-alignment heuristics that need nothing extra installed. The best-quality route for borderless tables is the neural one, installed as an extra, described as heavier and opt-in. Mixed or unknown documents get an automatic choice made per page. And the quality claim comes with a dataset and a metric: on a public financial table dataset it roughly doubles a structure score against the heuristics. That is a specific, checkable claim, and the repository does contain a benchmark directory.

## The bundle ships its own PDF backend, so there are no system packages

The last feature bullet says the default backend is bundled and there are no system dependencies. That is the single most consequential engineering decision in this library, because table extraction normally means a PDF engine installed on the host, and a library that needs one is a library that cannot go into a container without a base image chosen for it. The dependency list shows what the bundle is: a Python binding to a PDF rendering engine, plus an image library, a headless computer-vision build, a spreadsheet writer, an array library, a dataframe library and a table formatter. Reading a document and rendering a page are then both internal, so the same code path works on a laptop, in continuous integration and in a notebook.

## Every table carries its own quality report

Extraction that cannot be checked is extraction you cannot use, and this library takes that seriously. Each table has a parsing report with an accuracy figure, a whitespace figure, its order and its page number, and the sample output in the readme shows an accuracy just under ninety-nine and a whitespace score in double digits. There is also a confidence score per table and a filter on the returned collection so you can drop the noise. Two derived features fall out of the same idea. Continuations across page boundaries can be stitched into one table with a single call, which is what a real document does to you, and the export layer offers six formats including a spreadsheet, a database table and markup rather than only comma-separated values.

## Two extras carry version floors split by Python version

The optional dependency groups are where the maintenance history shows. The plotting extra declares two different minimum versions of the same library, one for Python below a given version and one for the version and above, which is how you support a library that dropped support for an older Python. The neural extra pulls a deep learning framework and a transformers library, and its inline comment explains the constraint: the fifth major of that library validates configurations strictly and rejects the checkpoint this project uses, so the dependency is handled in the file rather than with an upper bound. There is also a comment explaining that the neural backend is imported lazily, so the core install and the other four parsers never load a deep learning framework at all. That is the right way to ship a heavy optional path, and it is documented in the file where a maintainer will find it.

## The package name, the import name and the repository name are three things

The repository is called one thing, the distribution published to the package index is called another with a suffix, and the module you import is the first of them. So the install command, the import statement and the repository address differ, and a user searching for the repository name will find the wrong thing. The metadata makes it explicit with two URL fields, one pointing at the repository under the other name and one at the documentation, plus a separate changelog field that points at the releases rather than at a file in the tree, even though a changelog file is committed. It is a small piece of friction in an otherwise carefully built package, and it is the kind of thing that produces a wrong-installation question at least once a year.

## A security policy for the dependency scanner and a policy file nobody reads

The repository root carries a file whose name is a configuration for an automated vulnerability scanner's policy, which is not a common sight and tells you the project has thought about what its automated dependency checks are allowed to fail. It also carries a benchmark directory, a test session runner file, a lockfile, a requirements file and a project manifest all at once, which is four ways to describe dependencies for a library that publishes one. The three example notebooks are committed, including a step-by-step walkthrough of one specific parser and a comparison of all of them, and two of the three are also linked as hosted notebooks. So the documentation a user is sent to is versioned in the repository, which is the right call for a library whose output people check.

## The argument for using it stops mid-word

The section that answers why you would choose this library over anything else begins with a configurability bullet and stops partway through the first sentence, with the verb in the wrong form. Everything before it is precise: five parsers and how to choose between them, the guarantee about invented values, the bundled backend, the metrics, the exports, the page stitching, the flexible input including raw bytes and any binary stream. Everything after it, including the licence statement and the acknowledgements, is absent from the visible document. So the last thing a reader meets in this project is a broken sentence, after a hundred lines of unusually substantive documentation. It is the cheapest thing on the list to fix and the only one that costs a reader their confidence in the rest. The whole read function in the readme's own example is three lines:

```python
>>> import camelot
>>> tables = camelot.read_pdf('foo.pdf')
>>> tables
```

## Conclusion

Camelot is worth using if your PDFs contain tables and you want them as data frames rather than screenshots, because it ships five parsers with a documented choice between them, reports a quality figure per table so you can filter bad extractions, and stitches tables that continue across pages. Two things to know. The neural backend is opt-in and pulls a deep learning framework, and the readme is explicit about what it is for: it supplies table structure only, while the cell text always comes from the document's own text layer or from optical recognition, so it cannot invent or alter a value. That guarantee is the reason to prefer it over a purely model-based extractor on financial data. And built-in parsers need a text-based PDF, so a scan needs the extra install before anything else will work.

## FAQ

### how to install camelot in python

Install the published distribution with the package manager. The default backend is bundled and there are no system dependencies, so there is nothing else to install for the built-in parsers. Two optional extras exist: one for plotting, and one for the neural backend with recognition, which is what you need for scanned or image-only files.

### how to use camelot python

Import the module and call the read function with a path, a URL, raw bytes or any binary file-like object. The result is a collection you can filter, index, export to six formats, or turn into a dataframe, and each table carries a parsing report with accuracy and whitespace figures.

### Which Camelot parser should I use for borderless tables?

Two text-alignment heuristics are fast and need nothing extra installed. For the best structure quality, install the neural extra and use that flavour, which on a public financial table dataset roughly doubles a structure score against the heuristics and requires a deep learning framework.

### Can Camelot read a scanned PDF?

Not with the built-in parsers, which need a text-based file; the readme borrows another tool's test for that, which is whether you can click and drag to select text in a viewer. For scans, install the extra that combines the neural backend with recognition and read with the model flavour.

### Does the Camelot neural backend invent cell values?

The readme states that it does not. The model supplies only the table structure, while cell text is taken from the document's own text layer, or from optical recognition for scans, so the model is never given the job of reading values.

## Sources

- [camelot-dev/camelot on GitHub](https://github.com/camelot-dev/camelot)
- [License: MIT](https://github.com/camelot-dev/camelot/blob/master/LICENSE)
- [Project website](https://camelot-py.readthedocs.io)
- [README](https://github.com/camelot-dev/camelot/blob/master/README.md)
- [Releases](https://github.com/camelot-dev/camelot/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/camelot-dev-camelot
