# ExtractThinker's documented install stops at a release from June 2025, and the OCR and PDF paths need packages no list declares

> ExtractThinker wraps a document loader, an LLM client and a Pydantic contract into a single extract call. The version on PyPI is sixteen months behind the branch, four dependency files describe three different dependency sets, and two of the paths the readme tells you to use are not in the declared dependencies.

**enoch3712/ExtractThinker** — ExtractThinker is a Document Intelligence library for LLMs, offering ORM-style interaction for flexible and powerful document workflows.

- Repository: https://github.com/enoch3712/ExtractThinker
- Website: https://enoch3712.github.io/ExtractThinker
- Stars: 1,600 · Forks: 154
- Language: Python
- License: Apache-2.0
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/enoch3712-extractthinker

## The documented install stops at a release from June 2025

The install line in the readme is one command, and it resolves to a release that predates most of the repository:

```bash
pip install extract-thinker
```

The tags tell the story. v0.1.12 was published 2025-04-16, v0.1.13 on 2025-04-24 and v0.1.14 on 2025-06-09, and the version field in pyproject.toml still reads 0.1.14. The default branch, meanwhile, was pushed on 2026-09-16, so roughly fifteen months of commits sit above the newest tag without a release above it.

The readme is explicit about the gap rather than hiding it. A section headed 2026 additions on main names page retrieval with SQLite, parallel field extraction, configurable model routing, local entity masking, page events, PyMuPDF, Camelot, Tabula and Adobe loaders, and an MCP service with Docker Compose, then states that these changes are not yet a new PyPI release and that installing from a checkout is the way to reach them. The three commands it gives for that are git clone, cd ExtractThinker and pip install -e .

So there are two different products behind one project name, and the wheel and the branch are not interchangeable. A reader who follows the quickstart and a reader who follows the 2026 section end up with different libraries under the same import name.

## Four dependency files, and the container image reads two of them

The repository root holds requirements.txt, requirements-server.txt, requirements-test.txt and requirements-docs.txt next to a pyproject.toml that declares its own set through Poetry. They do not agree.

requirements.txt names pydantic, instructor, litellm, pytesseract, Pillow, python-docx, xlrd, pypdfium2 and pytesseract again, with pytesseract written twice in the same list and no version bound on any entry. Compared with the Poetry dependencies it drops python-dotenv, cachetools, pyyaml, tiktoken, python-magic, playwright and libmagic, and adds pytesseract, python-docx and xlrd, which the Poetry block does not mention at all.

Which file matters depends on how you install. The Dockerfile copies pyproject.toml, README.md and requirements-server.txt, then runs pip install --no-cache-dir . together with that server requirements file, so the image is built from the Poetry metadata plus one of the loose files, and requirements.txt takes no part in it. A developer following the readme installs only the package, which means the Poetry block alone. Whoever wrote requirements.txt was describing a different environment again, and nothing in the build points at it.

The Poetry block itself is more disciplined. It sets python to >=3.9,<3.14, gives lower bounds on pydantic at 2.11.5, litellm at 1.71.1, instructor at 1.8.3 and pypdfium2 in the 4.30 to 6 range, and caps pillow below 12.0.

## Python 3.13 is in range while tiktoken steps out of it

The readme states that the core supports Python 3.9 to 3.13 and that the optional MCP service requires 3.10 or newer. The Poetry metadata agrees on the ceiling, with python set to >=3.9,<3.14.

One dependency does not. tiktoken carries a marker of python = ">=3.9,<3.13", which means the resolver drops it entirely on 3.13 rather than failing loudly. So the version the readme calls supported loses a declared dependency without any note in the prose, and nothing in the file says what tiktoken was used for, so a reader cannot tell whether that is a missing feature or a missing import. The dev group makes the same split in a second place, pinning numpy to ^1.26.4 below 3.13 and to >=2.1,<3 from 3.13 upward.

The container sidesteps the question by pinning an interpreter. The Dockerfile starts from python:3.12-slim, which sits inside every range the project claims and inside the range where tiktoken still installs. Anyone running 3.13 natively is the only person exposed to the marker, and the readme does not mention the difference.

The same pattern appears in the MCP requirement. A 3.9 install is documented as valid for the core and invalid for the service, so the floor for the service is a separate number from the floor for the library, and neither the compose file nor the Dockerfile records which one it needs.

## OCR is a documented requirement that no declared list installs

Scanned documents are named as a case the library does not handle on its own: the readme says scanned documents need OCR or a vision-capable model. No OCR package appears among the Poetry dependencies. pytesseract appears in requirements.txt, twice, and nowhere else.

That leaves the common case of a scanned invoice or receipt without a path from the documented install. A wheel install brings pydantic, litellm, pillow, pypdfium2, instructor, python-dotenv, cachetools, pyyaml, tiktoken, python-magic, playwright and libmagic, and a user with a pile of scanned pages has to find a Tesseract binding and a working tesseract binary themselves. The readme does not name a package or give a command for it, and the quickstart is pointed at instead.

The model route is the documented alternative, and it has its own shape. EXTRACT_THINKER_MODEL selects a model available to the operator and that provider's API key is configured separately, with the readme noting that extraction uses the configured provider and may incur charges. A local option exists through Ollama, which the workflow table links as its own page. So the choice for a scanned document is an external OCR stack or a bill, and the library contributes neither.

One related detail from the same dependency block: playwright is a hard requirement at >=1.52.0, so a text extraction library installs a browser automation package by default, and no browser download step is written down anywhere in the readme.

## Six ways to read a PDF, and the readme points at the undeclared one

PDF handling is where the dependency lists and the documentation part company most visibly. The Poetry block declares pypdfium2 in the range >=4.30.2,<6. The readme tells you to install pypdf and use DocumentLoaderPyPdf. The 2026 section on main adds PyMuPDF, Camelot, Tabula and Adobe loaders.

So the PDF path the readme documents is not the PDF path the package installs, and the declared one is the path the readme does not mention. Both pypdf and pypdfium2 end up needed or unused depending on which you follow, and the four loaders on main are available only from a checkout.

MIME detection is the other loader-level dependency, and it is the one place the project reaches outside Python. python-magic is declared at >=0.4.27 alongside an entry called libmagic pinned to the wildcard version, which is the system library rather than the Python binding. The readme spells out the consequence: system MIME detection needs libmagic installed through brew on macOS or apt-get on Debian and Ubuntu, and the Dockerfile does exactly that with libmagic1 before switching to a non-root user.

Together that is six PDF routes plus a system-level prerequisite, with the documented install covering one of the six and requiring a manual package for the MIME layer.

## The MCP container is read-only, loopback-bound and starts with no key

The optional service ships as a container, and its configuration is the most defensive part of the repository. The Dockerfile sets EXTRACT_THINKER_DOCUMENT_ROOT to /data and EXTRACT_THINKER_HOST to 0.0.0.0, creates a user with uid 10001, chowns /data to it and switches to that user before the entry point. The command is python -m extract_thinker.mcp_server, port 8000 is exposed, and a health check polls http://127.0.0.1:8000/healthz every thirty seconds after a twenty-second start period.

The compose service then narrows that back down. It builds from the Dockerfile, publishes 127.0.0.1:8000:8000 so the port stays on loopback, mounts the document directory read-only from ${EXTRACT_THINKER_DOCUMENTS:-./examples/service-data}, sets read_only: true with a tmpfs on /tmp, and restarts unless-stopped. So the container listens on every interface internally while the host exposes it on one, which is the right way round.

One setting deserves a second look: the env_file entry for server.env is marked required: false, with a server.env.example at the repository root to copy from. The container therefore starts with no credential file at all, and since a model is selected through EXTRACT_THINKER_MODEL, the service comes up without one. The readme does not say what the health check asserts beyond reachability, nor what the MCP service does on a request with no provider configured, and that is the gap worth filling before anyone runs it.

## A ruff config sits in the root while the dev group installs flake8 and black

The root holds .flake8, .ruff.toml and .pre-commit-config.yaml. The Poetry dev group installs flake8 at ^7.1.2, black at ^24.10.0, ipykernel at ^6.29.5 and pytest at ^8.2.0. Ruff is not among them, so the config file for the faster linter is present in a checkout that cannot run it without an extra install.

Two formatters and two linters in one repository is the kind of overlap CONTRIBUTING.md would settle, and that file plus the pre-commit configuration is where the answer lives rather than in pyproject.toml, where no tool section appears at all. ipykernel in the dev group lines up with the notebooks directory under examples/, and the other example files are plain scripts including a receipt processor, a resume processor and a basic extractor, with an invoice PDF and a service-data directory alongside them.

The contribution guidance has one line worth repeating: the offline core suite runs without provider credentials, and good contributions are described as reduced document fixtures, loader compatibility fixes, examples with expected outputs and honest reports of model or parser limits. That is a sensible standard for a library whose hardest failures are upstream. It is also the one place the project says its output can be wrong, alongside the warning that a validated schema does not guarantee a factually correct result.

Two descriptions of the project do not match each other either. The repository summary calls it a document intelligence library for LLMs offering ORM-style interaction, while the Poetry description reads a library to extract data from files and documents agnositicaly using LLMs.

## Conclusion

ExtractThinker is worth reading if you want the ORM shape for LLM extraction, because the contract, loader and completion strategy are separated cleanly and the offline test suite is designed to run without a provider key. Two things decide whether it is usable for you today. If you need the 2026 capabilities, page retrieval with SQLite, parallel field extraction, model routing, entity masking, page events, the four added loaders or the MCP service, then a wheel install is the wrong artifact and you are installing from a checkout that carries no release tag. And if your documents are scanned rather than digital, the OCR dependency is present in one requirements file and absent from the package metadata, so you are choosing and pinning that yourself. Verify first which version your environment actually resolved, then check that the loader you intend to use is in the dependency set you installed. Teams that want a released, reproducible artifact rather than a moving branch should hold off until a tag catches up.

## FAQ

### How do I install ExtractThinker?

From PyPI with pip install extract-thinker, which resolves to release 0.1.14 published on 2025-06-09. The features described as 2026 additions on main, including the MCP service and the extra loaders, are stated to be not yet a PyPI release, so reaching them means cloning the repository and installing the checkout in editable mode.

### What does ExtractThinker need in order to read a PDF?

The readme says to install pypdf and use DocumentLoaderPyPdf, and pypdf is not among the declared dependencies. The Poetry block declares pypdfium2 instead, and the 2026 work on main adds PyMuPDF, Camelot, Tabula and Adobe loaders. System MIME detection additionally needs libmagic present as a system package.

### Which Python versions does ExtractThinker support?

The core is documented as Python 3.9 to 3.13 with pyproject setting >=3.9,<3.14, and the optional MCP service is documented as needing 3.10 or newer. One declared dependency, tiktoken, carries a marker limited to >=3.9,<3.13, so it is not installed on 3.13.

### Does ExtractThinker verify that extracted values are correct?

It verifies shape, not truth. Fields are declared as a Pydantic contract, so types and constraints such as a non-negative total are enforced on the result. The readme states plainly that schema validity does not guarantee factual accuracy and advises evaluating results on your own documents.

### What does the ExtractThinker MCP container expose?

python -m extract_thinker.mcp_server on port 8000, bound to 0.0.0.0 inside the container while compose publishes it to 127.0.0.1 only. It runs as uid 10001 with a read-only root filesystem, a read-only document mount and a tmpfs on /tmp, and the compose env_file for server.env is marked required: false.

## Sources

- [enoch3712/ExtractThinker on GitHub](https://github.com/enoch3712/ExtractThinker)
- [License: Apache-2.0](https://github.com/enoch3712/ExtractThinker/blob/main/LICENSE)
- [Project website](https://enoch3712.github.io/ExtractThinker)
- [README](https://github.com/enoch3712/ExtractThinker/blob/main/README.md)
- [Releases](https://github.com/enoch3712/ExtractThinker/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/enoch3712-extractthinker
