# textract: extract text from documents in Python without writing a parser per format

> The Python library that wraps format-specific extractors behind one call, what its dependency list and Dockerfile reveal about the trade-offs, and when you should reach for something else.

**deanmalmgren/textract** — extract text from any document. no muss. no fuss.

- Repository: https://github.com/deanmalmgren/textract
- Website: http://textract.readthedocs.io
- Stars: 4,726 · Forks: 720
- Language: HTML
- License: MIT
- Published: 2026-09-23 · Updated: 2026-09-23 · Language: en
- Canonical page: https://hysenlabs.com/projects/deanmalmgren-textract

## The format-dispatch problem textract is built to remove

Extracting text from a folder of mixed documents is rarely one problem. A PDF needs one library. A .docx needs another. A spreadsheet needs a third, and legacy .xls needs a fourth because the modern library dropped that format. HTML needs a parser, and an audio file needs a speech recogniser plus a codec. Each of those libraries has its own entry point, its own return type and its own way of failing.

textract's stated goal is to collapse that into a single call. The README puts it as "Extract text from any document. No muss. No fuss." The audience is Python developers doing text mining or natural language processing who receive documents they did not produce and do not want to maintain a dispatch table themselves. The project is tagged data-mining, natural-language-processing, text-mining and python, which matches that reading.

## How the dependency list encodes the format matrix

The mechanism is visible in pyproject.toml rather than in prose. Each supported format maps to a declared dependency: pdfminer.six for PDF, docx2txt for Word, python-pptx for PowerPoint, openpyxl and xlrd for spreadsheets, odfpy for OpenDocument, extract-msg for Outlook mail, beautifulsoup4 with lxml for HTML, chardet for encoding detection, Pillow for images, and SpeechRecognition for audio.

The comments in that list are the most informative part of the repository. openpyxl "Replaces xlrd for Excel file support (xlrd 2.0+ dropped xlsx support)" while xlrd remains "Supports legacy .xls files (openpyxl does not support .xls)". That is two spreadsheet libraries kept side by side because neither covers the other's territory. odfpy is annotated "Supports .ods files". Nothing about that is elegant, but it is honest about what the format matrix actually costs.

The architecture follows: textract is a dispatcher, not an extractor. It decides which backend handles a file and normalises the result into text. That means its output quality is bounded by the backends, and a bug in pdfminer.six surfaces to you as a textract bug. The upside is that you inherit upstream fixes for free.

## Installing textract with uv and extracting your first file

The repository uses uv as its build backend and lockfile tool, and the Makefile's sync target runs uv with all extras. Python 3.10 or newer is required according to pyproject.toml.

```bash
uv sync --all-extras
```

On macOS the sync target does one extra thing: it codesigns the pocketsphinx shared object inside the virtual environment, matching the pattern _pocketsphinx.cpython-*-darwin.so. Without that step, audio extraction on macOS can fail at load time. The Makefile wraps it in a shell conditional on uname so it is a no-op elsewhere.

The package also exposes a command line entry point, registered in pyproject.toml as textract.bin.textract:main. The Dockerfile uses it directly, with ENTRYPOINT ["textract"] and CMD ["--help"], so running the image with no arguments prints usage rather than doing work.

```dockerfile
FROM python:3.13-slim-trixie AS builder
COPY --from=ghcr.io/astral-sh/uv:0.12.3 /uv /uvx /usr/local/bin/
ENV UV_PYTHON_DOWNLOADS=0
WORKDIR /app
COPY pyproject.toml uv.lock README.rst ./
COPY textract ./textract
RUN uv sync --locked --no-dev --no-editable --extra pocketsphinx
```

The Dockerfile pins Python 3.13 rather than the newest interpreter, and the comment explains why: pocketsphinx has no prebuilt wheel for 3.14 on Linux yet, and building it from source needs a C toolchain. That is a concrete constraint, not a stylistic choice.

The runtime stage installs the system binaries the Python packages shell out to: ghostscript, poppler-utils, tesseract-ocr, unrtf, sox, libsox-fmt-mp3 and libportaudio2. This is the part people miss. Installing the Python package alone does not give you PDF, OCR or audio extraction; those tools have to exist on the machine. The Dockerfile deliberately excludes LibreOffice, and its comment points to docs/installation.rst#converting-legacy-doc-files for anyone who needs legacy .doc support.

## Where textract breaks and where it is the wrong choice

The first limitation is environmental. Because extraction depends on external binaries, textract is not a pure-Python install you can drop into a serverless function and forget. A minimal container needs poppler, tesseract and the rest, which inflates image size and gives you more surface to patch. If your deployment target forbids apt packages, this library is the wrong tool regardless of how good the Python API is.

The second is that the dependency list is long and each entry is a maintenance commitment. Pillow, lxml, pdfminer.six, python-pptx, openpyxl, xlrd, odfpy, extract-msg and SpeechRecognition all move independently. The pyproject.toml even carries a vulnerable_packages group pinning floor versions for Pygments, cryptography, idna, requests and urllib3, which is a signal that transitive dependencies need watching.

The third is scope. textract extracts text. It does not preserve layout, tables or reading order in any documented way, and the README does not describe how it handles scanned PDFs versus text PDFs, or what happens when a backend returns nothing. If your PDFs are scans, the OCR path depends on tesseract being installed and configured, and the documentation gives no accuracy guidance. For a single format you already handle, adding textract adds a dispatcher and eight dependencies to solve a problem you do not have.

## textract versus tesseract, and against parsing formats yourself

The most common comparison is textract against tesseract, and the two are not at the same level. Tesseract is an OCR engine: you give it an image or a rasterised page and it returns text. textract is an orchestration layer that may call tesseract for you but also handles formats where OCR is irrelevant, such as .docx, .xlsx and .html. Choosing tesseract directly makes sense when every input is an image or a scan and you want control over page segmentation and language models. Choosing textract makes sense when the inputs are heterogeneous and OCR is one branch among many.

The other alternative is writing the dispatch yourself: pdfminer.six for PDFs, docx2txt for Word, openpyxl for xlsx. That gives you exact control over each backend's options and a smaller dependency footprint for the formats you actually receive. The cost is that you now own the format detection, the encoding handling and the error paths, which is precisely the code textract already contains. The honest split is whether your format list is stable and short, in which case roll your own, or long and growing, in which case the dispatcher earns its keep.

## Maintenance, releases and what the MIT licence covers

The repository is not archived, and the last push was on 2026-09-01, the same day v2.1.0 was released. Before that, v2.0.0 and v2.0.0rc1 both landed on 2026-04-27. That is a real release cadence rather than a dormant project, but it is also a project whose version history shows a major bump this year, so pin your version and read the changelog before upgrading across it.

The licence is MIT, stated in both the README badge block and the license field of pyproject.toml. MIT is permissive: it allows commercial use and modification with attribution and no warranty. The practical implication for adopters is that the dependencies, not textract itself, carry the licence risk. pdfminer.six, openpyxl, xlrd, odfpy and the rest each ship under their own terms, and some have changed licences across major versions in the past. Check those separately; this is not legal advice, and the MIT grant on textract says nothing about what its backends permit.

Upgrade cost is dominated by the system binaries. When tesseract or poppler changes behaviour, textract's output can change without any Python package moving. That makes version pinning at the container level, not just the lockfile level, the safer habit.

## Conclusion

Adopt textract when you have a mixed pile of office documents, PDFs and HTML and you want one Python call instead of five libraries, and when you can accept the system-level dependencies that come with it. Do not adopt it if your input is a single format you already parse well, or if you cannot install poppler, tesseract and the other binaries on your target machines. Before committing, run the library against one real file of each format you care about and confirm the extracted text, because the README does not document per-format accuracy or failure behaviour.

## FAQ

### What is textract used for?

It extracts text from documents in many formats through one Python interface, which makes it useful for text mining and natural language processing pipelines that receive mixed file types. The README describes it as extracting text from any document with no fuss.

### How do I install textract in Python?

The repository uses uv, and the Makefile's sync target runs uv sync --all-extras. Python 3.10 or newer is required according to pyproject.toml. On macOS, that target also codesigns the pocketsphinx shared object in the virtual environment.

### Is textract free?

The library is MIT licensed, so it is free to use and modify under those terms. That covers textract itself; its backends such as pdfminer.six, openpyxl and odfpy ship under their own licences.

### What is textract?

It is a Python library and command line tool for extracting text from documents in many formats, built as a dispatcher over format-specific backends. The project is MIT licensed and hosted on GitHub under deanmalmgren/textract.

## Sources

- [deanmalmgren/textract on GitHub](https://github.com/deanmalmgren/textract)
- [License: MIT](https://github.com/deanmalmgren/textract/blob/master/LICENSE)
- [Project website](http://textract.readthedocs.io)
- [README](https://github.com/deanmalmgren/textract/blob/master/README.md)
- [Releases](https://github.com/deanmalmgren/textract/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/deanmalmgren-textract
