# Unstructured: Open-Source ETL for Feeding PDFs and Documents to Language Models

> Unstructured is a Python library and an optional ETL pipeline service for converting PDFs, Word documents, HTML, images, and 60-plus other file formats into clean text elements that language models can consume. The open-source library runs locally; a commercial platform and an MCP server extend it to production workflows.

**Unstructured-IO/unstructured** — Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models.  Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.

- Repository: https://github.com/Unstructured-IO/unstructured
- Website: https://www.unstructured.io/
- Stars: 15,442 · Forks: 1,329
- Language: HTML
- License: Apache-2.0
- Published: 2026-09-09 · Updated: 2026-09-09 · Language: en
- Canonical page: https://hysenlabs.com/projects/unstructured-io-unstructured

## What Unstructured Solves and Who It Is For

Raw documents do not feed cleanly into language models. A PDF may contain scanned images, multi-column layouts, embedded tables, and mixed encoding. The Unstructured library addresses this by providing partition functions for each document type that extract text, table data, and metadata as a list of typed elements.

The README describes the library's purpose as providing open-source components for ingesting and pre-processing images and text documents such as PDFs, HTML, Word docs, and many more. The use cases revolve around preparing data for LLM pipelines, retrieval-augmented generation, and vector database ingestion. The pyproject.toml classifies the project under Intended Audience :: Developers, Intended Audience :: Education, and Intended Audience :: Science/Research, and lists NLP, PDF, HTML, CV, XML, parsing, and preprocessing as keywords.

The project requires Python 3.11 or newer (the pyproject.toml specifies >=3.11, <3.14) and supports Python 3.11, 3.12, and 3.13.

## Installing the Library and Running the First Partition

The simplest install covers all document types:

```bash
pip install "unstructured[all-docs]"
```

For plain text files, HTML, XML, JSON, and emails that require no extra dependencies:

```bash
pip install unstructured
```

For a specific document type, install only the required extras:

```bash
pip install "unstructured[docx,pptx]"
```

Once installed, the partition functions map to document types:

```bash
python3
```

```python
>>> from unstructured.partition.pdf import partition_pdf
>>> elements = partition_pdf(filename="example-docs/layout-parser-paper-fast.pdf")

>>> from unstructured.partition.text import partition_text
>>> elements = partition_text(filename="example-docs/fake-text.txt")
```

The README notes that system dependencies such as poppler, tesseract, and LibreOffice may be needed depending on the document types being parsed. The Dockerfile shows the full set installed on the base image: libxml2, glib, mesa-gl, cmake, libmagic, wget, git, openjpeg, poppler, poppler-utils, poppler-glib, libreoffice, and tesseract with tessdata language models.

## Running Unstructured in Docker

The README provides a Docker path for teams that prefer a pre-configured environment. Pull the official image:

```bash
docker pull downloads.unstructured.io/unstructured-io/unstructured:latest
```

Create a running container and open a shell inside it:

```bash
docker run -dt --name unstructured downloads.unstructured.io/unstructured-io/unstructured:latest
docker exec -it unstructured bash
```

The Dockerfile uses a Chainguard wolfi-base image as its foundation. The base image is updated regularly, and the README warns that local docker build could fail due to upstream changes in wolfi-base. Multi-platform images are built for both x86_64 and Apple silicon hardware; the --platform flag can override the default.

Alternatively, build the image locally:

```bash
make docker-build
make docker-start-bash
```

## The Unstructured Transform MCP Server

The README describes a newer offering called Unstructured Transform, which is an MCP server that brings document processing to AI agents. The setup requires picking an MCP-compatible client (Claude Code, Cursor, Codex CLI, or similar), adding the Transform MCP server to the client's configuration via the mcp add command or the client's settings file, authenticating once, and then pointing the agent at a file or URL.

The README describes the MCP server as handling 60-plus formats including PDFs, emails, images, and scanned files. The agent describes the processing intent in plain language and Transform runs the partition, enrichment, chunking, and embedding steps. This is a separate product from the open-source library; the README links to transform.unstructured.io to get started.

## Dependency Structure and Optional Extras

The pyproject.toml lists core dependencies that apply to all installations, including beautifulsoup4, langdetect, lxml, numpy, spacy, and rapidfuzz. Each document type is an optional extra: csv requires pandas, docx requires python-docx, pdf requires pdfminer.six and pdf2image, xlsx requires openpyxl, and so on.

The all-docs extra bundles all document type dependencies for convenience, at the cost of a larger install. The Makefile shows separate test targets for each extra (test-extra-csv, test-extra-docx, test-extra-epub, test-extra-markdown, and others), which reflects the optional nature of each dependency group. The uv lock file is the source of truth for the full dependency tree, managed with uv sync --locked --all-extras --all-groups.

## When Unstructured Is the Wrong Tool

Unstructured is a preprocessing step; it does not store, index, or search documents. Teams looking for a complete RAG pipeline need to pair it with a vector database and an embedding model, neither of which is included.

The full Docker image is large because it bundles Tesseract with trained models, LibreOffice, and poppler. Teams that only need plain text or HTML parsing can install the bare unstructured package and avoid that overhead, but document-heavy workflows with scanned PDFs will need the heavy dependencies.

Apache Tika is an alternative document parser from the Apache Software Foundation that also handles a broad set of formats but is written in Java and runs as a server process. The architectural difference is that Tika is a standalone Java service that clients call over HTTP, while Unstructured is a Python library that runs in the same process as the calling code, making it easier to integrate into a pure Python data pipeline.

## Maintenance, Licence, and the Commercial Platform

The last push to the main branch was on 2026-09-15. The three most recent releases are 0.27.8 (2026-09-22), 0.27.9 (2026-09-26), and 0.27.10 (2026-09-27), showing active release activity. The project requires Python 3.11 or newer.

The open-source library is released under the Apache-2.0 licence. The pyproject.toml classifies the project as Development Status :: 4 - Beta. The commercial platform at unstructured.io is a separate product offering production-grade workflows with chunking, embedding, image and table enrichment, and a low-code UI. A demo can be requested from the sales team. The README links to both the open-source library documentation and the commercial platform, keeping the two clearly separated.

## Conclusion

Unstructured fits teams that need to extract clean text from diverse document formats before feeding them to a language model, particularly when Docker-based deployment or a pip extra install is acceptable. Projects requiring only plain text, HTML, XML, JSON, and email need no extra dependencies; other formats require the relevant pip extras and, for the Docker path, a large image that bundles Tesseract and LibreOffice. The open-source library is under Apache-2.0; the commercial platform and MCP server are separate products. The last push to main was on 2026-09-15, with v0.27.10 released on 2026-09-27.

## FAQ

### Is Unstructured free to use?

The Unstructured open-source library is released under the Apache-2.0 licence, which permits free use and modification. The commercial Unstructured Pipelines platform at unstructured.io is a separate product; the README links to a sales demo request for pricing details.

### How do I install the Unstructured library in Python?

Install with pip install "unstructured[all-docs]" to include support for all document types. For plain text, HTML, XML, JSON, and email only, pip install unstructured requires no extra system dependencies. Additional system packages such as poppler and tesseract may be required for PDF and image parsing.

### How do I use the Unstructured library?

Import the partition function for your document type, for example from unstructured.partition.pdf import partition_pdf, and call it with a filename argument. The function returns a list of typed elements containing the extracted text and metadata.

## Sources

- [License: Apache-2.0](https://github.com/Unstructured-IO/unstructured/blob/main/LICENSE)
- [Project website](https://www.unstructured.io/)
- [README](https://github.com/Unstructured-IO/unstructured/blob/main/README.md)
- [Releases](https://github.com/Unstructured-IO/unstructured/releases)
- [Unstructured-IO/unstructured on GitHub](https://github.com/Unstructured-IO/unstructured)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/unstructured-io-unstructured
