Unstructured: Open-Source ETL for Feeding PDFs and Documents to Language Models
Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.
At a glance
- What is it?
- Unstructured is a Python library and an optional ETL pipeline service for converting PDFs, Word documents, HTML, images, and 60-plus other file formats into clean text elements that language models can consume. The open-source library runs locally; a commercial platform and an MCP server extend it to production workflows.
- Who is it for?
- Unstructured fits teams that need to extract clean text from diverse document formats before feeding them to a language model, particularly when Docker-based deployment or a pip extra install is acceptable. Projects requiring only plain text, HTML, XML, JSON, and email need no extra dependencies; other formats require the relevant pip extras and, for the Docker path, a large image that bundles Tesseract and LibreOffice.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 15 days ago.
- What is it written in?
- Mainly HTML, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 17, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Unstructured Solves and Who It Is For
Raw documents do not feed cleanly into language models. A PDF may contain scanned images, multi-column layouts, embedded tables, and mixed encoding. The Unstructured library addresses this by providing partition functions for each document type that extract text, table data, and metadata as a list of typed elements.
The README describes the library's purpose as providing open-source components for ingesting and pre-processing images and text documents such as PDFs, HTML, Word docs, and many more. The use cases revolve around preparing data for LLM pipelines, retrieval-augmented generation, and vector database ingestion. The pyproject.toml classifies the project under Intended Audience :: Developers, Intended Audience :: Education, and Intended Audience :: Science/Research, and lists NLP, PDF, HTML, CV, XML, parsing, and preprocessing as keywords.
The project requires Python 3.11 or newer (the pyproject.toml specifies >=3.11, <3.14) and supports Python 3.11, 3.12, and 3.13.
Installing the Library and Running the First Partition
The simplest install covers all document types:
pip install "unstructured[all-docs]"For plain text files, HTML, XML, JSON, and emails that require no extra dependencies:
pip install unstructuredFor a specific document type, install only the required extras:
pip install "unstructured[docx,pptx]"Once installed, the partition functions map to document types:
python3>>> from unstructured.partition.pdf import partition_pdf
>>> elements = partition_pdf(filename="example-docs/layout-parser-paper-fast.pdf")
>>> from unstructured.partition.text import partition_text
>>> elements = partition_text(filename="example-docs/fake-text.txt")The README notes that system dependencies such as poppler, tesseract, and LibreOffice may be needed depending on the document types being parsed. The Dockerfile shows the full set installed on the base image: libxml2, glib, mesa-gl, cmake, libmagic, wget, git, openjpeg, poppler, poppler-utils, poppler-glib, libreoffice, and tesseract with tessdata language models.
Running Unstructured in Docker
The README provides a Docker path for teams that prefer a pre-configured environment. Pull the official image:
docker pull downloads.unstructured.io/unstructured-io/unstructured:latestCreate a running container and open a shell inside it:
docker run -dt --name unstructured downloads.unstructured.io/unstructured-io/unstructured:latest
docker exec -it unstructured bashThe Dockerfile uses a Chainguard wolfi-base image as its foundation. The base image is updated regularly, and the README warns that local docker build could fail due to upstream changes in wolfi-base. Multi-platform images are built for both x86_64 and Apple silicon hardware; the --platform flag can override the default.
Alternatively, build the image locally:
make docker-build
make docker-start-bashThe Unstructured Transform MCP Server
The README describes a newer offering called Unstructured Transform, which is an MCP server that brings document processing to AI agents. The setup requires picking an MCP-compatible client (Claude Code, Cursor, Codex CLI, or similar), adding the Transform MCP server to the client's configuration via the mcp add command or the client's settings file, authenticating once, and then pointing the agent at a file or URL.
The README describes the MCP server as handling 60-plus formats including PDFs, emails, images, and scanned files. The agent describes the processing intent in plain language and Transform runs the partition, enrichment, chunking, and embedding steps. This is a separate product from the open-source library; the README links to transform.unstructured.io to get started.
Dependency Structure and Optional Extras
The pyproject.toml lists core dependencies that apply to all installations, including beautifulsoup4, langdetect, lxml, numpy, spacy, and rapidfuzz. Each document type is an optional extra: csv requires pandas, docx requires python-docx, pdf requires pdfminer.six and pdf2image, xlsx requires openpyxl, and so on.
The all-docs extra bundles all document type dependencies for convenience, at the cost of a larger install. The Makefile shows separate test targets for each extra (test-extra-csv, test-extra-docx, test-extra-epub, test-extra-markdown, and others), which reflects the optional nature of each dependency group. The uv lock file is the source of truth for the full dependency tree, managed with uv sync --locked --all-extras --all-groups.
When Unstructured Is the Wrong Tool
Unstructured is a preprocessing step; it does not store, index, or search documents. Teams looking for a complete RAG pipeline need to pair it with a vector database and an embedding model, neither of which is included.
The full Docker image is large because it bundles Tesseract with trained models, LibreOffice, and poppler. Teams that only need plain text or HTML parsing can install the bare unstructured package and avoid that overhead, but document-heavy workflows with scanned PDFs will need the heavy dependencies.
Apache Tika is an alternative document parser from the Apache Software Foundation that also handles a broad set of formats but is written in Java and runs as a server process. The architectural difference is that Tika is a standalone Java service that clients call over HTTP, while Unstructured is a Python library that runs in the same process as the calling code, making it easier to integrate into a pure Python data pipeline.
Maintenance, Licence, and the Commercial Platform
The last push to the main branch was on 2026-09-15. The three most recent releases are 0.27.8 (2026-09-22), 0.27.9 (2026-09-26), and 0.27.10 (2026-09-27), showing active release activity. The project requires Python 3.11 or newer.
The open-source library is released under the Apache-2.0 licence. The pyproject.toml classifies the project as Development Status :: 4 - Beta. The commercial platform at unstructured.io is a separate product offering production-grade workflows with chunking, embedding, image and table enrichment, and a low-code UI. A demo can be requested from the sales team. The README links to both the open-source library documentation and the commercial platform, keeping the two clearly separated.
Editorial conclusion
Unstructured fits teams that need to extract clean text from diverse document formats before feeding them to a language model, particularly when Docker-based deployment or a pip extra install is acceptable. Projects requiring only plain text, HTML, XML, JSON, and email need no extra dependencies; other formats require the relevant pip extras and, for the Docker path, a large image that bundles Tesseract and LibreOffice. The open-source library is under Apache-2.0; the commercial platform and MCP server are separate products. The last push to main was on 2026-09-15, with v0.27.10 released on 2026-09-27.
Frequently asked questions
Is Unstructured free to use?
The Unstructured open-source library is released under the Apache-2.0 licence, which permits free use and modification. The commercial Unstructured Pipelines platform at unstructured.io is a separate product; the README links to a sales demo request for pricing details.
How do I install the Unstructured library in Python?
Install with pip install "unstructured[all-docs]" to include support for all document types. For plain text, HTML, XML, JSON, and email only, pip install unstructured requires no extra system dependencies. Additional system packages such as poppler and tesseract may be required for PDF and image parsing.
How do I use the Unstructured library?
Import the partition function for your document type, for example from unstructured.partition.pdf import partition_pdf, and call it with a filename argument. The function returns a list of typed elements containing the extracted text and metadata.
Official sources
Where this project is recommended
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/unstructured-io-unstructured)