Model or dataset
Zipstack/unstract avatar
Zipstack/unstract

Unstract: LLM Extraction of Structured JSON from PDFs, Deployed as an API or ETL Pipeline

LLM-Driven Extraction of Unstructured Data — Built for API Deployments & ETL Pipeline Workflows

7,266 stars720 forksPythonAGPL-3.0

At a glance

What is it?
Zipstack's Unstract turns PDFs, scans and images into structured JSON using natural language prompts, then serves the result over a REST API, an ETL pipeline, an MCP server or an n8n node. The Docker quickstart is short, but the AGPL-3.0 licence and the encryption key handling are the parts worth reading before you commit.
Who is it for?
Adopt Unstract if you have a document type that changes often enough that per-vendor templates keep breaking, and you are comfortable running a multi-container Docker stack on your own hardware or paying for the managed cloud. Do not adopt it if a single deterministic parser already handles your documents, or if AGPL-3.0 obligations are incompatible with how you ship.
Can I use it commercially?
Yes, with strict conditions. AGPL-3.0 is a network copyleft licence: if people use a modified version over a network, for example as a hosted service, you must offer them its source code under the same licence.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem Unstract targets: schema drift across vendors

The README frames the problem as a maintenance cost. Without a tool like this, extracting fields from documents means writing regex and building a template per vendor. Every new document type costs days of development. Unstract's claim is that you write a prompt once and it handles variations, and that a new document type takes minutes inside Prompt Studio rather than days.

That framing tells you who the project is for. The README names finance, insurance, healthcare and KYC/compliance, which are domains where the same logical field (an invoice total, a policy number, a patient identifier) arrives in a different visual layout from every counterparty. If your documents come from one generator with one layout, you do not have this problem, and an LLM in the loop adds cost and nondeterminism for no benefit. If your documents come from hundreds of senders, the per-vendor template approach is the thing that breaks.

The second half of the pitch is deployment shape. The README lists API Deployment, ETL Pipeline, an MCP Server for AI agents such as Claude, and an n8n node. That is a deliberate choice to sit behind whatever orchestration you already run rather than asking you to replace it.

How Unstract is put together: four services and a tool registry

The architecture diagram in the README shows four named components: Frontend, Backend, Worker and Platform Service. The repository layout confirms this and adds more: backend/, frontend/, workers/, platform-service/, runner/, tool-sidecar/, tools/, x2text-service/ and a unstract/ directory holding internal packages.

The root pyproject.toml is the clearest evidence of how the pieces relate. It declares the project as unstract version 0.1.0 with requires-python >=3.12,<3.13 and an empty dependencies list, then wires the real code together through [tool.uv.sources] as editable path dependencies: unstract-filesystem, unstract-workflow-execution, unstract-tool-sandbox, unstract-tool-registry, unstract-flags, unstract-core, unstract-connectors and unstract-workers. The root package is a workspace, not a library.

The tool-sandbox and tool-registry packages are the interesting part. Extraction tools run sandboxed and are registered, which is what lets the platform plug in different LLM providers (the README names OpenAI, Anthropic, Bedrock and Ollama) without the workflow layer knowing which one it is talking to. The x2text-service directory points at the text extraction side, and the topics list on the repository includes ocr and pdf-extraction alongside llm and structured-output.

One consequence of this layout: there is no pip install unstract that gives you the platform. The dependencies list is empty on purpose. You run the containers.

Installing Unstract with Docker Compose

The README states Linux or macOS (Intel or M-series), Docker and Docker Compose, 8 GB RAM minimum and Git as prerequisites. The quickstart is a clone followed by a script.

bash
git clone https://github.com/Zipstack/unstract.git
cd unstract
./run-platform.sh

After the script finishes, the README says to visit http://frontend.unstract.localhost and log in with username unstract and password unstract. Those are the defaults shipped with the local setup, so change them before the instance is reachable by anyone else.

run-platform.sh takes flags for the non-default paths. The README documents -v for a version tag, -u to upgrade an existing setup, -b to build images locally, -e to only set up environment files, -p to only pull images, and -d for detached mode. Combining them is supported, for example upgrading to a locally built current tag:

bash
./run-platform.sh -u -b -v current

The README does not document a rollback command for a failed upgrade. If you need to return to a previous release, the mechanism is not described, so plan a database backup before running -u rather than assuming the script reverses itself.

Before you start, read the warning about the encryption key. The README says the ENCRYPTION_KEY value lives in backend/.env or platform-service/.env and that losing it makes existing adapters inaccessible. Copy it somewhere durable at install time, not after you have configured connectors.

Where Unstract is the wrong tool

The clearest limitation is that Unstract is not a library you embed. The root pyproject.toml has an empty dependencies list and resolves everything through editable path sources, so the supported consumption path is the platform: containers, a frontend, a database, a worker queue. If you wanted a function that takes bytes and returns JSON inside an existing Python process, this is a heavier thing than you are looking for.

The second limitation is determinism. Extraction goes through an LLM, and the project explicitly supports swapping providers including a locally hosted Ollama. That flexibility means output quality is a property of your prompt and your chosen model, not a fixed guarantee of the tool. For a field that must be extracted identically every time, such as a regulatory identifier, a deterministic parser is still the safer component, and Unstract is better used around it than instead of it.

The third is the licence. Unstract is AGPL-3.0. That is a strong copyleft licence with a network clause, and it is a different proposition from a permissive licence if you intend to offer the software as a service or link it into a proprietary product. The README links an Enterprise option at unstract.com/pricing, which is the usual shape for an open core project. Whether your specific use triggers the network clause is a question for your own counsel; the repository does not answer it for you.

Finally, the README documents no rollback path for the upgrade flag, and it does not describe what happens to in-flight jobs during an upgrade. Treat upgrades as a maintenance window.

Unstract compared with building on a raw text-extraction layer

The related searches show people comparing Unstract with Unstructured and with LLMWhisperer, and the repository itself has an x2text-service directory, so the honest comparison is between Unstract and a text-extraction library or service that you wrap yourself.

A raw extraction layer does one job: it takes a PDF, scan or image and returns text or layout-aware text. Everything above that (deciding which fields matter, calling a model, validating the JSON, retrying on a bad response, loading the result into a warehouse) is yours to write. You get full control over the prompt, the model, the retry policy and the output schema, and you get a dependency that is small enough to vendor into an existing service.

Unstract bundles that upper layer. Prompt Studio holds the schema as a prompt, the tool registry abstracts the provider, and the deployment targets (REST API, ETL pipeline, MCP server, n8n node) are pre-built. The cost is the platform: four services, Docker Compose, a frontend you have to secure, and an upgrade script with no documented rollback. If your team already has a document pipeline and only lacks the extraction step, the raw layer is the smaller commitment. If you do not have that pipeline and the README's list of deployment shapes matches what you need, Unstract saves you from writing it.

Maintenance, releases and what the licence means in practice

The repository is not archived and the last push was on 2026-09-09, which is recent. Releases are frequent: v0.188.0 on 2026-09-09, v0.187.2 on 2026-09-03 and v0.187.1 on 2026-09-01. A version cadence that tight, with a minor number still below 1.0, means you should expect the platform to move and should pin a tag with -v rather than tracking latest in production.

The upgrade path is the maintenance cost that matters. run-platform.sh -u pulls a newer version over an existing setup, and the README documents no way to go back. Combined with the encryption key warning, that gives you two operational obligations: keep ENCRYPTION_KEY outside the repository, and take a database backup before every upgrade. Neither is optional if you have live adapters configured.

On licensing, AGPL-3.0 is the licence in the LICENSE file. The practical implication for most engineering teams is that modifications you distribute, including over a network in some interpretations, carry source obligations. The README points to a separate Enterprise offering, which typically exists to sell relief from exactly those obligations. Read the licence text and the pricing page together before you build a product on top of the platform, and take your own advice on the network clause rather than an article's.

Editorial conclusion

Adopt Unstract if you have a document type that changes often enough that per-vendor templates keep breaking, and you are comfortable running a multi-container Docker stack on your own hardware or paying for the managed cloud. Do not adopt it if a single deterministic parser already handles your documents, or if AGPL-3.0 obligations are incompatible with how you ship. Before anything else, copy the ENCRYPTION_KEY out of backend/.env and platform-service/.env, because the README states that losing it makes existing adapters inaccessible. Then confirm which LLM provider you will point the platform at, since the extraction quality and the cost per page both come from that choice rather than from Unstract itself.

Frequently asked questions

How do I install Unstract?

Clone the repository, change into it and run ./run-platform.sh. The README lists Linux or macOS, Docker and Docker Compose, 8 GB RAM minimum and Git as prerequisites, and says the platform is then reachable at http://frontend.unstract.localhost.

What is Unstract?

Unstract is a Python project from Zipstack that uses LLMs to extract structured JSON from PDFs, images and scans. The README describes defining the schema with natural language prompts and deploying the result as an API, an ETL pipeline, an MCP server or an n8n node.

Is Unstract free?

The repository is licensed AGPL-3.0 and the README includes a self-hosted quickstart with run-platform.sh, so the platform can be run without paying for it. The README also links a separate Enterprise option at unstract.com/pricing, and the documentation does not describe what that tier includes.

Is Unstract open source?

Yes. The repository is public, the LICENSE file is AGPL-3.0, and the README documents building images locally with ./run-platform.sh -b -v current. Note that AGPL-3.0 carries source obligations that differ from permissive licences.

How does Unstract compare with Unstructured?

The repository does not document a comparison. Based on the layout, Unstract includes an x2text-service for text extraction plus a tool registry, a workflow layer and deployment targets, whereas a text-extraction layer alone returns text and leaves schema definition, model calls and loading to you.

Official sources

  1. License: AGPL-3.0
  2. Project website
  3. README
  4. Releases
  5. Zipstack/unstract on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/zipstack-unstract.svg)](https://hysenlabs.com/projects/zipstack-unstract)