Model or dataset
ucbepic/docetl avatar
ucbepic/docetl

DocETL: Declarative, Agentic Map-Reduce for LLM Data Processing

A system for agentic LLM-powered data processing and ETL

4,114 stars445 forksPythonMIT

At a glance

What is it?
DocETL turns LLM document processing into a declared pipeline of map, reduce, filter and resolve operations, with an optimizer that rewrites prompts and swaps models. It is a good fit for Python teams with a real corpus and a real API budget.
Who is it for?
Adopt DocETL if you already have a corpus of tickets, PDFs or transcripts, you are comfortable in Python, and you want the map/reduce/filter plumbing and the MOAR optimizer instead of hand-written LLM calls. Skip it if your transformation is a single prompt over a handful of rows, or if you cannot send the data to a hosted provider, since the README only shows hosted keys and the Docker service.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 26 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem DocETL addresses: LLM calls as pipeline stages, not scripts

The README frames the alternative plainly: without DocETL you write each LLM call yourself, wire them together, and tune accuracy, cost and latency by hand. That is the actual pain. Once a corpus grows past a few hundred documents, a notebook full of prompt strings stops being maintainable, because retries, rate limits, batching and schema validation all leak into the calling code.

DocETL's answer is to make the pipeline the unit of work. You describe each stage in natural language, as the README puts it, something like pulling out every complaint in a ticket, and the framework supplies the operators and orchestrates them across your data. The audience is therefore specific: engineers who already have documents and a working API key, and who want the plumbing handled. It is not aimed at someone who needs one classification over twenty rows; a single API call does that.

Map, reduce and the rest: how a DocETL pipeline is actually structured

The mechanism is a chain of typed operations over a table. In the Python API, docetl.read_json returns a pipeline object, and each method appends an operation and returns a new pipeline. The README example maps over tickets with a prompt and an output schema declaring category and priority as strings, then reduces on reduce_key="category" so that each group is summarized separately. The reduce prompt uses a Jinja-style loop over inputs, which means the grouping is done by the framework and the prompt only sees one group at a time.

Two details matter more than they look. First, output schemas are declared, so the framework can validate what the model returns rather than trusting free text. Second, pipeline.schema() reports the resulting columns, which the README shows as {'category': 'str', 'summary': 'str'}, and pipeline.show() runs on five documents before you pay for the full set. That dry-run step is the part most hand-rolled scripts lack.

The YAML path expresses the same thing declaratively: a datasets block, a default_model, an operations list, and a pipeline with steps and an output. The entry points in pyproject.toml enumerate the operation types under the docetl.operation group, including map, parallel_map and filter, and the documentation index lists map, filter, reduce, resolve, split, gather and extract. Optimization is handled by MOAR, described in the README as automatic cost-accuracy optimization and covered by a separate guide and a VLDB 2026 paper.

Installing DocETL and running a first pipeline

Installation is a pip install plus one provider key. The README gives exactly this pair, and notes that any LLM provider key works because the dependency list includes litellm.

bash
pip install docetl
export OPENAI_API_KEY=your_key   # or any LLM provider key

After that, build a pipeline in Python. The following is the README's own example, trimmed to the classification step. Reading tickets.json, appending a map with a schema, then printing the schema is the smallest thing that proves the wiring works.

python
import docetl
docetl.default_model = "gpt-4o-mini"

pipeline = docetl.read_json("tickets.json")
pipeline = pipeline.map(
    prompt="Classify this support ticket: {{ input.text }}",
    output={"schema": {"category": "str", "priority": "str"}},
)
print(pipeline.schema())

Before a full run, call pipeline.show() to execute on five documents and print the results, then pipeline.collect() for the whole dataset. The README follows collect with a cost print, pipeline.total_cost, which is the number to watch while you calibrate. If you prefer a config file, the README's YAML form declares the same map operation with a schema of category and priority, and runs it with the CLI.

bash
docetl run pipeline.yaml

For interactive prompt work there is DocWrangler, either at docetl.org/playground or run locally. The repository ships a docker-compose.yml that builds the image and maps host ports 3031 to the frontend and 8081 to the backend by default, with a healthcheck against http://localhost:8000/health.

Where DocETL gets expensive or awkward

The honest limitation is that every operator is an LLM call, and the framework's value proposition depends on that being acceptable. Rate limiting is something you configure rather than something you get: the README sets docetl.rate_limits with an llm_call entry and an llm_tokens entry, both expressed as a count per unit of time. If you skip that, a large collect() run is bounded only by your provider's own limits and your wallet. The dry-run helpers exist precisely because full runs are not cheap to repeat.

There is also a dependency weight problem. pyproject.toml lists a long runtime set including litellm, openai-agents, pandas, scikit-learn, matplotlib, nltk, scipy, modal and boto3, and the extras add more: parsing pulls in pymupdf, python-docx and paddlepaddle, while server pulls in fastapi, docling and playwright. Installing docetl into an existing slim environment is not a small change, and the parsing and server extras in particular are heavy.

Finally, the README does not document rollback or checkpoint resumption for a partially completed run, so a failure midway through collect() is not something you can plan around from the README alone. Treat that as unverified rather than as a guarantee either way.

DocETL versus writing your own LLM calls or using a general orchestrator

The realistic alternative is a general workflow tool such as a DAG orchestrator, or simply a Python script with your own retry loop. The difference is where the intelligence sits. A general orchestrator schedules tasks you have already defined; it does not look at your prompt and decide that part of it would be cheaper as code, or that a smaller model would do. That rewriting is DocETL's specific claim, and it is why the project has its own optimizer paper rather than being a thin wrapper.

The second difference is the operator vocabulary. Map, filter, reduce, resolve, split, gather and extract are semantic operations over documents, with reduce_key grouping and declared output schemas. Reproducing that on top of a generic task runner means writing the grouping, the schema validation and the batching yourself. If your pipeline is genuinely a sequence of shell commands and API calls with no document semantics, a general orchestrator is the smaller dependency. If it is document analysis, DocETL is closer to the shape of the problem.

Maintenance, licensing and what upgrading costs you

The repository is not archived, and the last push was on 2026-09-05, which is recent. Releases are spaced rather than continuous: 0.2.5 on 2025-08-09, 0.2.6 on 2025-12-28, and 0.3.0 on 2026-06-17. That cadence suggests you should pin a version rather than track the default branch, and expect to re-verify your pipeline after a minor bump, since the 0.2.x to 0.3.0 jump is a minor version and the package ships a CLI, a Python API and a server in one distribution.

Licensing is MIT, declared both in the README badge and in pyproject.toml under license = { text = "MIT" }. That is permissive and puts few obligations on how you redistribute or host the software. It says nothing about your data or your provider's terms, and the dependency list includes hosted services such as modal and boto3 alongside the LLM providers, so the compliance question that matters here is about data leaving your environment, not about the DocETL licence itself. This is not legal advice; check your own obligations.

Upgrade cost is dominated by the extras. Because parsing and server pull in large packages, moving between versions can change your environment substantially even when the API you use is stable.

Editorial conclusion

Adopt DocETL if you already have a corpus of tickets, PDFs or transcripts, you are comfortable in Python, and you want the map/reduce/filter plumbing and the MOAR optimizer instead of hand-written LLM calls. Skip it if your transformation is a single prompt over a handful of rows, or if you cannot send the data to a hosted provider, since the README only shows hosted keys and the Docker service. Before committing, verify the per-run cost on your own data by checking pipeline.total_cost, confirm which operator types you need are listed under docetl.operation entry points in pyproject.toml, and read the optimization guide to see whether MOAR is a separate step or part of the run.

Frequently asked questions

What is AI ETL and how does DocETL fit into it?

AI ETL applies language models to the transform step of extract, transform and load, so unstructured inputs become structured rows. DocETL supplies that transform layer as declared operations such as map, reduce and filter, and returns tables the README describes as easy to query in your favorite database.

Is Python used for ETL in DocETL?

Yes. DocETL is a Python package installed with pip install docetl, and the README calls the Python API the recommended path for production code, notebooks and scripting. A YAML configuration file is offered as the low-code alternative.

How do I install DocETL and run a pipeline?

Run pip install docetl, then export a provider key such as OPENAI_API_KEY. Build a pipeline in Python with docetl.read_json followed by map and reduce, or declare it in YAML and run it with docetl run pipeline.yaml.

Can I try a DocETL pipeline before paying for a full run?

The README shows pipeline.show(), which runs on five documents and prints the results, and pipeline.schema() to check the output columns first. After a full pipeline.collect() run, pipeline.total_cost reports what the run cost.

Which LLM providers does DocETL support?

The README says any LLM provider key works, and the .env.example lists OpenAI, Anthropic, Google and Cohere variables. The project depends on LiteLLM, which the environment file describes as supporting over 100 models.

Is there a DocETL UI?

Yes, DocWrangler is a visual playground for interactive prompt development where you edit prompts and see results in real time. The README points to docetl.org/playground or a local run, and the repository includes a docker-compose.yml for it.

Official sources

  1. License: MIT
  2. Project website
  3. README
  4. Releases
  5. ucbepic/docetl on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/ucbepic-docetl.svg)](https://hysenlabs.com/projects/ucbepic-docetl)