# spotify/luigi: a Python task scheduler for batch pipelines, not a data engine

> Luigi is a Python package for building dependency graphs of batch jobs, with a central scheduler, a web UI and Hadoop helpers. It coordinates work you already have; it does not process data itself.

**spotify/luigi** — Luigi is a Python module that helps you build complex pipelines of batch jobs. It handles dependency resolution, workflow management, visualization etc. It also comes with Hadoop support built in.

- Repository: https://github.com/spotify/luigi
- Stars: 18,782 · Forks: 2,467
- Language: Python
- License: Apache-2.0
- Published: 2026-09-21 · Updated: 2026-09-21 · Language: en
- Canonical page: https://hysenlabs.com/projects/spotify-luigi

## What spotify/luigi solves, and for whom

Luigi addresses the plumbing around long-running batch processes. The README is explicit about the target: you want to chain many tasks, automate them, and "failures will happen". The tasks are typically Hadoop jobs, database dumps, machine learning runs or anything else that takes minutes to days. Luigi does not replace Hive, Pig or Cascading; the README says it is not a framework to replace these lower level tools but one that stitches tasks together, where a single task can be a Hive query, a Hadoop job in Java, a Spark job in Scala or Python, a Python snippet, or a table dump.

The audience is therefore engineers who already own the execution layer. If you have a warehouse, a cluster and a pile of scripts that must run in a particular order, Luigi supplies the ordering, the retry behaviour, the command line entry points and a visualiser. Spotify states it runs thousands of tasks every day internally on complex dependency graphs, mostly Hadoop jobs, powering recommendations, toplists, A/B test analysis, reports and dashboards. That is a description of a company with an existing batch estate, not of a greenfield project.

## The dependency graph lives in Python, not in XML

The design decision that shapes everything else is stated in the README's philosophy section: the dependency graph is specified within Python rather than in XML or another external configuration format. There are similarities to GNU Make in the task-and-dependency model, and to Oozie and Azkaban in the workflow space, but Luigi is not built specifically for Hadoop and is meant to be extended with other kinds of tasks. Because dependencies are Python objects, they can involve date algebra or recursive references to other versions of the same task. A daily job can declare its dependency on yesterday's job without a separate templating layer.

The mechanism behind that is the task class. A task declares its parameters and its requirements, and the scheduler resolves the graph from those declarations. The repository ships a luigid entry point in pyproject.toml that starts the central scheduler, and the README notes that the Luigi server comes with a web interface so you can search and filter among all tasks. The visualiser renders the graph: completed tasks appear green, tasks yet to run appear yellow, and the README's example screenshot comes from a production graph where most nodes are Hadoop jobs and some run locally to build data files.

One consequence deserves emphasis. Luigi passes dependencies, not data. A downstream task knows that an upstream task finished; it does not receive the upstream output as an argument. The README does not describe a data-passing channel between tasks, and the file system abstractions for HDFS and local files are described as ensuring atomic operations, which is the mechanism that prevents a pipeline from crashing in a state containing partial data. If your design assumes rows flowing from one task into the next, that assumption is yours to build, not something the framework supplies.

## Installing Luigi and running a first pipeline

The README gives the install path directly. The stable release comes from PyPI, and there is a variant with TOML configuration support.

```bash
pip install luigi
```

If you want TOML-based configs, the README points at the extra:

```bash
pip install luigi[toml]
```

For the bleeding edge code the README offers a git install, and notes that bleeding edge documentation is hosted separately from the stable documentation:

```bash
pip install git+https://github.com/spotify/luigi.git
```

The repository contains an examples directory with a hello_world.py and a top_artists.py, which are the files to read first if you want a working task definition rather than a description of one. The packaging metadata declares Python >=3.10, <3.15, and lists the console scripts luigi, luigid, luigi-grep, luigi-deps and luigi-deps-tree. Running luigid starts the scheduler and the web interface described in the README; running luigi executes tasks against it. The examples/config.toml file in the tree is the reference for the TOML configuration format that the toml extra enables, and examples/dynamic_requirements.py and examples/per_task_retry_policy.py show the two features most teams need early: requirements that are computed at runtime, and retry behaviour set per task rather than globally.

## Where Luigi is the wrong tool

The clearest limitation is the one the README states as a design boundary rather than a defect: Luigi is not a framework to replace Hive, Pig or Cascading. It coordinates; it does not compute. Teams that adopt it expecting a processing engine will end up maintaining both Luigi and the engine underneath.

The second limitation follows from the dependency model. Since the graph expresses completion rather than data movement, a task that needs the output of another task has to agree on a location, typically a file or a table, and rely on the atomic file system operations the README describes. That works well for batch jobs that write durable artefacts. It fits badly with fine-grained in-memory handoffs, where a framework that passes records between stages is a better match.

The third is operational. The README does not document rollback, and it does not describe what happens to in-flight tasks when the scheduler process restarts. A scheduled graph that takes days or weeks to complete, which the README explicitly says is in scope, will outlive many scheduler restarts, so that behaviour matters more here than in a short-lived job runner. Treat it as something to establish from the source and the tests in the repository rather than from the README.

Finally, the packaging metadata lists optional extras for jsonschema, prometheus and toml, and dependency groups that pull in clients for HDFS, Azure Blob Storage, boto, Docker and others. Each of those is a separate integration surface with its own release cadence, so the real maintenance burden of a Luigi deployment is usually the contrib modules, not the core scheduler.

## Luigi against Apache Airflow and Dagster

The natural comparison is Apache Airflow. Airflow also builds a DAG of tasks, but it keeps the graph in Python files that the scheduler parses to construct DAG objects, and it ships an extensive operator library and a backfill-oriented UI. Luigi's graph is built from task classes that resolve their own requirements at runtime, which is what makes dynamic requirements and date-algebra dependencies natural in Luigi. The practical difference is where the graph is defined: a static DAG object in Airflow against a recursive set of task classes in Luigi.

Dagster takes a third position, modelling pipelines around typed data assets and their materialisations rather than around task completion. That is closer to the data-passing model Luigi deliberately avoids, and it is the reason teams that need to reason about the artefacts a pipeline produces often prefer it. Luigi's answer is the file system abstraction plus the visualiser: you inspect the graph and the files, not a typed asset catalogue.

None of these differences make one of the three correct in general. They map onto different questions. If your problem is "run these two hundred jobs in the right order and tell me which ones failed", Luigi's model is direct. If your problem is "tell me which dataset is stale and why", a data-asset model answers it more directly than a task graph does.

## Maintenance, licensing and upgrade cost

The repository is not archived, and the last push was on 2026-07-18. Release tags are frequent: v3.8.1 on 2026-05-07, v3.8.0 on 2026-03-06 and v3.7.3 on 2026-02-12. The project is licensed under Apache-2.0, and pyproject.toml declares the license by file reference, with the classifier "License :: OSI Approved :: Apache Software License". Apache-2.0 is permissive and includes an explicit patent grant; it also requires that you preserve notices and state changes. That is a general description of the licence, not legal advice, and anyone embedding Luigi in a distributed product should read the LICENSE file in the repository rather than this paragraph.

Upgrade cost is dominated by two things. The first is the Python floor: requires-python is >=3.10, <3.15, so an environment pinned to an older interpreter cannot take current releases at all, and a future interpreter release will require a Luigi release before you can move. The second is the dependency surface. The core dependencies are narrow (python-dateutil, tenacity, tornado, python-daemon, typing-extensions), but the optional extras and dependency groups pull in storage and platform clients whose own version ranges move independently. A team that uses only the core scheduler has a small upgrade surface; a team using the HDFS, Azure or Docker integrations inherits each of those clients' release cycles.

## Conclusion

Adopt Luigi when your pipeline is a graph of batch tasks that already exist and you need dependency resolution, retries and a scheduler process rather than a data engine. Do not adopt it if you expect per-task data passing, dynamic cluster provisioning or a scheduler that survives a restart with full state; the documentation does not promise those. Before committing, verify that your Python version falls inside the >=3.10, <3.15 range declared in pyproject.toml, that your worker daemon strategy matches what luigid exposes, and that the contrib module for your storage system is still present in the tree.

## FAQ

### How do I install spotify/luigi?

The README gives pip install luigi for the latest stable version from PyPI, and pip install luigi[toml] if you want TOML-based configuration support. For the bleeding edge code it gives pip install git+https://github.com/spotify/luigi.git.

### Which Python versions does spotify/luigi support?

The README states that Python 3.10, 3.11, 3.12, 3.13 and 3.14 are tested, and pyproject.toml declares requires-python as >=3.10, <3.15.

### Does spotify/luigi process data or only schedule it?

It schedules. The README says Luigi is not a framework to replace Hive, Pig or Cascading, and that it helps you stitch many tasks together, where each task can be a Hive query, a Hadoop job, a Spark job, a Python snippet or a table dump.

### Is there a web interface for spotify/luigi?

Yes. The README states that the Luigi server comes with a web interface so you can search and filter among all your tasks, and the visualiser renders the dependency graph with completed tasks in green and tasks yet to run in yellow.

### What command starts the spotify/luigi scheduler?

pyproject.toml declares the console script luigid, which maps to luigi.cmdline:luigid, alongside luigi for running tasks. The README refers to the scheduler as the Luigi server.

## Sources

- [Issues](https://github.com/spotify/luigi/issues)
- [License: Apache-2.0](https://github.com/spotify/luigi/blob/master/LICENSE)
- [README](https://github.com/spotify/luigi/blob/master/README.md)
- [Releases](https://github.com/spotify/luigi/releases)
- [spotify/luigi on GitHub](https://github.com/spotify/luigi)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/spotify-luigi
