Dagster: asset-based Python orchestration, and when Airflow is the better fit
An orchestration platform for the development, production, and observation of data assets.
At a glance
- What is it?
- Dagster defines pipelines as Python functions that return data assets, then tracks lineage and freshness around them. It suits teams that want testable, typed data definitions in Python; it is heavier than a cron job and a different mental model from Airflow's task DAGs.
- Who is it for?
- Adopt Dagster when your pipelines are Python functions that produce tables, models or reports and you want unit-testable definitions plus lineage in one place; the README's own quick start is enough to evaluate it in an afternoon. Skip it if your work is mostly shell steps and cron-style schedules with no data artifacts to track, or if your team has no Python.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 5 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 26, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Dagster solves, and who ends up using it
Most schedulers answer one question: did the job run? Dagster answers a different one: is this table, model or report up to date, and what produced it? The README frames the project as a cloud-native data pipeline orchestrator for the whole development lifecycle, with integrated lineage and observability, and says it is designed for developing and maintaining data assets such as tables, data sets, machine learning models, and reports. That phrasing is the whole pitch. The unit you define is the artifact, not the step.
The audience follows from that. If you write pandas or scikit-learn code and want the pipeline definition to live next to it in the same repository, in the same language, with the same test runner, Dagster is aimed at you. The README's own example is exactly this shape: a function reads a table from a URL with pd.read_html, a second fits a LinearRegression on a dummy-encoded continent column, a third groups the frame and attaches the model coefficients. Three functions, three assets, one dependency graph that the tool can render and reason about.
It is a worse fit for teams whose pipelines are mostly shell invocations, vendor CLI calls, or long-running Spark jobs whose output nobody catalogues. You can express those in Dagster, but you pay the modeling cost without collecting the lineage benefit.
The mechanism: Python functions become assets, dependencies become a graph
The programming model is declarative and the runtime is inferred. You decorate a plain Python function with @dg.asset. Its parameters are the assets it depends on, and its return value is the asset it produces. In the README example, continent_change_model takes country_populations: pd.DataFrame, so Dagster knows the model asset is downstream of the populations asset without any explicit edge declaration. continent_stats takes two parameters, so it sits downstream of both.
That inference is the design decision worth noticing. Airflow makes you write the dependency edges and pass data through an external store such as XCom or a file path; Dagster makes the function signature the edge list and lets values flow directly between in-process steps. The trade-off is real: direct passing is pleasant for pandas-sized data and awkward once an intermediate frame is too large to hold in memory between steps, at which point you are back to writing to storage and reading back anyway.
Around the graph sits the metadata layer. The README describes a unified control plane with built-in observability, diagnostics, cataloging, and lineage, and a web UI that renders the asset graph (the README shows an example lineage screenshot). The repository layout backs this up: python_modules/ holds the Python packages, js_modules/ holds the frontend, and helm/ plus examples/deploy_docker, examples/deploy_ecs and examples/deploy_k8s cover deployment shapes. The test configuration in pyproject.toml marks suites as sqlite_instance or postgres_instance, which tells you the instance storage is a swappable component rather than a fixed choice.
Installing Dagster and running a first asset
The README's quick start uses uv and installs three packages: the core library, the web server, and a CLI package. The README states Dagster is available on PyPI and officially supports Python 3.9 through Python 3.14, so check your interpreter before anything else.
uv add dagster dagster-webserver dagster-dg-cliAfter that resolves, you have the library and the webserver entry point available in the environment. The README points new users at the docs and the hands-on tutorial at docs.dagster.io/etl-pipeline-tutorial rather than walking through a project scaffold itself, so the tutorial is where the first-run instructions actually live.
A minimal asset file is short. This mirrors the README example, trimmed to a single asset so the first run is quick:
import dagster as dg
import pandas as pd
@dg.asset
def country_populations() -> pd.DataFrame:
df = pd.read_html("https://tinyurl.com/mry64ebh")[0]
df.columns = ["country", "pop2022", "pop2023", "change", "continent", "region"]
df["change"] = df["change"].str.rstrip("%").astype("float")
return dfWhat you should see is an asset named country_populations in the graph, with its materialization recorded when you run it. The README does not document the exact CLI invocation for launching the webserver against a specific file, so follow the tutorial for that step rather than guessing at flags. The README also does not document rollback of a materialization, so treat re-running as the recovery path until you confirm otherwise in the docs.
Where Dagster gets in the way
The asset model is opinionated, and the cost shows up in two places.
First, the model wants a Python function per asset. A pipeline that is genuinely a sequence of commands, or that shells out to a proprietary binary, has to be wrapped in Python to participate. The wrapper adds indirection without adding lineage, and reviewers will ask what the asset actually represents.
Second, the operational surface is bigger than a scheduler's. There is a webserver, an instance with a storage backend, code locations, and a deployment story. The repository ships helm/ charts and three deployment examples under examples/, which is convenient if you are on Kubernetes, ECS or Docker and less convenient if you are not. The pyproject.toml markers show the project tests against both SQLite and Postgres instances; running SQLite locally is a reasonable development choice, but the project's own test matrix implies Postgres is the configuration it takes seriously at scale.
A third limit is documentation shape rather than capability. The README is a landing page: it shows one example, links to docs, and lists features. It does not cover upgrade mechanics, and MIGRATION.md exists at the repository root without the README explaining when you would need it. If you are planning a version bump across a minor release, that file is the thing to read, not the README.
Dagster against Airflow, and against Prefect
The comparison people actually search for is Dagster versus Airflow, and the difference is in what the graph is made of. Airflow's DAG is a graph of tasks; the task is the first-class object and the data it moves is a side effect you manage yourself. Dagster's graph is a graph of assets; the artifact is first-class and the computation is attached to it. That single inversion changes how you name things (tables and models rather than extract and load), how you test (call the function and assert on the returned frame), and what the UI shows you by default (freshness and lineage rather than task durations).
Prefect is the closer neighbor in spirit: both are Python-first and both let you decorate ordinary functions. The distinction the Dagster README draws is the emphasis on a declarative programming model with a metadata and cataloging layer around it, whereas a flow-run-centric tool tends to center the execution record. If your team's problem is scheduling and retries, that difference will feel like overhead. If your team's problem is not knowing which dashboard broke because an upstream table is stale, the asset graph is the feature you are buying.
One caveat on the Airflow comparison: the repository contains examples/airlift-migration-tutorial and examples/airlift-mwaa-example, which suggests a documented migration path from Airflow exists. The README does not describe it, so read those example directories before assuming the port is mechanical.
Licence, release cadence and what an upgrade costs
Dagster is Apache-2.0 licensed, per the README and the LICENSE file at the repository root. Apache-2.0 permits commercial use and modification and includes an explicit patent grant; it also requires that you preserve notices and state changes. That is a permissive arrangement, and it means the open source project and the hosted offering are separate decisions. The README links to dagster.io and to a community page, and search results distinguish Dagster OSS from Dagster Plus and Dagster Cloud, but the README itself does not describe what the hosted tiers add, so pricing and feature boundaries are not something this repository answers. Nothing here is legal advice; if you are redistributing a modified Dagster, read the licence text.
The maintenance signal is straightforward. The repository is not archived, and the last push was on 2026-09-18. Releases are frequent and versioned in pairs: 1.13.23 (core) / 0.29.23 (libraries) on 2026-09-16, 1.13.22 on 2026-09-11, 1.13.21 on 2026-09-03. The core and library version numbers move together, so an upgrade plan has to account for both. The justfile at the repository root shows the project's own CI runs ruff check, ruff format --check, prettier and a ty type check via scripts/run-ty.py, which is the bar a contribution or a fork is held to. For an adopter, the practical cost is that a fast cadence means you should pin versions and read CHANGES.md before moving, and that MIGRATION.md is the file to consult when a change is not additive.
Editorial conclusion
Adopt Dagster when your pipelines are Python functions that produce tables, models or reports and you want unit-testable definitions plus lineage in one place; the README's own quick start is enough to evaluate it in an afternoon. Skip it if your work is mostly shell steps and cron-style schedules with no data artifacts to track, or if your team has no Python. Before committing, verify three things: that your Python version is within the supported 3.9 to 3.14 range, that the storage backend you intend to run (the repository's test configuration distinguishes SQLite and Postgres instances) matches your production expectations, and that the integrations you need exist on the integrations page rather than being assumed. The release cadence is the other number to check against your own upgrade budget: 1.13.23 shipped on 2026-09-16, two weeks after 1.13.21 on 2026-09-03.
Frequently asked questions
What is Dagster used for?
It is used to define and run data pipelines whose outputs are data assets such as tables, data sets, machine learning models, and reports. You declare those assets as Python functions and Dagster runs them at the right time and keeps them up to date.
Is Dagster an ETL tool?
The README describes it as an orchestration platform for the development, production, and observation of data assets rather than as an ETL product. You can express extract, transform and load work in it, but the unit it models is the asset and its lineage, not the ETL step.
Is Dagster Python based?
Yes. The repository's primary language is Python, assets are declared as Python functions, and the README states that Dagster officially supports Python 3.9 through Python 3.14.
How do I install Dagster?
The README's quick start installs it from PyPI with uv add dagster dagster-webserver dagster-dg-cli. After that, the README points to the documentation and the hands-on ETL pipeline tutorial for the first project.
What is the difference between Dagster and Airflow?
Airflow's graph is made of tasks; Dagster's graph is made of assets, with the Python function that produces each asset attached to it. That changes what the UI emphasizes, since Dagster tracks lineage and freshness of the artifacts rather than only task runs.
Does Dagster cost money?
The project itself is Apache-2.0 licensed, so the source is free to use and modify under that licence. The README links to dagster.io and a community page but does not describe paid tiers, so the boundary between the open source project and the hosted offerings is not covered by this repository.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/dagster-io-dagster)