# Apache Hamilton: Python DAGs Built From Ordinary Functions

> Apache Hamilton (incubating) turns plain Python functions into a directed acyclic graph of data transformations. It is a lightweight library, not an orchestrator, and its value and its limits both follow from that choice.

**apache/hamilton** — Apache Hamilton helps data scientists and engineers define testable, modular, self-documenting dataflows, that encode lineage/tracing and metadata. Runs and scales everywhere python does.

- Repository: https://github.com/apache/hamilton
- Website: https://hamilton.apache.org/
- Stars: 2,601 · Forks: 215
- Language: Jupyter Notebook
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/apache-hamilton

## The problem Apache Hamilton solves for data teams

Most Python data work starts as a script or a notebook and ends as something nobody wants to touch. The logic is correct, but the dependency order lives in the author's head, the intermediate steps are named inconsistently, and moving the code into a scheduled job means rewriting the plumbing. Apache Hamilton targets that specific gap. The README frames it as the distance between proof-of-concept and production, and between data science, engineering and ops as separate functions.

The intended user is a data scientist or data engineer who already writes Python and wants structure without adopting a scheduler. The library is described as lightweight and portable: the same DAG definition runs in a script, a notebook, an Airflow pipeline or a FastAPI server. That portability claim is the product. If your transformations only ever run in one place, the payoff is smaller.

A second audience is teams that need lineage and metadata without bolting on a separate catalog. The repository description states that Hamilton encodes lineage and tracing, and the UI is presented as a way to visualize, catalog and monitor execution. That is a different pitch from pure orchestration.

## How the DAG is derived from function parameters

The mechanism is deliberately small. You write normal Python functions, and each function's parameters name the things it depends on. A function C that takes A and B as arguments has just declared two edges. Hamilton reads the module, matches parameter names to other functions or to inputs you supply, and assembles the graph. There is no separate DAG specification file and no decorator required to register a node.

The README's worked example is three functions, A, B and C, where B and C refer to A through their parameters. The documentation states that Hamilton loads that definition and builds the DAG automatically, and that the same one-line driver call works whether the graph has ten nodes or a thousand.

Two design consequences follow. First, the graph is a byproduct of ordinary Python, so unit testing a node means calling a function. Second, the definition and the execution are separated: the functions describe what to compute, and a driver decides when and where to run it. The README calls this out as letting data scientists focus on problems while engineers manage production pipelines.

The repository layout supports the modularity claim. The examples directory contains separate folders for airflow, dagster, dask, dbt, ibis, caching, data quality, LLM workflows and more, which indicates that integration is expected to happen around the core library rather than inside it.

## Installing Apache Hamilton and running a first DAG

Installation is a single pip command. The README recommends the visualization extra so you can render the graph, and notes that Graphviz must be installed on your system separately for that rendering to work. The package name on PyPI is apache-hamilton, and pyproject.toml declares support for Python >=3.10.1, <4.

```bash
pip install "apache-hamilton[visualization]"
```

If you also want the Hamilton UI, the README gives a second extra pair. Note that these are two separate extras, both named in the same bracket.

```bash
pip install "apache-hamilton[ui,sdk]"
```

With the package installed, a first DAG is a Python file containing functions whose parameters name their dependencies. The README's illustration uses functions A, B and C, where B and C take A as a parameter. Save those functions in a module, then pass the module to a driver, which resolves the graph and executes it. The README does not print the driver call in the excerpt available here, so check the driver concept page in the documentation for the exact import path and argument names before you copy anything.

What you should see after a successful run is the computed result for whichever node you requested, plus, if visualization is installed and Graphviz is present, a rendered graph showing A feeding both B and C. If the render fails, the missing piece is almost always the system-level Graphviz binary rather than the Python package.

For a zero-install look at the same ideas, the README points to tryhamilton.dev, a browser playground.

## Where the function-modifier approach becomes a constraint

The README makes a strong claim that function modifiers keep large DAGs DRY, and that other frameworks end up with redundant code or bloated functions. That is a real benefit, and it is also where the sharpest limitation sits. Modifiers are a layer of indirection: a reader looking at a function may not see the full behaviour until they know which modifiers are applied and under what conditions. The reference for each modifier lives on the documentation site, not in the README, so the learning curve is real for anyone inheriting a pipeline built with them.

The more fundamental boundary is stated by the project itself. The README says Hamilton is great for DAGs, but that if you need loops or conditional logic to build an LLM agent or a simulation, you should look at the sister library Burr. That is an honest scoping statement and it should be read literally. A workflow whose control flow is the point, rather than whose data dependencies are the point, is the wrong fit. Retry loops, iterative agent steps and stateful simulations do not map cleanly onto a static graph.

There is also an operational boundary. Hamilton is a library, not a scheduler. Portability across Airflow, Dagster and similar tools is an advantage, but it means the library will not decide when your pipeline runs, how failures are retried or how backfills are handled. Those remain the responsibility of whatever executes the driver call. Teams expecting a batteries-included platform will be disappointed.

Finally, the project is in incubation at the Apache Software Foundation. The README's own disclaimer states that incubation indicates the project has yet to be fully endorsed by the ASF, and that the process continues until infrastructure, communications and decision making stabilize. That is a governance fact worth weighing, not a defect in the code.

## Apache Hamilton compared with a task orchestrator

The natural comparison is with an orchestrator such as Airflow, and the difference is one of layer, not of quality. Airflow schedules tasks, manages retries, tracks run history and coordinates work across systems. Apache Hamilton defines the data transformations inside a unit of work and derives their order from function signatures. The repository includes an examples/airflow directory, which is consistent with the two being used together rather than as substitutes.

A second comparison is with dataframe pipeline libraries that chain method calls. Those express a pipeline as a linear sequence of operations, which reads well until you need to reuse an intermediate result in two places or inspect a single step. Hamilton's graph is a graph: a node can feed several downstream nodes, and any node can be requested on its own. The trade-off is that you write more small functions instead of one fluent chain, and the payoff only appears once the pipeline has enough steps that reuse and testing matter.

Against a full feature store or catalog product, Hamilton's lineage comes from the code rather than from a separate registration step, which the repository description highlights as a core property. The counterpoint is scope: a dedicated catalog will do more with that lineage than the library itself does, and the README's UI is presented as the place where visualization, cataloging and monitoring happen.

## Licence, releases and the cost of upgrading

Apache Hamilton is licensed under Apache-2.0, and pyproject.toml declares license-files covering LICENSE, NOTICE and DISCLAIMER. For most teams that is a permissive licence with a patent grant, but the NOTICE and DISCLAIMER files exist for a reason and should be carried along if you redistribute the package. This is a description of what the repository states, not legal advice; check with your own counsel if redistribution matters to you.

The release history shows a steady cadence rather than a frozen artifact. Version 1.89.0-incubating was released on 2025-10-11, and apache-hamilton-v1.90.0-incubating-RC0 appeared on 2026-04-25. The last push to the repository was on 2026-09-10. Note the versioning convention: the incubating suffix is part of the release name, and pyproject.toml keeps its own version field in sync with hamilton/version.py, with a comment in the file saying so.

Upgrade cost is mostly the ordinary Python kind: pin the version, read the release notes, and run your node-level unit tests, which is the main advantage of a graph made of plain functions. The one thing to watch is the version pin in pyproject.toml, which sets requires-python to >=3.10.1, <4. If you run an older interpreter, the package will not install at all, and the README's badge listing starts at Python 3.10.

## Conclusion

Adopt Apache Hamilton if your team already writes Python functions that pass dataframes or dicts around and you want the dependency graph to be visible, testable and portable across a notebook, a script, Airflow or a FastAPI service. Skip it if your workflow is dominated by loops, conditional branching or agent-style control flow, which is why the project points those users at its sister library Burr. Before committing, verify two things yourself: that your Python version satisfies the requires-python range of >=3.10.1, <4 declared in pyproject.toml, and that the function modifiers you plan to lean on are documented for the release you install, since the README advertises them broadly while the per-modifier reference lives in the documentation site.

## FAQ

### How do I install Apache Hamilton?

The README gives pip install "apache-hamilton[visualization]" for the core library plus graph rendering, and pip install "apache-hamilton[ui,sdk]" if you also want the Hamilton UI. Graphviz must be installed on your system separately for the visualization extra to render graphs.

### What Python versions does Apache Hamilton support?

pyproject.toml declares requires-python as >=3.10.1, <4, and the README's badge lists Python 3.10 through 3.14. The README text elsewhere says Python 3.8+, which conflicts with the packaging metadata, so trust the pyproject.toml range when you pin your environment.

### Is Apache Hamilton a replacement for Airflow?

No. Apache Hamilton is a library that derives a data transformation DAG from Python functions and runs wherever Python runs, while Airflow schedules and coordinates work. The repository ships an examples/airflow directory, which indicates the two are meant to be combined rather than substituted.

### What is Apache Hamilton's licence?

It is licensed under Apache-2.0, and pyproject.toml declares license-files covering LICENSE, NOTICE and DISCLAIMER. The README also carries an incubation disclaimer stating that the project has yet to be fully endorsed by the Apache Software Foundation.

## Sources

- [apache/hamilton on GitHub](https://github.com/apache/hamilton)
- [License: Apache-2.0](https://github.com/apache/hamilton/blob/main/LICENSE)
- [Project website](https://hamilton.apache.org/)
- [README](https://github.com/apache/hamilton/blob/main/README.md)
- [Releases](https://github.com/apache/hamilton/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/apache-hamilton
