# dlt (data load tool): a Python library for loading messy sources into typed tables

> dlt is an Apache-2.0 Python library that turns APIs, SQL databases, buckets and DataFrames into typed tables in 20+ destinations. It is a library, not a platform, and that distinction decides who should use it.

**dlt-hub/dlt** — data load tool (dlt) is an open source Python library that makes data loading easy 🛠️ 

- Repository: https://github.com/dlt-hub/dlt
- Website: https://dlthub.com/docs
- Stars: 5,913 · Forks: 614
- Language: Python
- License: Apache-2.0
- Published: 2026-09-22 · Updated: 2026-09-22 · Language: en
- Canonical page: https://hysenlabs.com/projects/dlt-hub-dlt

## What dlt replaces, and who ends up writing the code

Every pipeline has a boring middle: paginate the API, guess the types, create the table, decide what happens when a record reappears with a changed field. dlt is aimed at that middle. The README frames it as a library that "automates all your tedious data loading tasks" and that can be dropped into a Colab notebook, an AWS Lambda function, an Airflow DAG, a laptop or an AI coding agent. That list is the audience statement. If your load step already lives inside Python, dlt attaches to it; if your load step lives inside a managed platform's UI, dlt has nothing to attach to.

The pyproject.toml describes the package as "an open-source python-first scalable data loading library that does not require any backend to run." That sentence is the whole positioning. There is no server to deploy, no control plane to authenticate against, and no state outside the files dlt writes. The trade-off is symmetrical: you get no hosted UI, no built-in scheduler and no catalog, and in exchange nothing runs that you did not start.

## How a dlt pipeline moves data: source, pipeline, destination

The mechanism is three objects. A source produces records, a pipeline normalizes and loads them, a destination writes them. In the README's declarative REST example the source is a dictionary describing a client, a base URL and a list of resources, each with a name and an endpoint path. From that description dlt handles requests, pagination, schema inference and typing.

The second shape is simpler and more revealing: a resource is just a generator. The README says so directly, and the example decorates a function with table_name, primary_key and write_disposition, yields two dictionaries, and passes the result to a pipeline. Because the input is a Python iterable, dlt infers column types from the values it sees and writes the table. Nothing about the destination dialect is in the source code.

That separation is why the destination is a string. The README's example swaps duckdb for snowflake, bigquery, postgres, redshift, databricks, athena, clickhouse, motherduck, filesystem targets on S3, GCS or Azure, iceberg, delta, and custom reverse-ETL destinations. Changing the string changes credentials, DDL dialect, staging and schema drift handling. The cost of that convenience is that destination-specific behaviour is mediated by dlt, so anything unusual about your target has to be expressible through dlt's configuration rather than through SQL you write yourself.

## Install dlt and run a first pipeline into DuckDB

dlt supports Python 3.10 through 3.14, with the README noting that some optional extras are not yet available for 3.14 and that support for that version is considered experimental. The base install is one command.

```bash
pip install dlt
```

Extras pull in what a specific source or destination needs. The README lists several, including duckdb for a local destination, bigquery for BigQuery, s3 for cloud filesystems, sql_database for reading any SQL database, and hub for data quality, transformations and AI features. If you use uv, the README gives the equivalent as uv add.

```bash
pip install "dlt[duckdb]"
```

With that installed, a resource is a decorated generator. This example is the README's, with the same table_name, primary_key and write_disposition arguments and the same two records.

```python
import dlt

@dlt.resource(table_name="tracks", primary_key="id", write_disposition="merge")
def tracks():
    yield {"id": 1, "title": "Yellow",       "artist": "Coldplay",   "streams": 4_200_000_000}
    yield {"id": 2, "title": "Shape of You", "artist": "Ed Sheeran", "streams": 3_900_000_000}

dlt.pipeline(
    destination="duckdb",
    dataset_name="spotify_data",
).run(
  source=tracks(),
)
```

Running that script should create a local DuckDB database, infer the columns from the yielded dictionaries, and load two rows into a table called tracks inside the spotify_data dataset. To read the result back without opening a SQL client, the README uses the pipeline's dataset accessor. In the REST example the call is pipeline.dataset().playlist_tracks.df(), which returns the loaded table as a pandas DataFrame.

## The declarative REST source and where it stops helping

The rest_api_source configuration is the most opinionated part of the library. You describe a client with a base_url and a paginator, then a list of resources. The README shows a cursor paginator with cursor_path set to next_cursor, and a resource whose endpoint path contains a placeholder, playlists/{playlist_id}/tracks.

```python
from dlt.sources.rest_api import rest_api_source

source = rest_api_source({
    "client": {
        "base_url": "https://api.spotify.com/v1",
        "paginator": {"type": "cursor", "cursor_path": "next_cursor"},
    },
    "resources": [
        {
            "name": "playlist_tracks",
            "endpoint": {"path": "playlists/{playlist_id}/tracks"},
        },
    ],
})
```

Processing steps are where the declarative approach meets its edge. The README attaches a filter step and a map step to a resource, with the map pointing at a plain Python function that flattens a track record. So the answer to an awkward API is not a configuration key but ordinary code. That is a reasonable design, and it also means the declarative layer is a convenience over Python rather than a replacement for it. If a source needs signing, retries against a nonstandard error format, or a two-phase lookup, you will be writing functions, and at that point the question is whether the declarative wrapper is still earning its place.

## Loading DataFrames, SQL tables and files without a platform

Three other source shapes are documented. The SQL database source reflects tables and types straight from the database, and the README's example is a single call with a SQLAlchemy URL.

```python
from dlt.sources.sql_database import sql_database

source = sql_database("mysql+pymysql://user:pass@host/spotify")
```

The filesystem source lists files and then parses them, composed with the pipe operator and renamed with with_name. The README's example reads tracks_*.csv from an S3 bucket through DuckDB's CSV reader.

```python
from dlt.sources.filesystem import filesystem, read_csv_duckdb

source = (
    filesystem(
        bucket_url="s3://my-bucket/spotify",
        file_glob="tracks_*.csv",
    ) | read_csv_duckdb()
).with_name("tracks")
```

DataFrames go in directly. The README shows a pandas DataFrame passed to run() with a table_name, and states that pandas, Polars and Arrow tables load directly, with Arrow-backed frames moving with zero copies. That last claim is about the Arrow path specifically, and the README does not put a number on it, so treat it as a design statement rather than a measurement.

## Where dlt is the wrong tool

The README is explicit that dlt is "a library, not a platform" and that you keep your workflow and the other tools you already use. Read that as a boundary rather than a boast. There is no scheduler, so a pipeline runs when you call run(), which means cron, Airflow, Lambda or a notebook is doing the triggering. There is no UI for browsing loaded data, so you inspect the destination with its own client or through the dataset accessor. There is no access control layer, because there is no service to control access to.

There is also a version boundary worth reading carefully. The package declares requires-python as >=3.10, <3.15, and the README says optional extras are not yet available for 3.14 and calls support for that version experimental. If your runtime is pinned to 3.14 and your pipeline depends on an extra, that combination is the one to test before you plan around it.

A subtler failure mode is schema drift. dlt infers and evolves schemas, and the README lists schema drift among the things it handles when you change the destination string. Inference that adapts is helpful until a source starts emitting a field with inconsistent types, at which point the behaviour you get is whatever dlt's normalization decides, not what you would have written by hand. The README does not document a rollback procedure for a bad load, so plan for that gap yourself.

## dlt against writing your own loader or using a hosted ELT service

The real alternative splits in two. The first is a hand-written loader: requests plus a database driver plus your own DDL. That gives you total control over retries, batching and error handling, and it gives you every one of those problems to solve again for the next source. dlt's bet is that schema inference, pagination and dialect-specific DDL are the same problem each time, and its declarative REST source and its 20+ destination strings are that bet made concrete. If you have exactly one source and one destination that will never change, a script is genuinely simpler.

The second alternative is a hosted ELT service, where connectors are configured in a web UI and the vendor runs the sync. The difference in approach is where the code lives. With dlt the pipeline is a Python file in your repository, reviewable in a pull request and runnable on your laptop; with a hosted service the pipeline is configuration in someone else's system. dlt's README also points at a large catalogue of prebuilt sources, so the connector argument is not as one-sided as it first appears. What you give up by staying in Python is the operational layer, and that is the honest cost.

## Licence, maintenance and the cost of upgrading

dlt is licensed Apache-2.0, stated in both the README's badge set and the pyproject.toml license field, and the package classifiers include "License :: OSI Approved :: Apache Software License". That is a permissive licence, which matters if you intend to embed the library in a product rather than only run it internally. It is not legal advice; if your organisation has a policy on dependency licences, that policy decides.

On maintenance, the repository is not archived and the last push was on 2026-09-21. The most recent release listed is 1.30.0 from 2026-08-11, preceded by 1.29.1 on 2026-07-24 and 1.29.0 on 2026-07-13. That is a steady release cadence across the two months before the last push, and the version in pyproject.toml matches 1.30.0.

The upgrade cost is the usual one for a library that owns your schema. Release notes are the place to check for normalization or destination changes before bumping, and the repository carries a compiled_packages.txt file alongside uv.lock and a Makefile with many test targets, which tells you the maintainers test against a pinned dependency set rather than an open range. If you pin dlt in your own lockfile, a version bump is a code change with a schema consequence, not a patch you apply blind.

## Conclusion

Adopt dlt if you already write Python and want the load step to live in your own code rather than in a vendor's scheduler, and start with the DuckDB destination because it needs no credentials. Do not adopt it if you want a hosted UI, a catalog or a scheduler; the README describes a library with no backend, so orchestration stays your problem. Before committing, verify the Python version you run (the package declares >=3.10, <3.15), the extras you need for your source and destination, and whether the write_disposition and primary_key settings you plan to use behave the way your target expects.

## FAQ

### What does dlt stand for in the dlt data load tool?

The README expands it as "data load tool (dlt)", and the package name on PyPI is dlt. The description in pyproject.toml calls it an open-source python-first scalable data loading library that does not require any backend to run.

### What is the dlt data load tool used for?

It loads data from APIs, SQL databases, cloud buckets and DataFrames into typed tables in destinations such as DuckDB, BigQuery, Snowflake, Postgres and Iceberg. The README describes it as a library rather than a platform, so it runs inside your existing Python code.

### How do I install dlt and run a first pipeline?

Install the base package with pip install dlt, or add an extra such as dlt[duckdb] for a local destination. Then point a pipeline at duckdb with a dataset_name and call run() on a source, which the README shows both as a decorated generator and as a declarative REST API description.

### Which Python versions does dlt support?

The README states support for Python 3.10 through 3.14, and pyproject.toml declares requires-python as >=3.10, <3.15. The README adds that some optional extras are not yet available for 3.14 and that support for that version is considered experimental.

### Can I use dlt without a hosted platform or backend?

Yes. The pyproject.toml description says dlt does not require any backend to run, and the README says it is a library, not a platform, that you install into your existing code. There is no scheduler or UI included, so triggering and browsing are left to the tools you already use.

## Sources

- [dlt-hub/dlt on GitHub](https://github.com/dlt-hub/dlt)
- [License: Apache-2.0](https://github.com/dlt-hub/dlt/blob/devel/LICENSE)
- [Project website](https://dlthub.com/docs)
- [README](https://github.com/dlt-hub/dlt/blob/devel/README.md)
- [Releases](https://github.com/dlt-hub/dlt/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/dlt-hub-dlt
