Self-hosted service
airbytehq/airbyte avatar
airbytehq/airbyte

Airbyte Open Source: self-hosted ELT with 600+ connectors

Open-source data movement for ELT pipelines and AI agents — from APIs, databases & files to warehouses, lakes, and AI applications. Both self-hosted and Cloud.

22,152 stars5,368 forksPythonNOASSERTION

At a glance

What is it?
Airbyte is an open-source data movement platform for ELT pipelines. This review covers what it solves, how a sync actually works, how to install it with Docker, and where the self-hosted edition stops being the right choice.
Who is it for?
Adopt Airbyte Open Source if you need many API and database sources landing in a warehouse and you are willing to run and upgrade a self-hosted deployment yourself. Do not adopt it if you want zero operational surface, since the README points that audience at Airbyte Cloud instead.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Airbyte solves, and the engineer it is aimed at

Every data team eventually owns the same chore: pulling records out of SaaS APIs, relational databases and object storage, then writing them into a warehouse, a lake or a database on a schedule. Writing one bespoke script per source is easy. Writing two hundred of them, each with its own pagination, rate limits, schema drift and incremental cursor logic, is a maintenance program nobody asked for. Airbyte's answer is a connector catalog. The README states the project provides a catalog of 600+ connectors for APIs, databases, data warehouses, data lakes and AI applications, and that the intent is to cover the long tail of data sources while letting data engineers customize existing connectors.

The audience is therefore a team that already has a warehouse and a scheduler and needs the extraction layer to stop being bespoke. The README splits that audience in two. If the goal is moving data into warehouses, lakes or databases, it points at Airbyte Open Source (this repository) or Airbyte Cloud. If the goal is giving AI agents, LLMs or MCP clients real-time access to business data, it points at Airbyte Agents or the separate open-source Agent SDK. Those are different products with different install paths, and conflating them is the most common way to end up reading the wrong documentation.

How a sync actually runs: connectors, the CDK and the platform

The repository layout is the clearest description of the architecture. Airbyte is not a single Python service. There is airbyte-cdk, airbyte-integrations, airbyte-ci, docker-images, a Gradle build (build.gradle, settings.gradle, gradlew) and Python tooling (poe_tasks.toml, pytest.ini, ruff.toml). Connectors live under airbyte-integrations, and the CDK is what they are built against. That split matters when you evaluate the project: the connector you depend on is versioned and released separately from the platform that schedules it.

The data flow the README implies is the standard ELT shape. A source connector reads from the upstream system and emits records; a destination connector writes them. The platform handles the connection definition, the schedule and the state that makes incremental syncs possible. Airbyte's own positioning is ELT rather than ETL: the README describes moving data into warehouses, lakes and databases, and the connector catalog is organized around sources and destinations rather than transformations, so transformation is expected to happen downstream in the warehouse.

Two extension paths exist for sources the catalog does not cover. The no-code Connector Builder is a UI for defining connectors, and the low-code CDK is a configuration-based route. Both are documented on the project's own docs site rather than in this repository, which is worth knowing before you promise a stakeholder a custom connector by Friday.

Orchestration is deliberately left to other tools. The README lists Airflow, Dagster, Kestra and the Airbyte API as ways to trigger syncs, which tells you the platform expects to be a step inside a pipeline rather than the pipeline itself.

Installing Airbyte Open Source with Docker and running a first sync

The README does not inline an install command. It links to a deployment quickstart at docs.airbyte.com/quickstart/deploy-airbyte for Airbyte Open Source, and to a separate Cloud getting-started page. Docker is the deployment path the community search data keeps asking about, and the repository ships docker-images and a .dockerignore, so the containerized route is a first-class one. Because the exact compose file and image tags are maintained in the docs rather than in this README, follow the quickstart page for the current commands instead of copying a compose file from a blog post.

Once the platform is up, the working loop is the same in any deployment. You define a source, define a destination, then create a connection between them and set a schedule. The README's own entry points are the no-code Connector Builder for sources that do not exist yet and the tutorials page for end-to-end examples.

If you want to see the shape of a sync before installing anything, the project publishes a demo app at demo.airbyte.io. The README describes the screenshot in the repository as taken from Airbyte Cloud, so the UI you see in the docs and the UI you get from a self-hosted deployment are presented as the same product surface.

For teams that want the connector layer inside an application rather than as a platform, the README gives a separate install line for the Agent SDK:

bash
uv pip install airbyte-agent-sdk

That package is for embedding type-safe connectors as LLM tools, and the README states it works with pydantic-ai, LangChain, OpenAI Agents and FastMCP, with built-in retry, exception translation and output-size guardrails. It is not the ELT platform and does not replace it.

The self-hosted trade-off: you own the upgrade

Self-hosting is the reason many teams pick Airbyte, and it is also the main cost. The repository carries a Gradle build, a Python toolchain with poetry.lock and ruff.toml, a Makefile with pre-commit targets, and a dedicated airbyte-ci directory for connector CI. That is a large surface, and the release cadence visible in the release list (v1.7.0 in June 2025, v1.8.0 in August 2025, v2.0.0 in October 2025) means a self-hosted instance drifts from current within a couple of quarters unless someone owns upgrades.

The Makefile is honest about this being a contributor-oriented repository rather than a deployment artifact. Its targets install pre-commit hooks and clean git hooks:

bash
make tools.git-hooks.install

On macOS that target installs pre-commit and maven through brew; on Linux it uses pip. The presence of a version target that curls a Maven metadata endpoint for the Bulk CDK shows how much of the toolchain is Java and Gradle rather than pure Python, despite Python being the primary language.

A second limitation is licensing, and it is not a footnote. The repository carries a LICENSE and a LICENSE_SHORT, and the README's own badges point at a licenses directory under docs/project-overview/licenses and display both MIT and ELv2. The GitHub API reports the license as NOASSERTION, which is consistent with a repository that is not under a single uniform licence. The practical consequence is that you cannot assume every connector is under the same terms as the platform core. Check the licence that applies to the specific connectors you depend on before you plan to fork, modify and redistribute them.

Where Airbyte is the wrong tool

Airbyte is a batch-oriented data movement platform, and the README never claims otherwise. If your requirement is sub-second replication into an operational store, a connector-based sync platform is the wrong shape, and nothing in the repository suggests the project targets that.

It is also the wrong tool if you do not want to operate anything. The README presents Airbyte Cloud as the managed option for the same ELT job, and the screenshot in the repository is explicitly taken from Cloud. Choosing self-hosted means owning the deployment, the upgrade path and the connector versions yourself.

A third case is transformation-heavy work. Airbyte moves data; the README describes destinations as warehouses, lakes and databases and points at downstream tutorials rather than an in-platform transformation layer. If your pipeline is mostly business logic over already-loaded data, you are buying a connector catalog you will barely use.

Finally, custom connectors are not free. The no-code Connector Builder and low-code CDK lower the barrier, but the README links both to external documentation, which means the learning material is not in this repository. Budget for that before committing to a source the catalog does not cover.

Airbyte compared with Airflow

The most common comparison in the search data is Airbyte versus Airflow, and the two solve different problems. Airflow is a workflow orchestrator: it schedules and sequences tasks, and a task that extracts data is code you write. Airbyte is the extraction layer itself, with connectors that already know how to talk to a given API or database.

The README supports this reading directly, because it lists Airflow as a way to orchestrate Airbyte syncs, alongside Dagster, Kestra and the Airbyte API. That is not a competitive relationship; it is a composition. A typical stack runs Airflow or Dagster as the scheduler and calls Airbyte to do the movement.

The difference shows up in effort. With Airflow alone, every new source is a new operator or script plus its own pagination, retry and incremental-state handling. With Airbyte, a source that exists in the catalog is a configuration entry. The trade is control: an Airflow-native extraction gives you exact control over the request pattern, while a catalog connector gives you whatever the connector implements. Teams that need unusual API behavior often end up writing a custom connector rather than an Airflow task, which is a different skill and a different artifact to maintain.

Who should adopt Airbyte, and what to check first

Adopt Airbyte Open Source if you have a warehouse or lake as a destination, a scheduler already in place, and a growing list of sources you do not want to hand-code. The connector catalog is the product, and the README's claim of 600+ connectors for APIs, databases, data warehouses and data lakes is the thing you are actually buying into.

Do not adopt it if you want a managed service with no operational surface. The README routes that need to Airbyte Cloud, and the repository itself is a build-and-contribute codebase, not a deployment bundle. Do not adopt it either if your pipeline is mostly transformation, or if you need streaming latency rather than scheduled syncs.

Before you commit, verify three things. First, check the connector registry report linked from the README to confirm your specific sources and destinations are covered, and in what state. Second, read the licences directory the README's badges point to and identify which of your connectors fall under MIT and which under ELv2, since that determines what you may redistribute. Third, decide who owns upgrades, because the release cadence means a self-hosted instance will fall behind within a couple of quarters.

The project is not archived and the last push to master was on 2026-09-10, so the repository is receiving changes. That says nothing about whether the individual connector you need is healthy; the registry report is where that question gets answered.

Editorial conclusion

Adopt Airbyte Open Source if you need many API and database sources landing in a warehouse and you are willing to run and upgrade a self-hosted deployment yourself. Do not adopt it if you want zero operational surface, since the README points that audience at Airbyte Cloud instead. Verify first which of your sources are in the connector registry and which are in the MIT-licensed subset versus the ELv2 subset, because that split decides whether you can fork and redistribute what you depend on.

Frequently asked questions

Is Airbyte ETL or ELT?

Airbyte positions itself as ELT. The README describes moving data from APIs, databases and files into warehouses, lakes and databases, and its connector catalog is organized around sources and destinations rather than transformations, leaving transformation to the destination.

What is the difference between Airbyte and Airflow?

Airflow is a workflow orchestrator, while Airbyte is the data movement layer with prebuilt connectors. The README lists Airflow as one of the tools that can orchestrate Airbyte syncs, along with Dagster, Kestra and the Airbyte API.

Is Airbyte free to use?

There is an open-source edition in this repository and a managed Airbyte Cloud option. The repository is not under a single uniform licence: the README's badges point to a licences directory showing both MIT and ELv2, and the GitHub API reports the licence as NOASSERTION, so the terms depend on the component.

How do I install Airbyte Open Source?

The README does not inline install commands. It links to a deployment quickstart at docs.airbyte.com/quickstart/deploy-airbyte for Airbyte Open Source, and the repository ships docker-images, so the containerized route is supported. Follow the quickstart page for current commands and image tags.

What is Airbyte used for?

It moves data from APIs, databases and files into warehouses, lakes, databases and AI applications. The README also describes a separate Agent SDK for giving AI agents, LLMs and MCP clients access to business data.

How do I set up Airbyte?

Deployment starts from the quickstart link in the README, after which you define a source, define a destination, and create a connection between them with a schedule. Airbyte Cloud is the managed alternative for the same job.

Official sources

  1. airbytehq/airbyte on GitHub
  2. Issues
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/airbytehq-airbyte.svg)](https://hysenlabs.com/projects/airbytehq-airbyte)