# Mage OSS: a self-hosted notebook for building data pipelines

> Mage OSS is the open source, self-hosted half of Mage: a local development environment where pipelines are assembled block by block in Python, SQL or R. It installs in one command, but the README points at a commercial platform for anything beyond local work.

**mage-ai/mage-ai** — 🧙 Build, run, and manage data pipelines for integrating and transforming data.

- Repository: https://github.com/mage-ai/mage-ai
- Website: https://www.mage.ai
- Stars: 8,826 · Forks: 989
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/mage-ai-mage-ai

## The local pipeline editor Mage OSS is trying to be

Most orchestration tools assume you already have a cluster, a scheduler and a metadata database. Mage OSS starts from the opposite end: one machine, one project directory, and a browser tab. The README describes it as a self-hosted development environment for creating production-grade data pipelines, and the feature table lists modular pipelines, a notebook UI, prebuilt connectors, scheduling, visual debugging and dbt support.

The intended user is a data engineer or analyst who writes Python or SQL and wants to see intermediate data while building, rather than pushing a DAG to a remote scheduler and reading logs afterwards. The README's example use cases are concrete about this: moving data from Google Sheets to Snowflake with a Python transform, scheduling a daily SQL pipeline to clean and aggregate product data, developing dbt models in a notebook-style interface, running simple ETL or ELT jobs locally. Nothing in that list requires a cluster.

The trade-off is stated openly. The same README that documents the local tool also advertises Mage Pro for enterprise orchestration, collaboration and AI-assisted workflows, and lists multi-environment orchestration, role-based access control and real-time monitoring as Pro features. Mage OSS is positioned as the on-ramp, not the destination.

## Blocks, notebooks and the pipeline directory

The unit of work is a block. A pipeline is a sequence of blocks, and each block holds logic in Python, SQL or R. That is the whole mental model the README offers, and it maps onto the repository layout: the mage_ai package sits next to mage_integrations, which is where connector code lives, and the top level also carries templates, integrations, kube, aws and scripts directories. The dependencies in requirements.txt tell you what a block can reach for without extra installation: pandas, polars, pyarrow, dask, scikit-learn, great-expectations, sqlglot, redis and SQLAlchemy are all direct requirements, while heavier clients such as boto3, google-cloud-bigquery, duckdb, clickhouse-connect and mysql-connector-python sit behind the extras section of the same file.

Execution is step by step. The README promises logs, live previews and step-by-step execution for debugging, which is the main reason to prefer this over a job you submit and wait on. Scheduling is described as manual or cron, and the cron-converter dependency in requirements.txt is consistent with a cron expression being parsed somewhere in the scheduler rather than a bespoke trigger format. The docker-compose.yml in the repository shows the server being started directly as python mage_ai/server/server.py with --host, --port, --project and --manage-instance arguments, which is the shape of the runtime: a Tornado server (tornado is a pinned requirement) serving the UI and managing the project it was pointed at.

What the README does not describe is the state layer. There is no explanation of where pipeline runs are recorded, how concurrent runs of the same pipeline are handled, or what happens to a partially completed run. If those questions matter to you, they are answered in docs.mage.ai or not at all; the README is silent.

## Installing Mage OSS and running a first pipeline

The README gives three installation routes and calls Docker the recommended one. The image is published as mageai/mageai, so pulling it is the first step.

```bash
docker pull mageai/mageai:latest
```

If you prefer to keep it inside a Python environment, the same README offers pip and conda. Note that setup.py declares python_requires='>=3.10' even though pyproject.toml still carries a wider range for the development tooling, so the Python 3.10 floor is the one that applies to the published package.

```bash
pip install mage-ai
```

The conda route uses the conda-forge channel.

```bash
conda install -c conda-forge mage-ai
```

Installing the package places a console script named mage, declared in setup.py as mage=mage_ai.cli.main:app. The README does not list its flags and points to docs.mage.ai for the full setup guide, so the exact start command is not documented in the repository's README.

After that, the workflow the README describes is browser-based: open the local server, create a pipeline, add a block, write Python or SQL in it, and run the pipeline manually. You should see per-block logs and a data preview for the block output. To put it on a schedule, the README says cron is supported; the exact configuration lives in the docs rather than the README.

For a first real use, the README's own example is the shortest path: a Python block that reads from an API or spreadsheet, followed by a SQL block that cleans and aggregates, with the pipeline triggered on a daily cron. Keeping the first pipeline to two blocks makes the preview feature useful rather than noisy.

## Where Mage OSS stops and Mage Pro begins

The clearest limitation is not technical, it is editorial. The README is a sales page for two products at once. It opens with the local tool, then repeatedly returns to Mage Pro, and the Pro bullet list includes AI-assisted development and debugging, multi-environment orchestration, role-based access control, real-time monitoring and alerts, CI/CD and version control, and deployment as fully managed, hybrid or on-premises.

Read that list as a map of what the open source project does not claim to provide. If your requirement is a permission model with roles, or an alert when a scheduled run fails, the README does not offer it in Mage OSS; it offers it in Pro. That is a legitimate open-core split, but it means an evaluation based on the README alone will overestimate the open source tool.

The second limitation is scope of documentation. The README is short by design and defers almost everything to docs.mage.ai. It does not document rollback, does not document how to recover a failed run, and does not document the persistence model for the project store. The repository contains alembic as a dependency and a migrations directory under mage_ai/orchestration/db, which implies a relational metadata store with schema migrations, but the README never says so. For a tool you intend to schedule daily, the upgrade path for that store is the thing to read about before you rely on it.

A third constraint is the dependency surface. requirements.txt pins aggressively (Werkzeug 3.0.3, tornado 6.4.2, pandas >=2.1,<3, protobuf >=6.0,<7, and many more exact pins). That is a deliberate choice for reproducibility, and it also means Mage OSS shares a Python environment poorly with anything that wants different versions of those libraries. The Docker image sidesteps the problem entirely, which is presumably why the README recommends it first.

## Mage OSS compared with Airflow and Dagster

The obvious comparison is Apache Airflow. Airflow's core abstraction is a DAG defined in Python code, scheduled by a central scheduler against a metadata database, with the UI as an observability layer over runs. Mage OSS inverts the emphasis: the UI is where you build, and the pipeline is assembled from blocks rather than declared as a graph in a file. For a team that already has Airflow running and a convention for DAG files, Mage OSS offers little, because the thing it improves is the authoring loop, not the scheduling semantics.

Dagster is the closer comparison, since it also treats local development as a first-class concern and has a UI for inspecting runs. The difference stated by the README is the notebook-style interface and the language mix: Mage blocks can be Python, SQL or R inside the same pipeline, and dbt models can be developed in the same interface. Dagster's model is software-defined assets declared in Python. Which is better depends on whether your team's centre of gravity is SQL or Python.

A third option worth naming is to write nothing at all and use cron plus a script. Mage OSS earns its install when you need the preview, the logs and the block-level retry; for a single daily query, it does not.

## Version cadence, the Apache-2.0 licence and what an upgrade costs

Releases are not on a fixed rhythm. The most recent release listed for the project is 0.9.79, published on 2026-01-21. Before that, 0.9.78 landed on 2025-09-18 and 0.9.77 on 2025-09-03. The last push to the default branch was on 2026-08-13, so the repository is being worked on between releases, but the gap between 0.9.78 and 0.9.79 is about four months. Plan upgrades around releases, not around the commit stream.

There is a versioning wrinkle worth knowing. setup.py sets version='0.9.79' and carries a comment saying that changing it requires changing VERSION in mage_ai/server/constants.py as well. Meanwhile pyproject.toml, which describes itself as the DevEx poetry configuration, still reads version = "0.8.75-dev". If you build from the repository rather than installing a release, the version you see in the UI and the version in the packaging metadata can disagree. That is a real source of confusion when filing or reading bug reports.

The licence is Apache-2.0, declared in the LICENSE file and in the setup.py classifier. Apache-2.0 permits commercial use and modification and includes an explicit patent grant, which matters if you are embedding this in a product. It also means the open-core boundary is a business decision rather than a licence restriction: nothing in Apache-2.0 stops you from adding your own role-based access layer, though the README does not describe any extension point for doing so. This is not legal advice; read the LICENSE file and your own counsel's view before shipping derivatives.

Upgrade cost is dominated by the pinned requirements. Moving between releases can move exact pins such as tornado, Werkzeug or pandas, and any of those can break a block that depended on the old behaviour. The Docker image is the lower-risk upgrade path because the environment moves with the image.

## Conclusion

Adopt Mage OSS if you want a local, Apache-2.0 workspace where a Python or SQL pipeline can be written block by block and run on a cron schedule without a cloud account. Do not adopt it expecting the enterprise feature list on the same page: multi-environment orchestration, role-based access control and real-time alerting are described as Mage Pro, a separate product. Before committing, verify two things in docs.mage.ai: whether your target database has a prebuilt connector, and how the project store is persisted, since the README does not document rollback or state recovery.

## FAQ

### Is mage ai open source?

The repository is licensed Apache-2.0, declared in the LICENSE file and in the setup.py classifier, and the README calls this project Mage OSS. The same README describes Mage Pro as a separate commercial platform with enterprise features.

### What is mage.ai used for?

The README lists building pipelines locally with Python, SQL or R, running jobs manually or on a cron schedule, connecting to databases, APIs and cloud storage, and developing dbt models in a notebook-style interface. Its stated example use cases include moving data from Google Sheets to Snowflake and scheduling a daily SQL pipeline.

### how to use mage ai

Install it with Docker, pip or conda, start a project, then build a pipeline block by block in the browser UI and run it manually or on a schedule. The README defers the detailed steps to docs.mage.ai.

### is mage ai free

Mage OSS is open source under Apache-2.0 and the README says no cloud account is required, so the local tool carries no licence fee. The README points to Mage Pro for advanced tooling and enterprise features, which is a separate product.

### Does Mage AI have an app?

The README describes a self-hosted development environment that you install with Docker, pip or conda and then open in a browser. It does not describe a mobile or desktop application.

## Sources

- [License: Apache-2.0](https://github.com/mage-ai/mage-ai/blob/master/LICENSE)
- [mage-ai/mage-ai on GitHub](https://github.com/mage-ai/mage-ai)
- [Project website](https://www.mage.ai)
- [README](https://github.com/mage-ai/mage-ai/blob/master/README.md)
- [Releases](https://github.com/mage-ai/mage-ai/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/mage-ai-mage-ai
