Amphi's real deliverable is Python code, not a pipeline definition
visual data prep powered by python
At a glance
- What is it?
- The visual pipeline is a front end that emits ordinary Python using pandas and an in-process query engine, and the readme makes that explicit. That is the difference between this and a pipeline tool: when you are done, you have a script, and the licence terms are the part to read carefully.
- Who is it for?
- Adopt Amphi if you build data pipelines for people who do not write Python, since the generated code is ordinary Python with well-known libraries and a self-hosted server you control, which is a much better position than a proprietary pipeline format.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 44 days ago.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 6, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The pipeline is a code generator, and that is the whole bet
The feature list has three items under data preparation and the second is the one that determines whether this tool is useful to you. A visual interface for low-code pipeline development, Python code generation that produces native Python using libraries such as pandas and an in-process query engine and that you can run anywhere, and a claim about privacy through self-hosting. Read the second item carefully: the deliverable is a script. When you draw a pipeline, you get Python, and the Python uses libraries you already know and can install anywhere. That is a categorically different design from a pipeline tool that stores your transformations in its own format and needs its own runtime to execute them. With the first design, the visual layer is disposable. If the product is discontinued, you still have working code. With the second, the visualisation is the program and losing it loses the work. The third feature follows from the second: self-hosting is possible on your laptop or in your own cloud because what runs is your machine, not a service. The note above the list is a positioning statement worth taking seriously, that this is focused on data transformation for preparation, reporting and lightweight extract-transform-load work, and that it aims to be quick to learn and usable with language models. Lightweight is the operative word, and it also bounds the ambition.
The component extension point is a code generator with a form, which is a good design
The extensibility section shows how to add a component, and the code in the readme is the most substantial thing in it. A component is a class that extends a base component, and it supplies a description, a default configuration, a form definition describing the interface fields, an icon, and a technical identifier, all passed to the base constructor in a fixed order. Then two methods matter. One returns the import statements the generated code needs, and the other returns the code itself as a template string that interpolates the configuration values and the output variable name. So a component is not a plugin that hooks into a runtime; it is a code emitter with a user interface attached. That has three consequences worth naming. A component written in one language can emit code in another, which is why the interface can be a component file in a modern web syntax while the output is Python. The form and the code are declared together, so a mismatch between what the interface collects and what the code uses is a bug you write rather than a configuration error you discover at run time. And because the base class is exposed on a global object, the extension mechanism is documented enough that someone can write a component without reading the source, which is what the readme invites you to do. The example is deliberately trivial, a date picker that makes a one-row data frame, which is the right choice for documentation.
A local server with a bind address flag, which is the deployment story
The usage section is short and describes a command with three parameters:
amphi start -w /your/workspace/pathamphi start -w /your/workspace/path -i 0.0.0.0 -p 8888You run a start subcommand, and you can pass a workspace path telling it where your files are and where it should create pipelines, an address to bind to, and a port. The two documented invocations make the security model explicit. On your own machine you pass only a workspace, which means it binds to a loopback address and nothing else can reach it. On a server you pass an address that binds to all interfaces, and the readme tells you to do that so you can reach it over the internet. That is a two-line security note in a readme that otherwise does not discuss security, and it is exactly the kind of thing a self-hosted tool with a web interface needs to say: binding to all interfaces with no authentication in front of it means anyone who finds the address can read your workspace. The readme does not mention authentication, a login, or a reverse proxy, and the examples pick a port that suggests a notebook server convention. If you deploy the server form, put it behind something that authenticates. If you use the local form, you are fine. There is also a hosted demo linked from the readme, which is the sensible first step.
Two packages, a scheduler, and a requirements file that installs itself
The repository layout describes three projects and one of them is a clue. There is the main package directory, a JupyterLab extension directory, a scheduler directory, and a directory of test assets. Two of those ship as separate installable packages, which the readme's table makes explicit: the standalone application installs under one name and the notebook extension under another, with an upgrade command for each. That is a sensible split, since a data scientist already inside a notebook environment and a team wanting a shared server have different deployment needs. The scheduler directory is the interesting one. Its presence tells you scheduled execution is a real feature rather than a roadmap item, and its absence from the readme tells you the readme is not a complete feature list. The requirements file at the top of the repository is worth reading as a maintenance signal: it pins a notebook environment version, pins the extension to a specific version, and then includes a bare dot, meaning it installs the repository itself. So the development environment is reproducible, and the extension is pinned to a published version rather than the local build, which is a slightly unusual arrangement that probably exists to test the extension against its published artefact. The other top-level files are the documents a well-run project has: a changelog, a building guide, a contributing guide, a releasing guide, a document about using git and the forge, and an agent instruction file.
Telemetry, a licence badge, and a hosted demo
Two paragraphs near the end of the readme cover the things that affect whether you can use this. The first is telemetry. The readme states that anonymous telemetry data is collected to understand users and their use cases, and that you can opt out in the settings and disable collection entirely. For a tool whose main selling point is that your data stays on your machine, a phone-home is a legitimate thing to be uneasy about, and the readme's position is that it is anonymous and switchable. The honest reading is that the opt-out is what makes it acceptable: the product is a server you run yourself, so it can reach out, and the setting is the mechanism that lets an operator who objects stop it. Check that the setting is where the readme says it is. The second is the licence, and the badge is the important artifact. It is not the plain Apache badge. It is a badge naming the Elastic licence variant of the Apache terms, and that licence permits use as a service but adds a condition restricting offering it as a hosted or managed product. So the tool is source available rather than genuinely open source, which is a meaningful distinction for anyone who was planning to run it as an internal platform for a business. The readme's own positioning as something you self-host to keep control of your data sits slightly awkwardly with a licence that limits the commercial hosting of it, and that tension is worth resolving before you adopt it organisationally. A hosted demo is linked for evaluation.
Editorial conclusion
Adopt Amphi if you build data pipelines for people who do not write Python, since the generated code is ordinary Python with well-known libraries and a self-hosted server you control, which is a much better position than a proprietary pipeline format. Do not adopt it if you need scheduled production pipelines, because the readme's stated focus is data transformation for preparation, reporting and lightweight work, and a scheduler is a separate package in the repository whose scope the readme does not describe. Four things to verify. The licence, which the badge identifies as a source-available variant of the Apache terms with an additional condition, and that is a materially different thing from the Apache licence itself. Where your data goes, since the readme states anonymous telemetry is collected and that you can opt out in settings, so check the setting exists in your version. What the generated code looks like for your own pipelines, because the value proposition is that it is readable and portable and the readme only shows a toy example. And whether the JupyterLab extension or the standalone server fits your team, because they are different packages with different surfaces. The last push was on 2026-08-24 and no releases are published under a version scheme you can pin to.
Frequently asked questions
What does Amphi ETL produce when I build a pipeline?
Ordinary Python code using libraries such as pandas and an in-process query engine, which the readme says you can run anywhere. The visual layer is a code generator rather than a proprietary pipeline format, so the output is portable and the tool is disposable.
How do I install and run Amphi?
It is available as a standalone application or a JupyterLab extension, each with its own install and upgrade command. To start the standalone version you run a start subcommand, optionally with a workspace path, a bind address and a port, and the readme distinguishes a local invocation from a server one that binds to all interfaces.
How do I add a custom component to Amphi?
Write a component file that extends the exposed base component class, supplying a description, default configuration, a form definition, an icon and an identifier, then implement a method returning the import statements and a method returning the code template. Save the file in your workspace, right-click and choose to add the component, and the palette updates.
What licence is amphi-etl released under?
A source-available variant of the Apache 2.0 terms, identified by the badge as the Elastic licence variant, which adds a condition on offering the software as a hosted or managed product. It is not the plain Apache licence, which matters if you plan to run it as an internal platform.
Does Amphi collect telemetry?
Yes, the readme states it collects anonymous telemetry data to understand users and their use cases, and that you can opt out in the settings and disable collection entirely. For a tool that emphasises self-hosting for privacy, the opt-out is the part to verify in your own installation.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/amphi-ai-amphi-etl)