# open-data-pipeline: three server apps, four AI config dirs, one logged file

> HelicalInsight's OpenDataPipeline wraps Airflow, Spark, dlt and MongoDB behind a React interface, with LangChain-driven question answering over the transformation step and four LLM providers to choose from. What the README does not lead with is in the repository itself: a committed app.log at the root, four separate directories for configuring coding assistants, three server applications rather than one, and an admin role you grant by opening a MongoDB shell with a password printed in the documentation.

**helicalinsight/open-data-pipeline** — Open Data Pipeline is an AI powered data migration and data transformation tool

- Repository: https://github.com/helicalinsight/open-data-pipeline
- Stars: 656 · Forks: 633
- Language: Python
- License: Apache-2.0
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/helicalinsight-open-data-pipeline

## app.log is committed at the root of a data pipeline repository

The top level of this repository is twenty entries and one of them is a log file. app.log sits at the repository root, next to LICENSE and pytest.ini.

A committed log file is not a small thing in a project whose whole job is moving data. Logs are exactly the artefact you would expect to grow without bound, to contain whatever was in the payload at the time, and to be useless a week later.

There is a .gitignore in the tree, and there is a separate logs location documented for the running system, which is a directory under the setup path named after the organisation. So the application does have a proper log directory, and the root log file is something else.

The recorded counts are worth reading next to that. 656 stars, 633 forks, and zero open issues, on a repository with no tagged releases at all. A fork count that nearly matches the star count with an empty issue tracker suggests a deployment-shaped usage pattern, where people fork it to run their own pipelines rather than to contribute to it.

## Four directories exist to tell AI assistants how to read the codebase

Before you reach the source, there are four directories and one file whose only purpose is configuration for coding assistants. There is a .cursorrules file at the root, and there are .agent/, .askondata/ and odp_code_context/ directories.

That is four separate mechanisms for the same purpose, aimed at four different tools. It suggests an accumulation rather than a decision: each assistant convention was adopted as it became relevant and none of the earlier ones was removed.

The consequence is that the guidance can drift. If .cursorrules and odp_code_context disagree about where the code lives or how a task should be approached, there is no mechanism described that resolves it, and an assistant will pick one.

The rest of the root is the application itself, and it is more interesting for the shape than the contents. There is an airflow directory, which suggests the scheduler configuration is versioned alongside the code rather than being generated by the setup script.

## Three server applications sit beside the orchestration directory

The repository root contains what look like three separate services rather than one application. There is a dlt_server_app, a spark_server_app, and an opendatapipeline directory, alongside an airflow directory and a shared core.

That maps onto the architecture the README describes, which is four components: a React frontend, an Apache Airflow scheduler and webserver for orchestration, Apache Spark with dlt and Pandas for the compute, and MongoDB for metadata storage.

Two separate server applications for two compute engines is the notable design choice. dlt is the extraction and loading library, and Spark is the heavy processing engine, and giving each its own server means the light path does not have to start a JVM cluster.

The remaining directories are client/ for the frontend, core/ for shared code, and audit_tracker/, which has no counterpart anywhere in the architecture section of the README. An audit tracker with no documented role is the kind of thing worth asking about before you rely on it.

The declared architecture is also four layers deep and nothing else: no gateway, no API service of its own, no queue. The React frontend talks to the two compute servers and to Airflow.

## Admin access is granted by editing MongoDB, with the password written in the README

There are no user roles in the configuration file. To make a local user an admin you go into the database, and the procedure is four steps with the credentials supplied in the text.

First find the primary MongoDB container, the instruction being to run docker ps and grep for mongo_primary. Then open a shell inside it with docker exec. Then connect with the mongo shell:

```bash
   mongosh --username askondata --port 27021 --authenticationDatabase user_sessions
   
```

The username is askondata, the port is 27021 rather than the default, and the authentication database is user_sessions. The password is given in the following line as askondata, the same word as the username.

The last step switches to the user_sessions database and runs an update setting the role field on one user record, matched by email address.

Two things follow from this. The users collection and its role field are the authorisation model, and it lives in a database also used for session state, so sessions and identity share one store. And the default credential pair is documented in a public README, which is acceptable for a local setup you run yourself and is not acceptable anywhere reachable from a network.

## Hot reload is enabled by editing a Python file in the setup directory

The developer tips describe enabling automatic reloading during backend development, and the method is to open a file in the setup directory and add a flag.

The file is run.py inside opendatapipeline_src under the setup path, and the flag goes immediately after the word gunicorn in the command arguments. So the backend runs under gunicorn, and the reload flag is added by editing a line in a Python source file rather than by setting an environment variable.

The reason is visible in the instruction: this is a file in your setup directory, not in the repository. The setup script generates it, so the edit is local and does not propagate. That is a defensible choice for a generated deployment file.

The trade-off is that it is not discoverable. There is no configuration key named anything like reload, and a developer who has not read this section will try to restart the container instead.

The application logs, for comparison, are documented as a normal path under the setup directory rather than as something to be edited, which is the distinction that makes the previous instruction feel unusual.

## Four providers, and the Ollama one is configured with a full URL

The provider configuration is one environment block, and it is where a setup goes wrong most often, because the four providers do not take the same shape.

```
You can also configure the LLM provider using the following environment variables:
- `LLM_PROVIDER`: `ollama`, `openai`, `anthropic`, `google`
- `OPENAI_API_KEY`: Your OpenAI API key
- `OPENAI_MODEL`: e.g., `gpt-4o`
- `ANTHROPIC_API_KEY`: Your Anthropic API key
- `ANTHROPIC_MODEL`: e.g., `claude-3-5-sonnet-20241022`
- `GOOGLE_API_KEY`: Your Gemini API Key
- `GOOGLE_MODEL`: Your Gemini model to use
- `OLLAMA_BASE_URL`: e.g., `http://localhost:11434`
- `OLLAMA_MODEL`: e.g., `openhermes`
- `LLM_TEMPERATURE`: e.g., `0`
- `LLM_MAX_TOKENS`: e.g., `1000`
```

OpenAI, Anthropic and Google all take a key and a model name, and the two examples given are gpt-4o and claude-3-5-sonnet-20241022. Ollama takes neither a key nor a model in the key-and-model sense: it takes a base URL and a model name, with the examples being a local port address and openhermes.

That is the right shape for a self-hosted model, since there is no hosted endpoint to authenticate against. It also means the URL has to be reachable from inside the container rather than from your machine, which is why it is a full URL and not a host name.

Two global settings sit alongside the per-provider ones: a temperature and a max tokens value. Both are given as examples rather than defaults, with the temperature example being zero.

One naming wrinkle: the prose names Google Gemini as the provider, while the accepted value for the provider variable is google. Everything else, ollama, openai and anthropic, matches the product name.

## No releases, and the documentation link points at an API endpoint

There are no GitHub releases for this repository, so there is no version to pin and no upgrade note to read. The last recorded push is 2026-08-04 and the repository is not archived.

The documentation pointer is a single link to an API documentation path on the hosted instance, and the sentence around it promises technical knowledge hubs in the plural. One URL cannot be several hubs, so either the others are unlinked or the wording is aspirational.

The easier entry point is the hosted one. The page says the easiest way to get started is the managed production instance, with a self-hosted Docker path as the alternative. That inverts the usual open source pitch, where the local install is the default and the hosted version is a convenience.

For a data migration tool that is arguably correct, since the local path needs more than eight gigabytes of RAM, more than fifteen gigabytes of free disk, Docker and Node, and five to twenty minutes of setup.

## One script does the whole local install and expects to be run from the root

There is one command for the entire local environment: bash on the setup script at the repository root. Everything else in the getting started is either a prerequisite or something you do after it finishes.

The prerequisites are Docker at version 24 or newer and Node.js at version 18 or newer. The script is expected to take five to twenty minutes depending on the machine and the internet connection, and the instructions say they were tested on WSL2 and should work on any Linux machine.

The script spins up three things: an Airflow webserver, a Spark cluster and MongoDB. The README points at a single command to start them, and verification is watching the container list until everything reports healthy, at which point the application is reachable over HTTPS on localhost.

Two things are left to the reader. The top-level configurations inside the script can be edited to supply custom values before running it, and the environment file is a copy of an example that ships ready to use. Neither is explained further.

## Conclusion

Use this if you want Airflow and Spark behind a web interface with a conversational transformation step and you already run containers, since the local path is a single script that needs Docker, Node, 8GB of RAM and 5 to 20 minutes. Skip it if you want a pinned version to deploy, because there are no releases at all. Before you run it anywhere shared, change the documented MongoDB credential, since the README prints the username and password in the open, and note that the local instance is served over HTTPS on localhost with no certificate guidance given.

## FAQ

### What is a data pipeline?

In this repository it is four layers: a React frontend for managing pipelines, an Apache Airflow scheduler and webserver for orchestration, Apache Spark with dlt and Pandas for compute, and MongoDB for metadata storage. Extraction and loading go through dlt, which handles REST APIs, databases and files, and the transformation step is where the LangChain-driven question answering sits.

### How do I run OpenDataPipeline locally?

Install Docker at version 24 or newer and Node.js at version 18 or newer, then run the setup script from the root of the cloned repository. It spins up an Airflow webserver, a Spark cluster and MongoDB, takes five to twenty minutes, and wants more than 8GB of RAM and more than 15GB of free disk. You then watch the container health status and open the application on https://localhost.

### Which LLM providers does OpenDataPipeline support?

Ollama, OpenAI, Anthropic and Google Gemini, selected with the LLM_PROVIDER environment variable whose accepted values are ollama, openai, anthropic and google. OpenAI, Anthropic and Google take an API key and a model name, while Ollama takes a base URL and a model name. Two shared settings, a temperature and a max tokens value, apply across providers.

### How do I make a user an admin in OpenDataPipeline?

There is no configuration flag for it. You find the primary MongoDB container, open a shell in it, connect with mongosh using the username askondata on port 27021 against the user_sessions authentication database, then run an update on the users collection setting the role field to admin for the matching email address.

### Does OpenDataPipeline have any releases?

No. The repository has no GitHub releases at all, so there is no version to pin. It is not archived and the last recorded push is 2026-08-04. The documented starting point is the hosted managed instance rather than a self-hosted install.

## Sources

- [helicalinsight/open-data-pipeline on GitHub](https://github.com/helicalinsight/open-data-pipeline)
- [Issues](https://github.com/helicalinsight/open-data-pipeline/issues)
- [License: Apache-2.0](https://github.com/helicalinsight/open-data-pipeline/blob/main/LICENSE)
- [README](https://github.com/helicalinsight/open-data-pipeline/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/helicalinsight-open-data-pipeline
