Self-hosted service
nextflow-io/nextflow avatar
nextflow-io/nextflow

Nextflow: a dataflow workflow engine for pipelines that outgrow a shell script

A DSL for data-driven computational pipelines

3,494 stars813 forksGroovyApache-2.0

At a glance

What is it?
Nextflow is a Groovy-based DSL and runtime for data-driven computational pipelines, installable with a single curl command and deployable from a laptop to HPC schedulers and cloud batch services. It is strongest when the work is many independent tasks over files, and weakest when you want a small, single-machine script.
Who is it for?
Adopt Nextflow when your pipeline is a set of independent tasks over files and you need to move the same script between a laptop, an HPC scheduler and a cloud batch service without rewriting it. Do not adopt it for a handful of sequential shell commands on one machine; a makefile or a script is less machinery.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Groovy, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem Nextflow solves, and who ends up using it

A bioinformatics pipeline usually starts as a shell script. One step produces files, the next step consumes them, and the loop over samples is written by hand. That works until the sample count grows, the steps need to run on a cluster, or someone else has to reproduce the result on different hardware. At that point the script has to encode scheduling, dependency ordering, retries, container selection and partial re-execution, and none of that is what the author wanted to write.

Nextflow targets exactly that gap. The README describes it as "a workflow system for creating scalable, portable, and reproducible workflows" based on the dataflow programming model, which the project says simplifies writing parallel and distributed pipelines so you can focus on the flow of data and computation. In practice the user writes processes, each a self-contained command with declared inputs and outputs, and the engine decides what can run at once and where.

The audience is narrower than the topic list suggests. The repository topics include bioinformatics, reproducible-research and nf-core, and the README points at the nf-core project as "a community effort aggregating high quality Nextflow workflows". If your work is batch computation over files, the model fits. If your work is a long-running service with request-response traffic, it does not.

How the dataflow model actually drives execution

Nextflow is written in Groovy and, according to the README credits, is built on Groovy and GPars. The README opens with a quotation about dataflow variables being expressive in concurrent programming, which is the design idea rather than a marketing line: a process does not call the next one, it declares what it emits, and the runtime connects producers to consumers.

That indirection is what buys portability. The same pipeline definition can be handed to different executors. The README lists local execution, HPC schedulers, AWS Batch, Azure Batch, Google Cloud Batch and Kubernetes, and the repository topics name SGE and Slurm explicitly. Software dependencies are handled separately from execution, with Conda, Spack, Docker, Podman and Singularity named in the README. The separation matters: the pipeline text describes the computation, while a configuration file describes where it runs and which container or environment each process uses.

The trade-off is that the engine owns the schedule. You cannot easily interleave arbitrary host commands between tasks, and debugging means reading the work directory the engine creates rather than stepping through a script. The Makefile in the repository shows the same convention from the developer side, with a clean target that removes .nextflow* and work directories.

Installing Nextflow and running a first pipeline

The README gives the installation as a single command. It downloads the launcher and creates a nextflow executable in the current directory, which you then move somewhere on your PATH.

bash
curl -fsSL https://get.nextflow.io | bash

There is a second supported route for people who already manage software with conda. The README lists it as an alternative, not a replacement.

bash
conda install -c bioconda nextflow

After either route, the launcher creates the nextflow executable in the current directory, and the README says you can move it to a directory in your $PATH to run it from anywhere. The repository keeps a VERSION file and a CHANGELOG.md at the top level, and releases are tagged with version numbers such as v26.09.0-edge, so the version string is the thing to record alongside your results.

For a first real pipeline, the repository ships an examples directory containing examples/agents and examples/data. The README does not walk through a hello-world script, so the honest starting point is the documentation site it links for stable and edge releases, plus the nf-core workflows if you want something already written. What you should expect from a first run is a work directory containing one subdirectory per task invocation, which is the unit the resume mechanism keys on.

Where Nextflow is the wrong tool

Nextflow assumes a particular shape of work: many tasks, each with declared inputs and outputs, executed by an external command. If your pipeline is five sequential steps that each depend on the previous one and take an hour, the dataflow machinery adds configuration surface without adding parallelism. A shell script with set -e is easier to read and easier to hand to a colleague.

The second limitation is operational. The engine creates and manages a work directory, and the Makefile's clean target shows how much state accumulates, removing .nextflow* and work. On a shared cluster, that directory can grow quickly and it is your responsibility, not the engine's, to decide what to keep. The README does not document rollback or a built-in cleanup policy, so lifecycle management is something you bring.

The third is the dependency on a JVM toolchain and on the Groovy-based runtime. The repository contains a gradlew wrapper, build.gradle, settings.gradle and a compile.sh, which tells you the project is a Gradle build with a launcher script. That is fine for users of the released launcher, but it means the project is not a single static binary, and environments that forbid a JVM are out of scope.

Finally, the release stream matters. The most recent releases listed are v26.09.0-edge, v26.08.0-edge and v26.07.0-edge, all edge builds. The README points separately at stable and edge documentation, so the distinction is deliberate. If you need a frozen toolchain for regulated work, verify which channel you are installing from rather than assuming the launcher gives you a stable release.

Nextflow versus Snakemake: two answers to the same question

The comparison people reach for is Snakemake, and the difference is architectural rather than cosmetic. Snakemake is a Python-based system where you declare rules with input and output file patterns, and the tool resolves the dependency graph from those filenames. Nextflow, as the README states, is based on the dataflow programming model, where processes emit values and channels carry them to the next process. File patterns are not the primary coordination mechanism; data channels are.

The practical consequence is where each one feels natural. If your pipeline is a set of file transformations over a directory tree, Snakemake's rule-per-file-pattern style maps directly onto how you already think about the problem, and you stay inside Python. If your pipeline has branching, fan-out over sample collections, or needs to move unchanged between a local machine and a cloud batch service, Nextflow's channel model and executor abstraction carry more of that weight for you.

Both are legitimate choices and both have large communities. Nextflow's is organized around nf-core, which the README describes as aggregating high quality workflows that everyone can use, and around a forum and Slack linked from the README. That ecosystem is a real factor if you would rather adopt a published pipeline than write one.

Maintenance, releases and what the Apache 2.0 licence means here

The repository is not archived, and the last push was on 2026-09-23. Releases arrive frequently: v26.09.0-edge on 2026-09-22, v26.08.0-edge on 2026-08-20 and v26.07.0-edge on 2026-07-15 are the three most recent. That cadence is a maintenance cost as much as a signal. A pipeline validated against one version should record that version, because the edge channel is where the changes land first.

Nextflow is released under the Apache 2.0 licence, and the repository carries COPYING and NOTICE files alongside the source. The README also states that Nextflow is a registered trademark, linking to a trademark policy. That distinction is worth understanding before you build a product around the name: the code licence and the brand are governed separately. This is a description of what the repository says, not legal advice; if you plan to redistribute or rebrand, read COPYING and the trademark policy.

Upgrade cost is mostly in your own configuration. The pipeline text is the stable part; executor settings and container or conda environment definitions are what you re-verify when the engine changes. Keeping those in a config file rather than scattered through the pipeline is what makes an upgrade a review rather than a rewrite.

Editorial conclusion

Adopt Nextflow when your pipeline is a set of independent tasks over files and you need to move the same script between a laptop, an HPC scheduler and a cloud batch service without rewriting it. Do not adopt it for a handful of sequential shell commands on one machine; a makefile or a script is less machinery. Before committing, verify three things in the documentation: which executor you will target, how conda, Docker, Podman or Singularity will supply your tools, and how the resume mechanism interacts with your working directory. The repository is not archived and the last push was on 2026-09-23, so the codebase is moving; pin the version you validate rather than tracking edge releases.

Frequently asked questions

What is Nextflow used for?

It is a workflow system for creating scalable, portable and reproducible workflows, based on the dataflow programming model. The README says it can deploy workflows on a local machine, HPC schedulers, AWS Batch, Azure Batch, Google Cloud Batch and Kubernetes.

Is Nextflow better than Snakemake?

The README does not rank them. The concrete difference it describes is the model: Nextflow is based on dataflow programming, where processes emit values onto channels, rather than on file-pattern rules. Choose based on whether channel-based fan-out or file-pattern rules match your pipeline.

Is Nextflow free?

Nextflow is released under the Apache 2.0 licence. The README also notes that Nextflow is a registered trademark, which is a separate matter from the code licence.

How do I install Nextflow on Linux?

The README gives a single command, curl -fsSL https://get.nextflow.io | bash, which creates a nextflow executable in the current directory that you can move onto your PATH. It can also be installed from Bioconda with conda install -c bioconda nextflow.

What is a Nextflow pipeline in bioinformatics?

It is a workflow definition written in Nextflow's DSL, where each process declares its inputs and outputs and the runtime connects them. The repository topics include bioinformatics and reproducible-research, and the README points to nf-core as a community collection of ready-made Nextflow workflows.

How do I install Nextflow on Windows or macOS?

The README documents the curl installer and the Bioconda package but does not give platform-specific instructions for Windows or macOS. On macOS the curl command works in a shell; on Windows the README is silent, so check the documentation site it links for the current guidance.

Official sources

  1. License: Apache-2.0
  2. nextflow-io/nextflow on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/nextflow-io-nextflow.svg)](https://hysenlabs.com/projects/nextflow-io-nextflow)