Open-source project
nextgenusfs/funannotate avatar
nextgenusfs/funannotate

funannotate: a fungal genome annotation pipeline you run as a sequence of subcommands

Eukaryotic Genome Annotation Pipeline. Caveats are that GeneMark is not included in the docker image (see licensing below and you can complain to the developers for making it difficult to distribute/use).

401 stars95 forksPythonBSD-2-Clause

At a glance

What is it?
funannotate is a Python pipeline for eukaryotic genome annotation, built with fungi in mind. It chains cleaning, training, gene prediction and functional annotation into named subcommands, and its dependency story is the main thing to plan around.
Who is it for?
Adopt funannotate if you annotate fungal or other small eukaryotic genomes and you are willing to manage a conda environment plus a manually installed GeneMark-ES/ET for training-based prediction. Do not adopt it if you need a single self-contained container with every predictor included, or if you expect pip to pull in the external aligners and gene callers for you.
Can I use it commercially?
Yes. BSD-2-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 6 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gap funannotate fills for fungal genomes

A draft fungal genome arrives as a FASTA assembly. Turning that into a GFF3 file with gene models, product descriptions, GO terms and InterPro domains means running several unrelated programs, converting between their output formats, and reconciling disagreeing gene calls. funannotate exists to hold that sequence together. The README describes it as "a pipeline for genome annotation (built specifically for fungi, but will also work with higher eukaryotes)", and the subcommand structure reflects that: each stage is a command, and each command expects the previous stage's output.

The audience is a bench or core-facility bioinformatician who has an assembly and needs a published-quality annotation without writing the glue code. It is not a web service and not a library you import into an analysis script. You install it, you run subcommands, and you inspect files on disk between them.

How the stages chain: clean, train, predict, annotate

The pipeline is a linear data flow rather than a scheduler with a database behind it. Each subcommand reads files from a working directory and writes new files into it, so the state of a run is the contents of that directory.

funannotate clean takes the assembly and removes duplicate and short contigs. funannotate train uses RNA-seq alignments to produce training parameters for the ab initio gene predictors. funannotate predict then calls genes, combining the trained models with evidence such as protein alignments. funannotate annotate attaches functional information to the resulting models. The README's own smoke test names two of these stages directly, predict and train, which tells you they are the load-bearing ones.

The consequence of this design is that a failure in the middle leaves you with a partially populated directory and no rollback command. The README does not document rollback. You re-run the stage, or you start from a fresh directory, and you keep the intermediate files because they are the only record of what happened.

Installing funannotate with conda or Docker

The README gives two supported routes. The conda route is the one that puts funannotate on your $PATH as a normal command. You add the channels, then create an environment with a pinned Python range:

bash
conda config --add channels defaults
conda config --add channels bioconda
conda config --add channels conda-forge
conda create -n funannotate "python>=3.6,<3.9" funannotate

The version constraint is not cosmetic. setup.py declares REQUIRES_PYTHON as ">=3.6.0, <3.12", so the package installs across a wider range than the README's environment, but the README's own recipe narrows it to below 3.9. Follow the README rather than the metadata if you want the dependency set the maintainer tested. If the solve takes a long time, the README suggests mamba as a drop-in replacement:

bash
conda install -n base mamba
mamba create -n funannotate funannotate

For a container, the README documents pulling the image and an optional wrapper script that handles user and volume bindings:

bash
docker pull nextgenusfs/funannotate
wget -O funannotate-docker https://raw.githubusercontent.com/nextgenusfs/funannotate/master/funannotate-docker
chmod +x /path/to/funannotate-docker

After that the wrapper is invoked as if it were the funannotate executable. The README's example is `funannotate-docker test -t predict --cpus 12`, and it notes the wrapper may need to be on your $PATH. A slim image without the bundled databases is published as nextgenusfs/funannotate-slim. Two things to know about the full image: it is built from the latest code in master, so it can be ahead of the tagged releases, and GeneMark is not included in it.

GeneMark is the dependency you install by hand

This is the constraint that shapes real deployments. GeneMark-ES/ET is not distributed with funannotate, and the README is blunt about it: the docker image omits it, and the reader is invited to complain to the developers about the difficulty of redistributing it. Licensing is the reason.

So you obtain GeneMark from the developers at Georgia Tech, following their own instructions and accepting their licence terms. The README adds two mechanical steps that trip people up. First, change the shebang line of every Perl script in GeneMark to use `/usr/bin/env perl`. Second, make GeneMark findable, either by putting `gmes_petap.pl` on your $PATH or by setting the environment variable `$GENEMARK_PATH` to the gmes_petap directory.

If you skip GeneMark, you lose the training-based prediction path, which is the part of the pipeline that makes it worth running over a single gene caller. The README does not describe a supported no-GeneMark mode for prediction, so treat the manual install as part of the setup rather than an optional extra.

A first real run, and what the test command is for

Before you point funannotate at your own assembly, run the built-in test. It exercises the pipeline against bundled data and tells you whether the environment and the external binaries are wired up correctly:

bash
funannotate test -t predict --cpus 12

That is the README's example invocation, with the thread count set to the number of CPUs you want to give it. Expect it to take real time and to write output into a test directory; a clean exit means the predictors were found and ran. If it fails, the usual causes are a missing GeneMark install, a Perl shebang that still points at a specific interpreter, or a Python environment outside the documented range.

Once the test passes, the working pattern is one subcommand at a time against a directory you keep: clean the assembly, train on your RNA-seq, predict, then annotate. Each stage leaves files the next one consumes, so do not delete intermediates until the annotation is finished and you are satisfied with it.

Where funannotate is the wrong tool

The dependency surface is the honest limitation. funannotate orchestrates external aligners and gene predictors, and pip only installs the Python package. The README shows `python -m pip install funannotate` as a way to get the Python package, but the pipeline also needs the non-Python tools, and the README does not present pip as a complete installation. If you install that way and then run predict, missing binaries are the expected failure.

The second limitation is the environment. The README pins python>=3.6,<3.9, and the Dockerfile pins scikit-learn below 1.0.0 and biopython below 1.80. Those pins are what make the pipeline reproducible, and they are also what make it awkward to drop into a modern shared environment where something else already owns those packages. A conda environment per project is the practical answer.

Third, this is a fungal-scale tool. The README says it will work with higher eukaryotes, but the design target is small genomes. For a large plant or vertebrate genome, the training and prediction stages are not sized for that, and a tool built around that scale from the start is the better fit.

Alternatives and how they differ in approach

The most direct comparison is MAKER. MAKER is also an evidence-driven annotation pipeline, but it is configured through control files that you edit to declare which evidence sources and which ab initio predictors to use, and it is typically run through MPI across many machines. funannotate instead exposes a fixed sequence of subcommands with a smaller set of choices, and it is aimed at fungal genomes rather than being genome-size agnostic. If you want to hand-tune the evidence weighting and distribute the run across a cluster, MAKER's control-file model gives you that. If you want a documented path from assembly to annotated GFF3 for a fungal genome with fewer decisions, funannotate's subcommands are the shorter route.

A second alternative is to assemble the pieces yourself: run a gene predictor, run an alignment tool, and merge the results. That gives you full control and no pipeline version to track, at the cost of writing and maintaining the format conversions. funannotate's value is precisely that it has already made those conversions, and the price is that you inherit its pinned dependencies and its GeneMark requirement.

Maintenance, versioning and licence

The repository is not archived, and the last push was on 2024-03-01, which is also the date of the v1.8.17 release. The cadence before that was roughly annual: v1.8.13 in 2022-08-10 and v1.8.15 in 2023-04-11. Plan for a slow-moving project where you pin a version and upgrade deliberately rather than tracking master.

That matters because the Docker image is built from master, so the container can contain code newer than any tagged release. For reproducibility, pin the conda package version or build from a tag rather than pulling the image and assuming it matches v1.8.17.

funannotate itself is BSD-2-Clause, which is permissive and places few obligations on how you redistribute or use the pipeline. That licence does not extend to GeneMark, which you install separately under its own terms from Georgia Tech. The two licences are independent, and the README's note about redistribution difficulty is a licensing statement, not a technical one. Read GeneMark's terms yourself before you build it into anything you distribute.

Editorial conclusion

Adopt funannotate if you annotate fungal or other small eukaryotic genomes and you are willing to manage a conda environment plus a manually installed GeneMark-ES/ET for training-based prediction. Do not adopt it if you need a single self-contained container with every predictor included, or if you expect pip to pull in the external aligners and gene callers for you. Before committing, verify three things on your own data: that the conda solve resolves under python>=3.6,<3.9, that GeneMark is reachable through $GENEMARK_PATH or gmes_petap.pl on your $PATH, and that funannotate test -t predict runs to completion with the CPU count you intend to use.

Frequently asked questions

How do I install funannotate?

The README gives two routes: create a conda environment with the bioconda, conda-forge and defaults channels and the constraint python>=3.6,<3.9, or pull the nextgenusfs/funannotate Docker image and use the funannotate-docker wrapper script. A pip install of the funannotate package gets you the Python code but not the external tools the pipeline calls.

What is the funannotate train step for?

funannotate train produces the training parameters used by the ab initio gene predictors, which funannotate predict then uses when calling genes. The README's own test invocation targets the predict stage, and both stages are named in the pipeline's subcommand list.

Why is GeneMark missing from the funannotate Docker image?

The README states that GeneMark is not included in the docker image because of licensing, and points readers to the GeneMark developers for the download. You install GeneMark-ES/ET manually, then make it available either by adding gmes_petap.pl to $PATH or by setting $GENEMARK_PATH.

Which Python version does funannotate require?

The README's conda recipe pins python>=3.6,<3.9, while setup.py declares REQUIRES_PYTHON as >=3.6.0, <3.12. The README's narrower range is the one to follow if you want the dependency set the maintainer tested.

Is funannotate only for fungal genomes?

The README describes it as built specifically for fungi but says it will also work with higher eukaryotes. The design target is small genomes, so large plant or vertebrate assemblies are outside the intended scale.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
  4. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/nextgenusfs-funannotate.svg)](https://hysenlabs.com/projects/nextgenusfs-funannotate)