Biopython: a Python toolkit for sequence files, alignments and phylogenies
Official git repository for Biopython (originally converted from CVS)
At a glance
- What is it?
- Biopython is a long-running Python library for computational molecular biology. This review covers what it does, how to install it with pip or conda, where it stops being the right tool, and what to check before adopting it.
- Who is it for?
- Adopt Biopython if your work is file parsing, sequence handling, alignment I/O or tree manipulation in Python, and you are willing to read the Tutorial and Cookbook rather than guess at API behaviour. Do not adopt it as a wrapper around external programs you have not installed, and do not expect it to replace a dedicated aligner, a structural biology suite or a statistics package.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What the Biopython project actually provides
Biopython is not one tool. It is a collection of Python packages under the Bio namespace, assembled by what the README calls an international association of developers of freely available Python tools for computational molecular biology. The pyproject.toml package list shows the shape of it: Bio.Align and Bio.AlignIO for alignments, Bio.Blast and Bio.Emboss and Bio.Compass for talking to external search programs, Bio.Entrez for NCBI access, Bio.Phylo for trees, Bio.Graphics for drawing, BioSQL for database-backed storage, and more.
The audience is narrow and clear. If you write Python and your data is FASTA, GenBank, PDB, Newick or alignment files, this is the library the ecosystem expects you to reach for. If your work is wet-lab bench protocol or a web dashboard, nothing here is aimed at you.
One thing worth stating plainly: the repository description says it is the official git repository, originally converted from CVS. That history is visible in the API. Some modules carry decades of accumulated interface decisions, and the DEPRECATED.rst file exists precisely because breakages accumulate. The project treats that file as part of its public documentation, which is a reasonable signal about how it manages change.
How the pieces fit together: parsers, records and external programs
The mechanism most users meet first is the parser. You hand Bio a file handle and a format name, and you get back an iterator of record objects with attributes. That is the whole data flow for a large part of the library: file in, structured Python objects out, your code in between.
The second mechanism is the external program interface. Bio.Blast, Bio.Emboss and Bio.Compass do not reimplement BLAST, EMBOSS or ClustalW. The README lists standalone NCBI BLAST, EMBOSS and ClustalW as useful third party tools you may wish to install. That phrasing matters: the Python module builds a command line, runs the binary, and parses the output. If the binary is not on your PATH, the module has nothing to call. This is the single most common source of confusion for newcomers, and the README does not spell it out in the install section.
The third mechanism is the wrapper packages. Bio.Entrez ships DTDs and XSDs inside the package so that responses from NCBI can be validated locally. BioSQL is a schema plus adapters, and it needs a database driver: psycopg2 or PyGreSQL for PostgreSQL, MySQL Connector/Python or mysqlclient for MySQL. Those are optional dependencies, installed separately, and the README is explicit that they are only needed if you use those parts.
NumPy is the one hard dependency. Everything else is opt-in.
Installing Biopython with pip, and a first parse
The README's impatient section gives the install path directly. Pre-compiled binary wheels have been published on PyPI for Linux, macOS and Windows since Biopython 1.70, so pip should be quick and should not require a compiler.
pip install biopythonTo move to a newer release later, the README gives the upgrade form, and the uninstall form if you need to back out.
pip install --upgrade biopython
pip uninstall biopythonConda users have a second channel: the README carries a conda-forge badge pointing at the conda-forge Biopython package on anaconda.org, so the equivalent install is a conda-forge install of the biopython package rather than pip. Pick one and stay with it; mixing the two in a single environment is how you end up with two copies on the path.
For a first real use, the pattern is a file handle plus a format string. The README does not print a parse example, so treat the following as the shape of the call rather than a copied snippet: import the parser module, open your file, pass the handle and the format name, and iterate the records. The Tutorial and Cookbook at biopython.org/docs/latest/ is where the project puts the worked examples, and it is generated from the repository with Sphinx, so it tracks the code rather than drifting from it.
Before you write any of it, check the Python version. pyproject.toml sets requires-python to >=3.10, and the README recommends Python 3.13 while listing tested support for 3.10 through 3.14 plus the 3.15 release candidate, and PyPy3.10 v7.3.17 or later. If your environment is on 3.9 or older, pip will refuse the install, and that is the correct outcome.
Where Biopython is the wrong choice
The library is a toolkit, not a pipeline engine. It parses and represents biological data; it does not decide what analysis to run, and it will not scale a whole-genome workflow for you. If your job is aligning ten thousand short reads against a reference, you want an aligner, and Biopython's role shrinks to reading the output file afterwards.
External program wrappers are the sharpest failure mode. A module that shells out to BLAST will fail at runtime, not at import time, when the binary is absent or when its output format has changed between versions. The README lists those programs as third party tools you may wish to install and gives no version pinning guidance. That gap is real, and it means your reproducibility story depends on how you manage those binaries outside Python.
The optional dependencies are a second trap. Bio.Phylo plotting needs matplotlib; certain niche tree functions need networkx plus pygraphviz or pydot; the CDAO parser needs rdflib; Bio.Graphics needs ReportLab. None of these arrive with pip install biopython. A script that imports cleanly on your machine can fail on a colleague's because you installed matplotlib months ago and forgot.
Finally, consider the API surface itself. The package list in pyproject.toml runs to dozens of top-level modules, and the DEPRECATED.rst file exists because interfaces do get retired. A codebase that reaches into many corners of Bio at once will feel upgrade friction that a codebase using two modules will not.
Biopython against Bioconda and plain NumPy parsing
The most direct alternative for installation is Bioconda, the conda channel that ships bioinformatics software. The difference is not the library, it is the dependency solver. Conda resolves the whole environment, including compiled third party tools, from one channel with binary compatibility guarantees. pip resolves Python packages and leaves system binaries to you. If your workflow depends on BLAST or EMBOSS matching the version Biopython expects, conda's model is the better fit; if you are deploying into a container built around pip and a Python base image, pip is simpler and the README's wheel support makes it fast.
A second alternative is to skip the library and parse formats yourself with plain Python and NumPy. For a single FASTA file with simple headers, that is a dozen lines and no dependency. Biopython earns its place when the format is complex (GenBank, PDB, alignment formats with per-column metadata), when you need round-trip writing as well as reading, or when you want the parser maintained by people who track format revisions. For one-off scripts against a stable, trivial format, the library is overhead.
A third comparison is with domain-specific Python packages that wrap a single tool. Those tend to track one program's output format closely and change when it changes. Biopython's bet is breadth and a stable record interface across formats. That bet is why the DEPRECATED file is long.
Licence, maintenance and the cost of upgrading
The licence is the item to resolve before anything else. pyproject.toml declares license = "LicenseRef-Biopython-License-Agreement" with license-files = ["LICENSE.rst"], and the repository's top level carries both LICENSE and LICENSE.rst. GitHub reports the licence as NOASSERTION, meaning the platform's classifier could not map it to a standard identifier. The README calls the terms generous and points at LICENSE.rst, but generous is not a licence name. Read LICENSE.rst and, if you are distributing Biopython inside a product, get a real opinion on it. Nothing here is legal advice.
On maintenance: the repository is not archived, and the last push was on 2026-09-22. Activity alone does not tell you whether the parts you use are maintained. The project's own signals are better: NEWS.rst summarises changes per release, DEPRECATED.rst notes API breakages, and CONTRIBUTING.rst and CONTRIB.rst describe how to take part. Those three files are where upgrade cost actually lives.
Upgrade cost has two components. The Python side is the deprecation cycle: modules and call signatures get retired, and DEPRECATED.rst is the record. The binary side is the external programs, which Biopython does not pin. If you upgrade Biopython and your BLAST version has moved independently, output parsing is where it will hurt. Budget for testing the parsers you depend on, not the whole library.
Editorial conclusion
Adopt Biopython if your work is file parsing, sequence handling, alignment I/O or tree manipulation in Python, and you are willing to read the Tutorial and Cookbook rather than guess at API behaviour. Do not adopt it as a wrapper around external programs you have not installed, and do not expect it to replace a dedicated aligner, a structural biology suite or a statistics package. Before you commit, verify the licence text in LICENSE.rst against your distribution plans, confirm that the Python version you target is one of the supported ones listed in pyproject.toml, and check DEPRECATED.rst for the modules you plan to call.
Frequently asked questions
What is Biopython used for?
It is a set of Python packages for computational molecular biology, covering sequence and alignment parsing, phylogenetics, graphics, NCBI access through Bio.Entrez, and interfaces to external programs such as BLAST and EMBOSS. The README describes the project as an international association of developers of freely available Python tools for computational molecular biology.
Is Biopython free?
The README states that the package is open source software made available under generous terms, and points to LICENSE.rst for details. Note that pyproject.toml declares the licence as LicenseRef-Biopython-License-Agreement rather than a standard identifier, so read the licence file itself.
How to install Biopython?
The README's impatient section gives pip install biopython, with pip install --upgrade biopython to upgrade and pip uninstall biopython to remove it. Pre-compiled binary wheels have been provided on PyPI for Linux, macOS and Windows since Biopython 1.70, so a compiler is not normally required.
What is the difference between Python and Biopython?
Python is the language and runtime; Biopython is a library that runs on it. pyproject.toml requires Python 3.10 or later and lists NumPy as the only mandatory dependency, with everything else optional.
How to learn Biopython?
The README points to the user-centric documentation called The Biopython Tutorial and Cookbook, plus API documentation, at biopython.org/docs/latest/. That documentation is generated from the repository using Sphinx.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/biopython-biopython)