# microsoft/Biodiversity is a hub repo, not a toolkit: what it actually contains

> Microsoft's Biodiversity repository is an index for MegaDetector, Pytorch-Wildlife, SPARROW and the rest of the AI for Good Lab's conservation work. Here is what the hub does, how to install the one package that actually ships from it, and where you should go instead.

**microsoft/Biodiversity** — Microsoft AI for Good Lab — Biodiversity research hub. Open-source AI models, edge devices, and tools for biodiversity monitoring and conservation. Your source for MegaDetector, SPARROW, PytorchWildlife, Bioacoustics, and more.

- Repository: https://github.com/microsoft/Biodiversity
- Website: https://microsoft.github.io/Biodiversity/
- Stars: 1,076 · Forks: 297
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/microsoft-biodiversity

## What microsoft/Biodiversity is, and what it is not

The README is explicit that this repository is a hub. It says the work was deliberately broken into focused, dedicated repositories, one per project, and that this repo ties them together. That means the top-level directory is mostly an index: megadetector.md, a docs/ tree, mkdocs.yml, and a table of links. The table lists seven projects, including microsoft/MegaDetector for camera-trap detection of animals, people and vehicles, microsoft/MegaDetector-Acoustic for audio classification and species identification, microsoft/MegaDetector-Classifier for fine-tuning classifiers on your own datasets, microsoft/MegaDetector-Overhead for point-based wildlife localization from aerial views, microsoft/MegaDetector-Sonar for sidescan sonar imagery, microsoft/Pytorch-Wildlife as the deep learning framework and model zoo, and microsoft/SPARROW as a solar-powered acoustic and remote recording edge device.

The audience is conservation technologists, ecologists and research engineers who already know they want to run detection or classification on field data and need to find the right component. The hub solves a discovery problem, not a modelling one. If you arrived expecting a single installable library that does camera traps and bioacoustics, you will be disappointed: the only package this repository publishes is PytorchWildlife, and the README points you elsewhere for everything else.

That split is defensible. The README says keeping everything in one repository made code harder to find, harder to maintain, and harder to extend. The cost is that version numbers, issue trackers and release cadences now live in seven places, and the hub does not track them.

## How the hub is organised and how PytorchWildlife is packaged

The repository layout tells you what is local and what is remote. Alongside the docs and the Dockerfile there are three working directories: PytorchWildlife/, PW_Bioacoustics/ and PW_FT_classification/, plus PW_FT_detection/. The demo/ directory holds runnable entry points, including image_demo.py, video_demo.py, gradio_demo.py, image_demo_OWL.py, image_demo_herdnet.py and a detection_classification_pipeline_demo.py, with matching notebooks for Colab.

setup.py is the clearest statement of what ships. It reads the version from version.txt, names the distribution PytorchWildlife, and declares python_requires of at least 3.10. The dependency list is where the real engineering constraints show: supervision is pinned to exactly 0.23.0, and gradio is constrained to the 6.x line with a comment noting that the pin exists to include the fix for GHSA-6655-8ph2-63j3 and earlier arbitrary-file-read advisories. The remaining dependencies are the usual deep learning stack (torch, torchvision, torchaudio, ultralytics, yolov5, timm, lightning, omegaconf, scikit-learn) plus wget, chardet and Pillow.

requirements.txt is broader than setup.py because it also carries the documentation toolchain (mkdocs, mkdocs-material, mkdocstrings) and the bioacoustics dependencies (librosa, soundfile, pyyaml, torchmetrics, onnxruntime). That is a development file, not a runtime one. If you copy it wholesale into a production image you will install a documentation site you do not need.

## Installing PytorchWildlife and running a first detection

The package is published on PyPI under the name PytorchWildlife, and setup.py confirms the distribution name. The README links the PyPI badge and a Hugging Face demo space, so pip is the documented route. The setup.py metadata requires Python 3.10 or newer.

```bash
pip install PytorchWildlife
```

After installation, the demo/ directory is the fastest way to see the library do something. image_demo.py is the still-image entry point, and the repository also ships image_detection_demo.ipynb and video_demo.py for the same pipeline over video. The demos are the only usage examples present in the repository, so treat them as the reference rather than looking for a documented API surface in the README.

If you prefer a container, the Dockerfile builds from python:3.11-slim, installs libgl1-mesa-glx, libglib2.0-0 and ffmpeg, copies the repository into /app, and runs pip install --no-cache-dir PytorchWildlife. It exposes port 80 and sets two environment variables.

```dockerfile
ARG GRADIO_SERVER_NAME="0.0.0.0"
ENV GRADIO_SERVER_NAME=${GRADIO_SERVER_NAME}

ARG GRADIO_SERVER_PORT="80"
ENV GRADIO_SERVER_PORT=${GRADIO_SERVER_PORT}
```

The Dockerfile comments are unusually direct about the risk here. They state that GRADIO_SERVER_NAME defaults to 0.0.0.0 so the container can serve on its published port, and that you should only publish that port on a trusted network or behind an authenticating reverse proxy because the Gradio demo has no built-in auth. Take that literally. The same file notes that the base image was moved off Python 3.8, which reached end of life in October 2024 and pinned gradio below the CVE fix line.

For bioacoustic work, requirements.txt lists librosa, soundfile, pyyaml, torchmetrics and onnxruntime, and the repository contains a PW_Bioacoustics/ directory. The README does not document a bioacoustics command line, so the code in that directory is the source of truth.

## Where this hub is the wrong tool

The most concrete limitation is that the hub does not version the projects it links to. If microsoft/MegaDetector-Overhead changes, nothing in this repository records it. You get a markdown table entry, not a pinned commit or a compatibility matrix. For a team that needs reproducible field pipelines, that is a real gap: your lockfile will capture PytorchWildlife, but the other six components have to be tracked separately.

Second, the packaging is honest about its maturity. setup.py declares Development Status 3 - Alpha. That is a self-assessment from the maintainers, and it should shape how much you build on top of the API without pinning. Combined with the exact pin on supervision==0.23.0, you should expect dependency resolution conflicts if your environment already uses a different supervision release.

Third, the Gradio demo is not a deployment target. The Dockerfile itself warns that it has no built-in auth and that port 80 should only be published on a trusted network or behind an authenticating reverse proxy. If your requirement is a multi-user web service with access control, this repository gives you a demo, not a product.

Fourth, the README does not document rollback, upgrade procedures, or a deprecation policy. If you need a supported upgrade path between releases, nothing here supplies one, and you should check the release notes page the README links under Previous versions before committing.

## How this differs from using MegaDetector directly

The real alternative is not a competing product; it is going straight to microsoft/MegaDetector and skipping the framework layer. The README describes MegaDetector as the camera-trap animal detection model that started the whole effort, and it is the component the hub's own citation guidance singles out: Beery, Morris, Yang 2019 is the paper to cite for any use of MegaDetector specifically, while Hernandez et al. 2024 covers the PyTorch-Wildlife framework and models accessed through it.

The difference in approach matters. MegaDetector is a detection model with a long track record in the conservation community; the README notes that a list of organisations using it across global conservation work is maintained on the MegaDetector repository. PytorchWildlife, by contrast, is described as the collaborative deep learning framework and model zoo, which means it wraps detection and classification behind a shared interface and brings the heavier dependency set (ultralytics, yolov5, lightning, timm, gradio).

If your pipeline is detection-only and you want the smallest dependency surface, going to MegaDetector directly is the leaner choice. If you need classification on top of detection, the repository's demo/detection_classification_pipeline_demo.py and PW_FT_classification/ directory show that the framework is built for that chained workflow, and the overhead is the point. For audio, neither of those applies: the relevant entry is microsoft/MegaDetector-Acoustic, and for hardware it is microsoft/SPARROW.

## Licence, citation and the cost of keeping up

The repository is MIT licensed, confirmed by the LICENSE file and by setup.py's license field and its OSI Approved :: MIT License classifier. MIT is permissive, so the practical implication is that you can use and redistribute the code with attribution and without a copyleft obligation on your own work. That is a statement about the licence text, not legal advice; if you are shipping a commercial product or need to know how the licence interacts with model weights distributed separately, ask your own counsel.

The citation obligation is separate from the licence and is easy to miss. The README asks that work using any repository under the umbrella cite Hernandez et al. 2024 for the framework or models accessed through it, and Beery, Morris, Yang 2019 for MegaDetector specifically. A citation.cff file is included for automated citation tools, which means reference managers can pull the metadata directly rather than you transcribing it.

Upgrade cost is the weak point. The last push to this repository was on 2026-08-25, and the most recent release listed is pw_v1.3.0 from 2026-04-22, following pw_v1.2.1 in April 2025 and pw_v1.2.0 in January 2025. The gap between v1.2.1 and v1.3.0 is roughly a year, so plan for infrequent but potentially large jumps. Because setup.py pins supervision to an exact version and constrains gradio to a major line, a single version bump can force changes across your environment. Budget time for that, and read the release notes the README links before upgrading in a working field pipeline.

## Conclusion

Adopt the hub as a map, not a dependency. If you are starting a camera-trap or bioacoustics pipeline, install PytorchWildlife from this repository and follow the links to microsoft/Pytorch-Wildlife and microsoft/MegaDetector for the code that is actually maintained there. If you need a versioned, supported library with a stable API, this is the wrong entry point: setup.py still declares Development Status 3 - Alpha, the Dockerfile ships a Gradio demo with no built-in authentication, and the README documents no rollback or upgrade path. Verify first that the sub-repository you need exists and is the one carrying the code you want, then check the citation requirements in citation.cff before you publish anything built on it.

## FAQ

### Is microsoft/Biodiversity a library I can install, or just a list of links?

It is primarily a hub. The README states the work was split into dedicated repositories, one per project, and this repository ties them together. The one package it publishes is PytorchWildlife, declared in setup.py.

### What Python version does PytorchWildlife need?

setup.py sets python_requires to >=3.10, and the Dockerfile builds from python:3.11-slim. The Dockerfile comments note the base image was moved off Python 3.8, which reached end of life in October 2024.

### Which projects are listed under the microsoft/Biodiversity hub?

The README table lists microsoft/MegaDetector, microsoft/MegaDetector-Acoustic, microsoft/MegaDetector-Classifier, microsoft/MegaDetector-Overhead, microsoft/MegaDetector-Sonar, microsoft/Pytorch-Wildlife and microsoft/SPARROW.

### Is the Gradio demo in the Dockerfile safe to expose publicly?

The Dockerfile comments say no. They state that GRADIO_SERVER_NAME defaults to 0.0.0.0 and that you should only publish that port on a trusted network or behind an authenticating reverse proxy, because the Gradio demo has no built-in auth.

### What should I cite if I publish work using these models?

The README asks for Hernandez et al. 2024 for any use of the PyTorch-Wildlife framework or models accessed through it, and Beery, Morris, Yang 2019 for any use of MegaDetector specifically. A citation.cff file is included for automated citation tools.

## Sources

- [License: MIT](https://github.com/microsoft/Biodiversity/blob/main/LICENSE)
- [microsoft/Biodiversity on GitHub](https://github.com/microsoft/Biodiversity)
- [Project website](https://microsoft.github.io/Biodiversity/)
- [README](https://github.com/microsoft/Biodiversity/blob/main/README.md)
- [Releases](https://github.com/microsoft/Biodiversity/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/microsoft-biodiversity
