bugbug: Mozilla's Machine Learning Platform for Bug Triage and Test Selection
Platform for Machine Learning projects on Software Engineering
At a glance
- What is it?
- bugbug is an open-source Python platform from Mozilla that applies machine learning to Bugzilla bug management and software quality tasks, including classifying defects by type, predicting which patches are likely to cause regressions, selecting relevant tests for a given change, and routing bugs to the right team member. It is designed for large software projects with substantial Bugzilla history and is used in production for the Firefox development pipeline.
- Who is it for?
- bugbug is best suited for large open-source projects with an established Bugzilla workflow, a significant body of labeled bug history, and engineering capacity to operate a Python ML service. Teams on GitHub with few hundred bugs in their history will not have enough training data to get useful model accuracy.
- Can I use it commercially?
- Yes, with conditions. MPL-2.0 is a weak copyleft licence: you can use it inside commercial and closed-source software, but if you distribute changes to its own files, you must publish those changes under the same licence.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The Problem bugbug Addresses
Large software projects on Bugzilla face two recurring costs: incoming bugs require manual triage before the right engineer sees them, and every code change requires a decision about which tests to run. Manual triage does not scale as bug volume grows, and running every test on every patch is too slow for a browser-engine project. The Mozilla Hacks blog posts from 2019 and 2020 referenced in the README describe both problems for Firefox specifically.
bugbug addresses both by training ML classifiers on the historical record of Bugzilla bugs and Mercurial commits. A classifier trained on past bugs can score new incoming bugs for type, assignee, component, and whether they need QA verification, without requiring a developer to read each one. A test selection model can narrow the test suite to the tests most likely to catch regressions in a particular change.
The 18 Classifiers and What Each Predicts
The README documents 18 classifiers, each a separately trainable model:
assignee predicts the appropriate engineer for a bug. backout detects patches likely to be reverted due to build or test failures. bugtype labels bugs as crash, memory, performance, or security. component assigns product and component to untriaged bugs. defect distinguishes real defects from feature requests and tasks, with reported accuracy around 93 percent and precision around 95 percent on the labeled dataset. defect vs enhancement vs task extends the defect classifier to a three-way split.
devdocneeded identifies bugs that require developer documentation. needsdiagnosis flags webcompat issues that are likely invalid. qaneeded detects bugs that require QA verification. regression vs non-regression detects regressions despite inconsistent use of Bugzilla's regression keyword. regressionrange detects regression bugs that include a specific revision range.
regressor identifies patches likely to cause regressions, for routing to additional review. spam catches spam bug reports. stepstoreproduce detects bugs with steps to reproduce. testfailure predicts which patches are likely to cause test failures. testselect selects the relevant tests for a given patch. tracking detects bugs that should be tracked. uplift detects bugs for which release uplift should be approved or rejected.
These classifiers reflect Mozilla's specific workflow. Projects with different bug taxonomies or different CI systems would need to retrain the models on their own data, and some classifiers (such as uplift) are specific to the Firefox release process.
Installing bugbug and Running a Classifier
bugbug uses uv (docs.astral.sh/uv) for dependency management. The README requires Python 3.12 or higher. Installing the base dependencies runs in one command:
uv syncFor test dependencies, add --group test. For NLP extras, add --extra nlp. The repository also documents an optional libgit2 requirement for the repository mining script. On Debian, that version is currently only available in the experimental repository:
sudo apt-get -t experimental install libgit2-devTraining a classifier requires running the trainer script with the model name. The README warns this takes 30 minutes or more:
python3 -m scripts.trainer defectTo classify a bug using an already-trained model (or by downloading a pre-trained model automatically on first run):
python -m scripts.bug_classifier defect --bug-id ID_OF_A_BUG_FROM_BUGZILLAThe repository also supports training on Mozilla's Taskcluster CI by adding a line to the pull request description:
Train on Taskcluster: spambugThis re-runs on every push to the linked branch, so the README advises limiting push frequency or temporarily removing the line to avoid wasting CI resources. The docker-compose.yml in the repository defines services for the HTTP service, a background worker, and commit retrieval, making containerized deployment straightforward for teams with Docker.
Limitations and When Not to Use bugbug
The practical ceiling of bugbug's value is the size and quality of the training data. The defect classifier's documented accuracy of around 93 percent comes from a dataset of 2,110 bugs. Projects with fewer labeled bugs will produce less reliable models. The README does not state a minimum dataset size for each classifier, but the principle is clear: the models learn patterns from historical data, and sparse history yields uncertain predictions.
Several classifiers are specific to the Mozilla/Firefox workflow. The uplift classifier is explicitly about Firefox release policy. The needsdiagnosis classifier targets the webcompat.com issue tracker. Teams outside Mozilla would either skip these or retrain them with their own data and labels, which requires annotating enough historical bugs to serve as a training set.
The repository mining script for commit-level features requires libgit2 v1.0.0, which the README notes is only available in Debian's experimental repository as of the documentation. This is a concrete operational constraint for teams on standard Debian or Ubuntu installations.
bugbug does not provide a UI for non-technical stakeholders to view classification results. The HTTP service exposes an API, but consuming it requires integrating with the project's workflow externally.
Comparison with SonarQube
SonarQube (sonarsource.com) is a widely used code quality and static analysis platform. It scans source code using deterministic rules for bugs, code smells, and security vulnerabilities, and it works without requiring any historical bug data.
The fundamental difference is the approach. SonarQube applies fixed rules derived from language specifications and known antipatterns; its findings are reproducible and explainable by the rule that fired. bugbug trains probabilistic classifiers on the actual history of bugs filed against a specific project; its predictions reflect the patterns in that project's specific bug record, not generic language rules.
For a new project or a project without much bug history, SonarQube is the practical choice because it needs no training data. For a large project like Firefox with decades of Bugzilla history, bugbug can generate predictions that reflect the specific failure modes of that codebase, such as which subsystem a new bug belongs to or which patches historically tended to be reverted. The two tools address overlapping but distinct problems and can be used together.
HTTP Service, Agents, and Licence
The repository includes an HTTP service (http_service/) that exposes the trained models as an API. The docker-compose.yml defines bugbug-http-service and bugbug-http-service-bg-worker, both reading BUGBUG_BUGZILLA_TOKEN and BUGBUG_GITHUB_TOKEN from the environment and exposing port 8000. A team can run the HTTP service alongside their Bugzilla instance to serve live classification results to integrations or dashboards.
The repository also contains an agents directory (agents/) with subprojects including autowebcompat-repro, autowebcompat-diagnosis, bug-fix, build-repair, test-repair, frontend-triage, and test-plan-generator, each with its own Docker Compose configuration. These suggest active expansion toward agentic software engineering workflows beyond classification.
The licence is MPL-2.0, which allows use in proprietary products as long as modifications to MPL-licensed files are released under the same licence. The last push was on 2026-09-28. The pyproject.toml requires Python 3.12 or higher and lists production dependencies including scikit-learn, langchain, langgraph, and several Anthropic and OpenAI SDK packages, indicating the project has grown beyond its original ML classifier scope to include LLM-based agents.
Editorial conclusion
bugbug is best suited for large open-source projects with an established Bugzilla workflow, a significant body of labeled bug history, and engineering capacity to operate a Python ML service. Teams on GitHub with few hundred bugs in their history will not have enough training data to get useful model accuracy. Before adopting it, verify that Python 3.12 and uv are available in your environment, confirm that the libgit2 v1.0.0 requirement can be satisfied (the README notes it is currently only in Debian's experimental repository), and assess whether running a full training cycle (30 minutes or more, per the README) fits your CI budget.
Frequently asked questions
Does bugbug work with GitHub issues, or only with Bugzilla?
The docker-compose.yml environment variables include both BUGBUG_BUGZILLA_TOKEN and BUGBUG_GITHUB_TOKEN, indicating the HTTP service supports both. However, the core classifiers and training documentation in the README are written for Bugzilla data. GitHub-specific support for the classifier training pipeline is not documented in the README.
Can I use bugbug's pre-trained models without training my own?
Yes. The README states that running the bug_classifier script without training a model first will automatically download an already-trained model. These pre-trained models are trained on Mozilla's own bug data, so their accuracy on a different project's bugs will depend on how similar that project's patterns are to Firefox's.
What is the minimum Python version required to run bugbug?
The README and pyproject.toml both require Python 3.12 or higher. The README notes that you can verify the exact version in use by checking pyproject.toml.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/mozilla-bugbug)