Open-source project
rafaelpadilla/Object-Detection-Metrics avatar
rafaelpadilla/Object-Detection-Metrics

Object-Detection-Metrics: PASCAL VOC mAP Without XML or JSON Conversion

Most popular metrics used to evaluate object detection algorithms.

5,104 stars1,027 forksPythonMIT

At a glance

What is it?
The repository that most papers cite for Precision x Recall curves and Average Precision in Python. It reads plain text files of bounding boxes, ships two sample folders, and has not been pushed since 2026-09-01.
Who is it for?
Adopt it when you need PASCAL VOC style Average Precision and your predictions already live in text files, and you want the metric definition to be readable rather than hidden inside a framework. Do not adopt it for COCO's twelve-metric report, for segmentation masks, or for video, because the README points those users to the successor tool.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 29 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Problem Object-Detection-Metrics Solves, and For Whom

The README opens with the motivation directly: there is no consensus on how object detection is scored, and researchers working outside a competition dataset end up writing their own metric code. When two papers implement Average Precision slightly differently, the resulting numbers are not comparable, and the README states that a wrong or different implementation can create biased results. This project packages the metrics used by PASCAL VOC, COCO, Open Images and ImageNet's localization challenge into one set of functions so that everyone computes the same thing.

The audience is narrower than the topic suggests. It is aimed at researchers who have detections and ground truth as bounding box coordinates and want a number without adopting a training framework. The README is explicit that the implementation avoids conversions to XML or JSON, which is the friction that pushes people toward the evaluation code bundled inside TensorFlow's object detection models or inside py-faster-rcnn. If you already train inside a framework that reports mAP for you, this project adds little. If you have a detector written in something else, or you want to check a framework's reported number against an independent implementation, it is the piece that was missing.

How the Precision x Recall Curve and Average Precision Are Computed

The mechanism is described in the README in three parts. First, matching: each detection is compared against ground truth boxes of the same class using Intersection Over Union, a measure the README derives from the Jaccard Index. A detection counts as a true positive when its IOU with a ground truth box exceeds the threshold, and the README's definitions section covers the standard 0.5 convention that PASCAL VOC uses. Second, ranking: detections are ordered by confidence score, and the curve is built by walking down that ranking and recomputing precision and recall at each step. Third, summarization: Average Precision is the area under that curve, and the README documents both interpolation schemes, the 11-point method used by older PASCAL VOC years and the interpolation over all points that later versions adopted.

That third point is where most disagreements between papers originate, and the repository treats it as a first-class choice rather than a hidden default. The data flow is file-based in both directions: ground truth boxes come from one directory, detections from another, and the output is written under results/. The top-level layout reflects this, with groundtruths/, detections/, and a parallel pair of groundtruths_rel/ and detections_rel/ directories, plus samples/sample_1/ and samples/sample_2/ as worked examples. Nothing in the pipeline requires touching your model, which is the design claim the README makes when it says no modifications to complicated input formats are needed.

Installing Object-Detection-Metrics and Running a First Evaluation

There is no packaged release to install. The repository ships a requirements.txt pinned to specific versions, including numpy==1.22.0, matplotlib==3.1.3, opencv-python==4.2.0.32 and PyQt5==5.12.3. Those pins date from the era of the v0.2 release and include PyQt5, which is a heavy dependency for what is fundamentally a metric script. Installing them into an existing project can downgrade numpy, so isolate the environment first.

bash
pip install -r requirements.txt

The evaluation entry point visible in the repository root is pascalvoc.py, alongside _init_paths.py and the lib/ package that holds the implementation. The README's how-to-use section walks through running it against the bundled samples, and the two sample folders exist precisely so you can confirm the tool works before pointing it at your own data. The expected outcome is a printed Average Precision per class plus a Precision x Recall plot, with artifacts landing in results/.

bash
python pascalvoc.py

The script name above is the file that appears in the repository root. The README does not spell out each flag, so consult the script itself for the arguments it accepts. Once the samples produce a plot, replace the groundtruths/ and detections/ directories with your own files and rerun. If a class you expect is missing from the output, the usual cause is that no ground truth file in your directory mentions that class.

Where Object-Detection-Metrics Stops Being the Right Tool

The README carries its own warning. A banner near the top states that a new version is available at rafaelpadilla/review_object_detection_metrics, and lists what that version adds: all COCO metrics, other file formats, a user interface to guide evaluation, and the STT-AP metric for object detection in videos. Read that as a boundary. This repository is the PASCAL VOC lineage, and the twelve COCO metrics, which vary IOU thresholds and object sizes, are not what it computes. If your paper reports COCO AP, you are in the wrong repository.

The second limitation is the dependency surface. PyQt5 and PyQtWebEngine appear in requirements.txt, which means a headless CI machine needs either those wheels or a trimmed install. The pins themselves are a maintenance signal: numpy==1.22.0 and opencv-python==4.2.0.32 will conflict with newer stacks, and the release history shows v0.1 in 2018 and v0.2 in 2019, so the versioned releases are old even though the repository has received commits since. The last push was on 2026-09-01, which is recent, but a recent push is not the same as a supported dependency matrix.

Third, the metric itself is unforgiving in ways the tool cannot fix. Average Precision at a single IOU threshold says nothing about localization quality at stricter thresholds, and a model that is strong at 0.5 can look weak at 0.75. Reporting one number from one threshold and comparing it to a paper that used another is the exact failure mode the README was written to prevent, and no script can detect that you did it.

The Alternative Inside the Same Project: review_object_detection_metrics

The honest comparison is not against a competitor but against the successor the README points to. review_object_detection_metrics is a separate repository by the same author, and the README describes the difference in approach: it computes all COCO metrics rather than the PASCAL VOC set, it accepts other file formats instead of the plain text layout used here, it adds a UI that walks through the evaluation, and it introduces STT-AP for video. That is a different scope, not a faster version of the same thing.

The trade-off is legibility. This repository is small enough that a reader can follow the precision and recall computation from the entry script into lib/ and check the interpolation against the paper the README cites, a 2021 Electronics article titled A Comparative Analysis of Object Detection Metrics with a Companion Open-Source Toolkit, and the earlier 2020 IWSSIP survey. The successor covers more ground and therefore hides more. If your goal is to report a COCO-style number, use the successor. If your goal is to understand or audit how a PASCAL VOC number was produced, this repository is the shorter path, and the two can coexist in the same environment since they are separate checkouts.

Licence, Maintenance and the Cost of Upgrading

The licence is MIT, declared in the LICENSE file at the repository root and shown by the badge at the top of the README. For a metrics library that gets vendored into research code, MIT is the permissive case: you can copy the functions into your own evaluation script or ship them inside a larger project. The one obligation MIT imposes is preserving the copyright notice and permission text, so if you paste code out of lib/ into your own file, keep the notice with it. This is a description of the licence terms, not legal advice; read LICENSE yourself if the distinction matters to your organisation.

The README also asks for citation if you use the code in research, and provides BibTeX entries for both the 2021 Electronics paper and the 2020 IWSSIP survey. That is a request rather than a licence condition, but it is the convention in this field and reviewers do check.

Upgrade cost is the practical concern. There is no package on an index to pin, so adoption means vendoring or a git submodule, and the pinned requirements become your problem. Moving from this repository to review_object_detection_metrics is not a version bump; it is a rewrite of your evaluation step, because the input format and the metric set both change. Budget for that as a separate task rather than assuming a drop-in replacement.

Editorial conclusion

Adopt it when you need PASCAL VOC style Average Precision and your predictions already live in text files, and you want the metric definition to be readable rather than hidden inside a framework. Do not adopt it for COCO's twelve-metric report, for segmentation masks, or for video, because the README points those users to the successor tool. Before trusting a number, verify which interpolation variant your comparison baseline used, confirm the IOU threshold you intend to report, and check that your ground truth file lists every class you care about, since an absent class produces no curve rather than an error.

Frequently asked questions

What is a good IoU score in Object-Detection-Metrics?

The README describes IOU as a measure based on the Jaccard Index and covers the standard 0.5 threshold convention used by PASCAL VOC, where a detection counts as a true positive when its overlap with a ground truth box exceeds the threshold. It does not define a universal good value, because the threshold is a reporting choice you make, not a property of the detector.

Does Object-Detection-Metrics compute COCO metrics?

No. The README states that the successor project, review_object_detection_metrics, is the one that includes all COCO metrics, along with other file formats, a user interface and the STT-AP metric for video. This repository covers the PASCAL VOC style Precision x Recall curve and Average Precision.

What does an accuracy of 1.0 mean in Object-Detection-Metrics?

The README does not document an accuracy field with that range, so a value of 1.0 has no defined meaning here. The reported quantities are precision and recall at a given IOU threshold, summarized as Average Precision, and a perfect score would require every detection to match a ground truth box while missing none.

Can I use Object-Detection-Metrics without converting my results to XML or JSON?

Yes, that is the stated design goal. The README says the implementation does not require modifications to complicated input formats and avoids conversions to XML or JSON files, using simplified ground truth and detected bounding box inputs instead. The repository layout shows plain groundtruths/ and detections/ directories.

Which interpolation should I use for Average Precision in Object-Detection-Metrics?

The README documents both the 11-point interpolation used by older PASCAL VOC years and the interpolation over all points adopted later, and the choice changes the resulting number. Pick the one your comparison baseline used, since mixing them is the source of the disagreement the project was created to reduce.

Official sources

  1. Issues
  2. License: MIT
  3. rafaelpadilla/Object-Detection-Metrics on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/rafaelpadilla-object-detection-metrics.svg)](https://hysenlabs.com/projects/rafaelpadilla-object-detection-metrics)