Library / SDK
obss/sahi avatar
obss/sahi

SAHI: find the small objects your detector walks past

Framework agnostic sliced/tiled inference + interactive ui + error analysis plots

5,528 stars789 forksPythonMIT

At a glance

What is it?
SAHI, Slicing Aided Hyper Inference, is an MIT-licensed Python library for large-scale object detection and instance segmentation that detects small objects in large images by slicing, framework-agnostic inference and result merging. It supports ultralytics, mmdetection, Hugging Face and torchvision models through seven CLI commands, integrates the Fiftyone app for interactive review, and counts more than 600 citing publications.
Who is it for?
Use SAHI when your detection problem is small objects in large imagery, aerial, medical or high-resolution photography, and you want the slicing, inference and merging handled by a maintained library instead of hand-rolled around your detector. Skip it when objects already fill a healthy fraction of the frame, where slicing adds cost without recall.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 3 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Slicing as a recall strategy

SAHI, which expands to Slicing Aided Hyper Inference, is a lightweight vision library for large scale object detection and instance segmentation, and its single organizing idea is that detectors trained on ordinary datasets miss small objects in large images because the objects occupy too few pixels of the input. The remedy is in the name, slice the image, run inference over the slices so small objects become proportionally large, and merge the results. The library is framework agnostic, supporting popular detection models from ultralytics, mmdetection, Hugging Face and torchvision through one API, and the underlying method is published in an IEEE paper linked from the repository. The community evidence is unusual in being enumerated, more than 600 publications cite SAHI, and a discussions thread lists competition winners who used it, which for a technique library is the strongest adoption signal available.

Seven verbs from predict to coco yolo

The command line interface organizes the workflow into seven verbs. predict performs sliced or standard prediction on video and images using any supported model, the general-purpose entry point. predict-fiftyone runs the same prediction and opens results in the Fiftyone app for interactive exploration. The coco family handles datasets, coco slice automatically slices COCO annotation and image files, coco fiftyone explores multiple prediction results ordered by number of misdetections, coco evaluate computes classwise COCO AP and AR against ground truth, coco analyse calculates and exports error analysis plots, and coco yolo converts any COCO dataset to the ultralytics format. The verbs sketch a complete evaluation loop, slice the data, predict, inspect the worst failures visually, quantify, analyze by error type, and hand off to training, without leaving the CLI.

pip install sahi and modest dependencies

Installation is the friendliest possible line:

bash
pip install sahi

The package requires Python 3.8 or newer and its dependencies are deliberately modest, OpenCV, matplotlib, shapely for geometry, tqdm, pillow, numpy, pyyaml, requests, and the fire and click CLI frameworks, with no deep learning framework among them, because the models arrive from whichever detection library the user brings. A conda-forge package exists alongside PyPI, and the dependency design means SAHI layers onto an existing detection stack rather than competing with it. The sahi command installs through the standard entry point, and releases are current, 0.12.5 in August 2026, 0.12.6 two weeks later and 0.12.7 in September, with the repository pushed on 2026-09-17.

A notebook per framework, including the new ones

The demo directory is the practical compatibility matrix, thirteen notebooks each wiring SAHI to one detection stack, ultralytics and its newer YOLOE variant, YOLOv5, torchvision, mmdetection, Hugging Face, Detectron2, Grounding DINO, RT-DETR and Roboflow, plus the technique-specific notebooks for slicing itself and batch slicing. The spread tells two stories, first that the framework-agnostic claim is exercised rather than asserted, and second that the project tracks the detection ecosystem's frontier, with Grounding DINO and RT-DETR representing the open-vocabulary and transformer ends of current practice. For a newcomer, finding a notebook for their exact stack converts evaluation from reading documentation to running a cell, and the batch slicing notebook addresses the throughput question everyone asks second.

Fiftyone for the part that is usually skipped

Object detection work fails most often not at inference but at the unexamined middle, where predictions are glanced at rather than analyzed. SAHI wires the Fiftyone app into that gap twice, predict-fiftyone for interactive exploration of a run's results, and coco fiftyone for comparing multiple prediction results on a COCO dataset with entries ordered by number of misdetections, so review effort concentrates where the failures are. Behind the visualization sits the quantitative half, coco evaluate for classwise AP and AR, and coco analyse for exported error analysis plots, the confusion patterns and failure taxonomies that turn a weak detector into a better training plan. Building the inspection loop into the same tool as the slicing is the difference between a library that improves a metric and one that improves the process.

Documentation aimed at AI assistants

The README carries a section titled Approved by AI Tools, and its content is concrete rather than aspirational, SAHI's documentation is indexed in Context7 MCP, giving AI coding assistants up-to-date, version-specific code examples and API references, and the project publishes an llms.txt following the emerging standard for machine-readable documentation, with an installation guide for wiring the docs into an agent workflow. A DeepWiki rendering and a Hugging Face Spaces demo running the YOLOX-backed pipeline round out the machine-facing surface. This is a small but telling investment, libraries whose examples can be retrieved correctly by coding agents get adopted inside those agents' outputs, and SAHI is notably early in treating that as a first-class documentation channel alongside humans.

Trilingual docs and a maintained core

The documentation effort extends across languages, with the README available in English, Simplified Chinese and Turkish, and the zensical configuration files for the docs site carrying the same trilingual commitment. The project is maintained by Fatih Cagatay Akyon and Onuralp Sezer under the obss organization, carries a CITATION.cff for academic use alongside the paper, and ships a code of conduct, security policy and pre-commit configuration. A markdown lint configuration keeps the multilingual docs tidy, and the CHANGELOG tracks the steady 0.12.x line. The honest framing for adopters is that SAHI is the maintained, cited implementation of a technique you could build yourself, slicing and merging logic over any detector, and the value purchased is not the idea but the edge cases, the merging, the evaluation tooling and the framework adapters that already survived hundreds of other teams' datasets.

Editorial conclusion

Use SAHI when your detection problem is small objects in large imagery, aerial, medical or high-resolution photography, and you want the slicing, inference and merging handled by a maintained library instead of hand-rolled around your detector. Skip it when objects already fill a healthy fraction of the frame, where slicing adds cost without recall. Verify first that your model framework is among the supported ones or covered by a demo notebook, budget for the extra inference passes slicing implies, and use coco evaluate before and after to confirm the recall gain on your own data rather than trusting the technique in the abstract.

Frequently asked questions

What is SAHI in object detection?

SAHI, Slicing Aided Hyper Inference, is an MIT-licensed Python library that improves detection of small objects in large images by slicing images, running framework-agnostic inference and merging results. It works with ultralytics, mmdetection, Hugging Face and torchvision models, and more than 600 publications cite it.

What is slicing aided hyper inference?

It is the technique of dividing a large image into slices, running a detector on each so small objects occupy proportionally more of the input, then stitching detections back together. SAHI implements it as a library and CLI, and the method is described in an IEEE publication linked from the repository.

Which detection frameworks does SAHI support?

It is framework agnostic: the predict command works with ultralytics, mmdetection, Hugging Face and torchvision models, and the demo notebooks additionally cover Detectron2, Grounding DINO, RT-DETR, Roboflow, YOLOv5 and YOLOE variants.

Official sources

  1. License: MIT
  2. obss/sahi on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/obss-sahi.svg)](https://hysenlabs.com/projects/obss-sahi)