# BiRefNet: high-resolution dichotomous image segmentation in Python

> BiRefNet is the official implementation of a CAAI AIR 2024 paper on bilateral reference for high-resolution dichotomous segmentation, published under MIT. It is a research repository first: you install it with pip, load weights from Hugging Face, and run inference.py, and you accept that the README documents no rollback, no versioning scheme and no ONNX export.

**ZhengPeng7/BiRefNet** — [CAAI AIR'24] Bilateral Reference for High-Resolution Dichotomous Image Segmentation

- Repository: https://github.com/ZhengPeng7/BiRefNet
- Website: https://www.birefnet.cv
- Stars: 4,243 · Forks: 336
- Language: Python
- License: MIT
- Published: 2026-09-23 · Updated: 2026-09-23 · Language: en
- Canonical page: https://hysenlabs.com/projects/zhengpeng7-birefnet

## What dichotomous image segmentation means here, and who needs it

Dichotomous image segmentation asks a narrow question: which pixels belong to the foreground object and which do not, at high resolution, with fine structures preserved. The README frames the repository as the official implementation of "Bilateral Reference for High-Resolution Dichotomous Image Segmentation" (CAAI AIR 2024), and the GitHub topics list background-removal, camouflaged-object-detection, salient-object-detection and high-resolution-image-segmentation. Those topics describe the same model pointed at four different jobs.

The audience is therefore not the same as the audience for a general segmentation toolkit. If you want to cut a person out of a photo, detect a camouflaged animal in a cluttered frame, or produce a matte for compositing, the output you need is a high-resolution alpha-like mask, not a set of labelled boxes and class ids. BiRefNet is trained for exactly that binary output. If you need instance masks across dozens of categories, the repository does not target that problem at all, and the paper title says so.

One detail in the README is easy to miss and matters for planning: the project is asking for GPU resources. A note dated 2024-04-08 states the maintainers "need more GPU resources" to push performance further, especially toward general use and higher-resolution images. That is a research project stating its own compute constraint in public, and it tells you the weights you download today are the current checkpoint of an ongoing training effort rather than a frozen product.

## The bilateral reference mechanism and the inference data flow

The name of the method is the mechanism. Bilateral reference means the network uses two reference signals, one from the image itself and one from the gradient or boundary information, and the paper title is the only place in the README where this is stated. The repository does not include an architecture diagram in the text given, so the honest description is limited to what the code layout shows: models/ holds the network definitions, config.py holds the configuration, image_proc.py handles image preparation and post-processing, and inference.py is the entry point that ties them together.

The data flow at inference is conventional for this class of model. An image is loaded, resized to the resolution the chosen checkpoint was trained at, passed through the network to produce a segmentation map, and that map is resized back and applied to the original image. image_proc.py is where that preparation and post-processing live, and the README's own news entries confirm the post-processing path is a performance concern: a June 30, 2025 entry reports that refine_foreground was accelerated by roughly 8 times, to about 80ms on a 5090, using a GPU implementation of fast-fg-est contributed in issue #226 by a user named @lucasgblu. The fact that a single post-processing step needed an 8x speedup tells you the mask refinement, not the network forward pass, was a real bottleneck.

Resolution is a first-class parameter in this design, not an implementation detail. The README lists separate checkpoints trained at different resolutions: BiRefNet_dynamic trained on a dynamic range from 256x256 to 2304x2304, BiRefNet_HR and BiRefNet_HR-matting trained at 2048x2048. Choosing the wrong checkpoint for your input size is the most likely way to get disappointing masks, and the README does not provide a decision table for that choice.

## Installing BiRefNet and running your first inference

There is no pip package for BiRefNet in the README. Installation means cloning the repository and installing requirements.txt, which pins two versions you should read before you run anything:

```bash
pip install -r requirements.txt
```

The file requires torch>=2.5.0, torchvision, numpy<2, opencv-python, timm, scipy, scikit-image, kornia, einops, tqdm, prettytable, tabulate, ipykernel, huggingface-hub>0.25 and accelerate. The numpy<2 pin is the one that breaks environments: if your project already runs on numpy 2.x, installing these requirements will downgrade it or fail, and the README does not document a workaround. The huggingface-hub and accelerate entries exist because weights are pulled from the Hugging Face model repository at ZhengPeng7/BiRefNet.

Once the environment is in place, inference.py is the entry point:

```bash
python inference.py
```

According to the repository layout, inference.py reads its settings from config.py, so the image path, the model variant and the output location are configured there rather than passed as flags. The README does not document the individual config keys, so open config.py and read it before assuming a command-line flag exists. The expected result is a foreground mask written out for the input image, with the refinement step handled by image_proc.py.

For a first run, pick the checkpoint that matches your input. If you are feeding images of many different sizes, the README describes BiRefNet_dynamic as trained from 256x256 to 2304x2304 and notes it shows performance on any resolution. If your inputs are consistently around 2048x2048, BiRefNet_HR is the general-use checkpoint and BiRefNet_HR-matting is the matting variant. There is also a Hugging Face Space at ZhengPeng7/BiRefNet_demo and three Colab notebooks linked from the README (multiple-images inference, inference and evaluation, and box-guided segmentation) if you want to see the output before installing anything locally.

## Where BiRefNet fails or is the wrong tool

The repository has one release, v1, dated 2024-05-13. The last push was on 2026-09-02, so the code is being touched, but there is no semantic versioning, no changelog beyond the news list, and no documented upgrade path between checkpoints. If you deploy a pinned checkpoint and the maintainers later publish a better one under the same model name on Hugging Face, the README gives you no procedure for validating that swap. The news entry from Sep 23, 2025 is a concrete example: the attention implementation in the Swin transformer was replaced with PyTorch's official SDPA, and the README states the models are "exactly same" with lower memory cost. That is a good change, but it is also the kind of change that arrives without a version number.

The second limitation is hardware. The README's efficiency figures are tied to specific GPUs: 17 FPS at 1024x1024 with 3.45GB GPU memory on a single RTX 4090, and the refine_foreground timing measured on a 5090. Nothing in the README describes a CPU path, a quantized path, or a mobile deployment. If your target is a phone, a browser, or a CPU-only server, this repository does not address it.

The third is scope. This is a dichotomous segmenter. It produces foreground against background. It is not a class-aware instance segmentation model, and the README makes no claim to be one. If your pipeline needs to know that the mask is a person versus a car, you are adding a second model, and BiRefNet's contribution is only the mask.

## BiRefNet against rembg, RMBG and SAM-style segmenters

The related searches for this project are dominated by comparison queries: birefnet vs rembg, birefnet vs rmbg, birefnet vs rmbg 2.0, birefnet vs sam, birefnet vs sam3, birefnet vs inspyrenet, birefnet alternative. The README does not contain benchmark comparisons against any of those, so the only honest comparison is architectural and operational.

rembg is a packaged tool: you install it and run it on an image, and it bundles several background-removal models behind one interface. BiRefNet is a research repository. You clone it, install a requirements.txt that pins numpy<2 and torch>=2.5.0, edit config.py, and run inference.py. The difference is not accuracy, which the README cannot settle. The difference is that rembg gives you a stable command and BiRefNet gives you the current state of a paper's implementation, including the checkpoint variants BiRefNet_dynamic, BiRefNet_HR and BiRefNet_HR-matting that you must choose between yourself.

SAM and SAM-style models are promptable segmenters: they take a point, a box or a mask and return a segmentation. The README links a box-guided segmentation Colab, so BiRefNet has a guided mode, but its default output is dichotomous, foreground against background, trained for high resolution and fine structure. RMBG is a background-removal model family, which overlaps with BiRefNet's background-removal topic but not with its camouflaged-object-detection or salient-object-detection topics. If your problem is a camouflaged object in a high-resolution frame, the comparison to a general background remover is not the right one.

## Licence, maintenance cost and what upgrading actually involves

The repository is MIT licensed, and the README carries an MIT badge linking to the LICENSE file. MIT is permissive: it allows commercial use, modification and redistribution provided the copyright notice and permission notice are included. This is not legal advice, and the licence covers the code in this repository, not necessarily the model weights hosted on Hugging Face, which are a separate distribution channel the README links to but does not describe in licensing terms. If you plan to ship the weights commercially, read the model repository's own terms rather than assuming the MIT badge on the code applies to them.

Maintenance cost is the interesting part. Because there is no package on a registry, upgrading means pulling the repository and re-reading the news list, which is where breaking-ish changes are announced. The Sep 23, 2025 attention change is the template: same models, different implementation, lower memory. The Jun 30, 2025 refine_foreground change is the other template: a user-contributed GPU implementation that made a documented step roughly 8 times faster, merged through an issue thread rather than a release. Neither arrived with a version number.

Practically, that means your upgrade procedure is: pin the commit you deployed, diff the news list and the files you actually touch (config.py, image_proc.py, models/), and re-run your own evaluation set. The repository includes eval_existingOnes.py and an evaluation/ directory, which suggests the project has its own evaluation path you can reuse for that check.

## Conclusion

Adopt BiRefNet if you need high-resolution dichotomous segmentation or matting and you are willing to run Python with torch>=2.5.0 and numpy<2 yourself. Do not adopt it if you need a packaged CLI, a supported ONNX or mobile runtime, or a documented upgrade path: the README describes none of these, and the only release is v1 from 2024-05-13. Verify first that your environment satisfies numpy<2 and torch>=2.5.0, then pick the weight variant that matches your input resolution (BiRefNet_dynamic for mixed sizes, BiRefNet_HR or BiRefNet_HR-matting for 2048x2048), and check the model zoo section for the efficiency figures before you size your GPU.

## FAQ

### What is BiRefNet?

BiRefNet is the official implementation of the paper "Bilateral Reference for High-Resolution Dichotomous Image Segmentation" (CAAI AIR 2024), written in Python and released under MIT. It performs dichotomous image segmentation, meaning it separates foreground from background at high resolution, and the repository also lists background removal, camouflaged object detection and salient object detection among its topics.

### How do I use BiRefNet?

Clone the repository, install requirements.txt, and run inference.py, which takes its settings from config.py. You also need to choose a checkpoint: BiRefNet_dynamic for images of varying resolution, BiRefNet_HR for general use at 2048x2048, or BiRefNet_HR-matting for matting at that resolution.

### Is BiRefNet open source?

Yes. The repository is licensed under MIT, and the README links the LICENSE file with an MIT badge. Note that the model weights are distributed separately through Hugging Face, and the README does not describe the licensing of those weights.

### What is the best model for image segmentation?

The README does not rank BiRefNet against other segmentation models, so it gives no answer to this. What it does state is that BiRefNet targets high-resolution dichotomous segmentation, foreground against background, with separate checkpoints at 256x256 to 2304x2304 and 2048x2048.

### What is a good IoU score for segmentation?

The README does not state an IoU threshold or target. It links an inference and evaluation Colab notebook and the repository contains eval_existingOnes.py and an evaluation/ directory, so you can run the project's own evaluation path rather than relying on a general figure.

## Sources

- [License: MIT](https://github.com/ZhengPeng7/BiRefNet/blob/main/LICENSE)
- [Project website](https://www.birefnet.cv)
- [README](https://github.com/ZhengPeng7/BiRefNet/blob/main/README.md)
- [Releases](https://github.com/ZhengPeng7/BiRefNet/releases)
- [ZhengPeng7/BiRefNet on GitHub](https://github.com/ZhengPeng7/BiRefNet)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/zhengpeng7-birefnet
