# HivisionIDPhotos: ID photo generation on CPU, and where it stops being the right tool

> HivisionIDPhotos is a Python pipeline that turns an ordinary portrait into a sized, background-replaced ID photo using ONNX matting and face detection models. It runs offline on CPU, which is its main advantage and also the source of its sharpest limits.

**Zeyi-Lin/HivisionIDPhotos** — HivisionIDPhotos: a lightweight and efficient AI ID photos tools. AI .

- Repository: https://github.com/Zeyi-Lin/HivisionIDPhotos
- Website: https://modelscope.cn/studios/SwanLab/HivisionIDPhotos
- Stars: 21,572 · Forks: 2,505
- Language: Python
- License: Apache-2.0
- Published: 2026-08-08 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/zeyi-lin-hivisionidphotos

## The problem HivisionIDPhotos solves, and for whom

Producing a compliant ID photo is a small, repetitive task with a fixed shape: detect the face, cut the person out of the background, place them on a solid colour, crop to a specification, and optionally tile several copies onto a print sheet. HivisionIDPhotos packages that sequence as one Python pipeline, and the README frames the goal as a systematic algorithm for ID photo production rather than a single matting model.

The audience is narrow in a useful way. It suits engineers who need to generate ID photos in bulk or inside an application without sending portraits to a third-party service. The README states that matting is fully offline and runs on CPU alone, and that the project supports either fully offline or client-cloud inference. That matters for anyone handling identity documents, where uploading a face to an external API is a compliance question before it is an engineering one.

It is not a consumer photo editor. The output is a cropped, background-replaced portrait at a chosen specification, plus print layouts. The README also lists beauty retouching and, as a pending item, automatic formal-clothing replacement, so the current feature set is narrower than the roadmap suggests.

## How the pipeline is assembled: matting, face detection, layout

The architecture is a chain of ONNX models rather than a trained end-to-end network. Matting models live in hivision/creator/weights, and the README names four: MODNet, a variant called hivision_modnet tuned for solid background replacement, rmbg-1.4 from BRIA AI, and birefnet-v1-lite from BiRefNet. Face detection is a separate stage with its own directory, hivision/creator/retinaface/weights, holding a RetinaFace ONNX file when you choose that detector.

MTCNN is the default detector and is described as an offline CPU model with millisecond-level inference but lower detection accuracy. RetinaFace is the higher-accuracy offline option, described as second-level on CPU. A third path, Face++, delegates detection to an online API. So the accuracy and latency trade-off is exposed as a configuration choice rather than hidden inside the code.

The README publishes a performance table measured on a Mac M1 Max with 64GB of memory and no GPU acceleration, at image resolutions of 512x715 and 764x1146. MODNet with MTCNN uses 410MB and takes 0.207s and 0.246s. MODNet with RetinaFace uses 405MB and takes 0.571s and 0.971s. birefnet-v1-lite with RetinaFace uses 6.20GB and takes 7.063s and 7.128s. Those numbers are the clearest statement of the design trade-off in the repository: the best segmentation quality costs roughly thirty times the inference time and fifteen times the memory of the default pairing.

Around the models sits the application layer. app.py runs a Gradio demo, deploy_api.py exposes an HTTP service, and demo/ contains the UI, configuration and processor modules. The README mentions a beast mode in the Gradio demo that changes the memory loading strategy, which is the project's own answer to the memory profile of the heavier models.

## Installing HivisionIDPhotos and generating a first photo

The README requires Python 3.7 or newer and notes that the project is mainly tested on Python 3.10, on Linux, Windows and macOS. It recommends creating a conda environment first. Clone the repository and install both requirement files; the split matters because requirements-app.txt carries the demo and API dependencies that requirements.txt does not.

```bash
git clone https://github.com/Zeyi-Lin/HivisionIDPhotos.git
cd HivisionIDPhotos
pip install -r requirements.txt
pip install -r requirements-app.txt
```

Weights are not in the repository. The README gives a script as the first download method, with an option to fetch a single model by name.

```bash
python scripts/download_model.py --models all
# or a single model
python scripts/download_model.py --models modnet_photographic_portrait_matting
```

The second method is manual download from the release links, placing each file in hivision/creator/weights. Note that two of the four models must be renamed after download: the BRIA ONNX file becomes rmbg-1.4.onnx and the BiRefNet file becomes birefnet-v1-lite.onnx. Getting those names wrong is a silent failure mode, because the pipeline looks for a fixed filename.

With weights in place, the README's demo command starts the Gradio interface on port 7860.

```bash
python app.py --host 0.0.0.0 --port 7860
```

For a service rather than a UI, the README documents running deploy_api.py, and the docker-compose.yml in the repository maps that service to port 8080. The compose file defines two services from the same image, linzeyi/hivision_idphotos: one running app.py on 7860 and one running deploy_api.py on 8080.

```yaml
services:
  hivision_idphotos:
    image: linzeyi/hivision_idphotos
    command: python3 -u app.py --host 0.0.0.0 --port 7860
    ports:
      - '7860:7860'
  hivision_idphotos_api:
    image: linzeyi/hivision_idphotos
    command: python3 deploy_api.py
    ports:
      - '8080:8080'
```

The Dockerfile is based on python:3.10-slim, installs ffmpeg, libgl1-mesa-glx and libglib2.0-0, copies both requirement files, and exposes 7860 and 8080. It does not download model weights, so a container built from it still needs the weights mounted or fetched at runtime.

## Where HivisionIDPhotos breaks down

The default configuration is the weak point. MTCNN is fast and offline, but the README itself describes its detection accuracy as lower. For a portrait with a clear, frontal face this is fine. For a tilted head, heavy glasses glare, or a face partly in shadow, detection quality determines whether the rest of the pipeline has anything to work with, and the matting model cannot compensate for a bad crop. Moving to RetinaFace triples to quintuples CPU latency according to the published table, which is a real cost if you are processing a queue.

The quality ceiling is the second constraint. The README states that birefnet-v1-lite has the best segmentation accuracy of the four models, which implies the others are worse. Choosing it means 6.20GB of memory and about seven seconds per image on the reference machine. The README says GPU acceleration currently applies only to birefnet-v1-lite and asks for around 16GB of VRAM. That combination rules out small containers and most shared CI runners.

There is also a scope limit worth stating plainly. The README lists automatic formal-clothing replacement as waiting, and the supported outputs are ID photos and print layouts. If your requirement is a full retouching workflow with manual masking, this is the wrong layer to build on. And because the repository does not include weights, the licence of each downloaded model is a separate question from the Apache-2.0 licence on the code.

## Compared with calling a hosted matting API

The obvious alternative is a hosted segmentation or ID photo API: send the image, receive a cutout. The difference is where the model runs and who holds the portrait. A hosted API gives you a maintained model without a download step, no weights directory to populate, and no 6GB memory floor, and it usually scales without you provisioning anything.

HivisionIDPhotos inverts those properties. The README states that matting is fully offline and CPU-only, so a portrait never leaves the machine, and there is no per-image cost or rate limit. The price is that you own the model lifecycle: downloading weights, matching filenames, and choosing between the fast and accurate detectors yourself. It also means latency is yours to budget, and the published table shows that budget ranges from a fifth of a second to seven seconds depending on the model pair.

The README also documents a middle path. Face++ is a hosted detection API listed alongside the two offline detectors, so detection can be remote while matting stays local. That is a reasonable split if detection accuracy is the binding constraint and the portrait itself is not the sensitive part of the request.

## Maintenance, versions and licence

The repository is not archived, and the last push was on 2025-01-21, which is the same date as the v1.3.1 release. The two releases before it are v1.3.0 on 2024-11-16 and v1.2.9 on 2024-09-25, so the release cadence in the recorded history is roughly every one to two months up to that point. There is no published upgrade path in the README, and no migration notes between v1.2.9 and v1.3.1.

Upgrade cost is mostly the weights and the environment. The dependency list pins numpy to 1.26.4 or below, which will conflict with projects that require a newer numpy, and onnxruntime is the execution engine for every model. The Dockerfile pins the base image to python:3.10-slim, so container users inherit that Python version. Because weights are downloaded rather than vendored, a version bump that changes a model filename would require re-downloading and possibly renaming files.

The code is Apache-2.0. The weights are not covered by that licence: the README links to MODNet, BRIA AI's RMBG-1.4 and BiRefNet releases, each with its own terms. Anyone shipping this in a product should read those terms rather than assume the repository licence extends to them. That is a factual boundary, not legal advice.

## Conclusion

Adopt HivisionIDPhotos when you need offline ID photo generation on CPU and can accept the quality ceiling of the default MODNet plus MTCNN combination, which the README measures at 0.207s and 410MB on a Mac M1 Max. Do not adopt it if you need the birefnet-v1-lite segmentation model in a small container: the README lists 6.20GB of memory and roughly 7 seconds per image for that combination, and GPU acceleration for it asks for about 16GB of VRAM. Before committing, verify which model weights are actually present under hivision/creator/weights, since the repository does not ship them, and check whether the Apache-2.0 licence covers the third-party weights you download.

## FAQ

### How can I turn a photo into an ID photo with HivisionIDPhotos?

Install the dependencies, download the matting and face detection weights into hivision/creator/weights, then run the Gradio demo with python app.py or the HTTP service with python deploy_api.py. The pipeline detects the face, cuts out the person, replaces the background and crops to the specification you choose.

### What does an ID picture produced by HivisionIDPhotos look like?

The output is a portrait cropped to the ID specification you select, with the original background replaced by a solid colour, optionally with beauty retouching applied. The README also documents print layouts in five sheet sizes, including six-inch, five-inch, A4, 3R and 4R.

### Are AI photos made by HivisionIDPhotos safe?

The README states that matting is fully offline and runs on CPU alone, so the portrait does not have to leave your machine for that stage. The project also supports a client-cloud path and an online Face++ detection API, both of which do send data outward, so safety depends on which configuration you choose.

### Can I use AI to identify a photo with HivisionIDPhotos?

HivisionIDPhotos does not identify people or match faces. It runs face detection to locate a face in the frame, then uses matting models to separate the person from the background and produce an ID photo. Identification is outside what the repository documents.

## Sources

- [Official documentation](https://modelscope.cn/studios/SwanLab/HivisionIDPhotos)
- [Official README](https://github.com/Zeyi-Lin/HivisionIDPhotos#readme)
- [Project repository](https://github.com/Zeyi-Lin/HivisionIDPhotos)
- [Release notes](https://github.com/Zeyi-Lin/HivisionIDPhotos/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/zeyi-lin-hivisionidphotos
