Library / SDK
ShiqiYu/libfacedetection avatar
ShiqiYu/libfacedetection

ShiqiYu/libfacedetection: YuNet face detection that compiles without dependencies

An open source library for face detection in images. The face detection speed can reach 1000FPS.

12,806 stars3,031 forksC++NOASSERTION

At a glance

What is it?
libfacedetection ships a CNN face detector as plain C source files, so a C++ compiler is the only build requirement. It is fast on x86 and much slower on ARM, and the newest YuNet ONNX model needs a fixed input shape under OpenCV DNN.
Who is it for?
Adopt libfacedetection when you need a small, dependency-free face detector that runs on CPU and you control the build flags: copy src/ into your tree, compile with -O3 or /O2, and call the detection function from your own threads. Skip it if you need landmarks beyond the five points it returns, if you need GPU inference, or if you are targeting a Raspberry Pi 4 and expect real-time 640x480 detection, where the documented single-thread figure is 2.47 FPS.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 94 days ago.
What is it written in?
Mainly C++, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What libfacedetection solves, and who it is actually for

Most face detectors arrive as a framework dependency. You install PyTorch or TensorFlow, pull a model, and inherit a runtime that may be larger than your application. libfacedetection takes the opposite route: the CNN model is converted into static variables inside C source files, and the README states the source code does not depend on any other libraries. What you need is a C++ compiler. That single sentence defines the audience. This is for embedded and native developers who want face detection inside a C or C++ program without shipping an inference engine, and for people building on Windows, Linux or ARM where adding a Python stack is unwelcome. It returns bounding boxes for faces. The repository describes it as a detector, not a recognizer, so it does not tell you who a person is, only where faces are. The README also notes a minimal face size of roughly 10x10 pixels, which sets a practical floor on how far you can downscale an image before small faces disappear from the output.

How the CNN is embedded and what the detection call returns

The mechanism is unusual and worth understanding before you adopt it. The trained network is serialized into C++ arrays in src/facedetectcnn-data.cpp. There is no model file to load at runtime and no deserialization step, which removes a class of deployment failures around missing or mismatched model paths. Detection runs through the functions declared in facedetectcnn.h, and the examples directory contains detect-image.cpp and detect-camera.cpp showing the intended call pattern. Speed comes from SIMD: the README says you can enable AVX2 on Intel CPUs or NEON on ARM. The published tables assume those paths are active, so a build without them is not the build being measured. There is a second front end. The same network is available as an ONNX model from OpenCV Zoo, and the repository provides scripts under opencv_dnn/ in both C++ and Python for running it through OpenCV DNN. That route trades the zero-dependency property for OpenCV, and it comes with a documented constraint covered below. The library was trained by the separate libfacedetection.train repository, so the training pipeline and the inference library live in different places.

Installing libfacedetection and running a first detection

There is no package manager step and no install command in the README. The integration model is source copying: you copy the files in src/ into your project and compile them alongside your own files. The README gives one setup detail that trips people up, pointing to issue #222: you must place facedetection_export.h where you copied facedetectcnn.h, and define FACEDETECTION_EXPORT in that header. The repository also ships COMPILE.md for building the code as a static or dynamic library if you prefer that over copying sources. On the compiler side, the README is explicit: add -O3 with g++, or choose Maximize Speed /O2 in Microsoft Visual Studio. Building without optimization gives you a debug build of a CNN, which is not a meaningful measurement of the library. To see it work before integrating anything, the README points to examples/detect-image.cpp and examples/detect-camera.cpp as the demonstration of how to use the library. The README also documents a second entry point, the Highway-based C API declared in facedetect_hw.h, whose README example reads:

C++
#include "facedetect_hw.h"

int* results = facedetect_hw_cnn(result_buffer, bgr_image_data,
                                 width, height, step);

That call returns a pointer into the buffer you supply, so the buffer lifetime is yours to manage. For the ONNX path, the repository keeps its own scripts under opencv_dnn/, and the README warns that OpenCV DNN does not support the latest YuNet with dynamic input shape. If you use that route, your input must match the shape baked into the ONNX file exactly. The README also suggests enabling OpenMP for speed but adds that the better solution is calling the detection function from different threads, which is the same advice the Highway variant repeats for its own API.

The speed claim depends on resolution, threads and instruction set

The headline number is 1000 FPS, and the tables show where it comes from. On an Intel i7-7820X with AVX2 and 16 threads, 640x480 runs at 152.65 FPS, 320x240 at 550.54 FPS, 160x120 at 1745.13 FPS and 128x96 at 2994.23 FPS. The four-digit figure belongs to very small frames. Single-threaded AVX2 at 640x480 is 19.99 FPS. The ARM table tells a sharper story: on a Raspberry Pi 4 B with 4 threads, 640x480 reaches 7.97 FPS and 320x240 reaches 30.32 FPS. If your plan is full-resolution detection on a Pi, the documented numbers do not support it. Accuracy is reported on WIDER Face at default settings with scales=[1.], confidence_threshold=0.02: AP_easy=0.887, AP_medium=0.871, AP_hard=0.768. The hard subset is where crowded and small faces live, and 0.768 is the honest weak point of the model. Note that these figures are the project's own published measurements on its own hardware, not independent verification.

The Highway variant is a separate build, not a runtime switch

An independent Highway-based implementation sits under highway/ and leaves the original code untouched, exposing its own C API through facedetect_hw.h. It follows the same deployment model as the rest of the project: build separately for each instruction set and platform. The README states plainly that it does not currently use Highway runtime dynamic dispatch, so you do not get one binary that adapts to the CPU it lands on. The current x86 path is a hybrid, with Highway kernels for pointwise work and guarded AVX2/FMA intrinsics for selected depthwise and maxpool kernels. Its API uses thread-local internal workspaces, so the recommended model is external parallelism: call facedetect_hw_cnn from multiple threads, one result buffer per calling thread. That is a real constraint on how you structure a server, because the buffer lifetime is your responsibility. If you want a single portable binary that picks the best instruction set at startup, this is not that, and the README does not claim otherwise.

Where libfacedetection is the wrong choice

Two boundaries matter most. First, it detects faces; it does not identify people. A team that needs recognition, embeddings or matching has to add a separate model, and the README offers nothing for that. Second, the output is boxes plus the small set of landmarks the model produces, not a dense facial mesh. If your pipeline needs eye, nose and mouth geometry for alignment or expression work, the detector alone will not carry it. There is also a maintenance reality to weigh: the most recent tagged release is v3.0 from 2021-09-24, and the last push to the repository was on 2026-06-28. The repository is not archived and work is still landing on the branch, but the release tags are old, so anyone pinning to a version number is pinning to something from 2021. The ONNX note adds a compatibility trap: the latest YuNet uses dynamic input shape and OpenCV DNN does not support that, so a pipeline that resizes frames freely before inference will fail or misbehave on that path. Finally, the licence file is present but the repository metadata reports the licence as NOASSERTION, which means the platform could not classify it automatically. Read LICENSE yourself and decide with your own counsel; nothing here settles that question.

How it compares to OpenCV's own detector and to ONNX Runtime

The closest alternative for most readers is OpenCV's built-in cascade classifier, which ships with OpenCV and needs no extra sources. The difference is in the model: a Haar cascade is a hand-designed feature classifier, while libfacedetection embeds a CNN, and the project's WIDER Face numbers reflect that gap on harder images. The trade is dependency weight. If OpenCV is already in your build, the cascade costs you nothing new; libfacedetection costs you a source copy and a compile flag. The second alternative is running YuNet through ONNX Runtime instead of OpenCV DNN. Both consume the same ONNX model from OpenCV Zoo, but ONNX Runtime handles dynamic input shapes, which is exactly the limitation the README flags for the OpenCV DNN path. Choosing ONNX Runtime means accepting a heavier runtime dependency in exchange for shape flexibility and a cleaner model upgrade path. If your priority is a small binary on a fixed platform with a known frame size, the embedded C arrays remain the leanest option of the three.

Maintenance, upgrades and what the licence question costs you

The upgrade surface is small, which cuts both ways. Because the model lives in source files, upgrading means replacing source files rather than swapping a model artifact, and there is no version negotiation at runtime to go wrong. The cost is that you cannot hot-swap models without a rebuild, and every model change is a code change in your tree. With v3.0 tagged in September 2021 and the last push on 2026-06-28, expect to track the branch rather than wait for a release if you want current code. The Highway directory is the clearest example: it is newer work that is not covered by any of the release tags. On licensing, the repository contains a LICENSE file but the platform reports NOASSERTION, so the classification is unresolved at the metadata level. That is a signal to open the file, not a conclusion about what it permits. If you are embedding the source in a commercial product, the terms in that file are the ones that apply, and this article does not interpret them.

Editorial conclusion

Adopt libfacedetection when you need a small, dependency-free face detector that runs on CPU and you control the build flags: copy src/ into your tree, compile with -O3 or /O2, and call the detection function from your own threads. Skip it if you need landmarks beyond the five points it returns, if you need GPU inference, or if you are targeting a Raspberry Pi 4 and expect real-time 640x480 detection, where the documented single-thread figure is 2.47 FPS. Before committing, check three things in your own tree: whether facedetection_export.h is in place next to facedetectcnn.h with FACEDETECTION_EXPORT defined, whether your ONNX input shape matches the model exactly if you go through OpenCV DNN, and whether your target CPU supports AVX2 or NEON, since the published speed tables assume those instructions are enabled.

Frequently asked questions

How accurate is YuNet in libfacedetection?

The README reports WIDER Face results at default settings with scales=[1.] and confidence_threshold=0.02: AP_easy=0.887, AP_medium=0.871, AP_hard=0.768. The hard subset, which covers crowded and small faces, is the weakest of the three. These are the project's own published numbers.

How accurate is face detection with libfacedetection?

Accuracy is published only as the WIDER Face AP figures above, with no separate per-image or per-dataset breakdown in the README. The README also states a minimal detectable face size of roughly 10x10 pixels, which effectively sets the accuracy floor for downscaled images.

Can you give an example of face detection with libfacedetection?

The repository ships examples/detect-image.cpp and examples/detect-camera.cpp, which the README names as the demonstration of how to use the library. The ONNX route has its own C++ and Python scripts under opencv_dnn/.

How does facial recognition know who you are, and does libfacedetection do that?

libfacedetection only locates faces in an image; the README describes it as a face detection library and offers nothing for identifying individuals. Recognition would require a separate model that the repository does not provide.

Official sources

  1. Issues
  2. README
  3. Releases
  4. ShiqiYu/libfacedetection on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/shiqiyu-libfacedetection.svg)](https://hysenlabs.com/projects/shiqiyu-libfacedetection)