Self-hosted service
yihong1120/Construction-Hazard-Detection avatar
yihong1120/Construction-Hazard-Detection

Construction-Hazard-Detection runs one stream process per camera over shared GPU workers

Enhances construction site safety using YOLO for object detection, identifying hazards like workers without helmets or safety vests, and proximity to machinery or vehicles. HDBSCAN clusters safety cone coordinates to create monitored zones. Post-processing algorithms improve detection accuracy.

350 stars46 forksPythonAGPL-3.0

At a glance

What is it?
A live-camera safety monitor that flags missing hard hats, missing vests, workers too near machinery, workers inside cone-derived zones and vehicles near utility poles. Redis holds keys, MediaMTX holds the video, and the production database schema is not in the repository.
Who is it for?
This fits a site operator who already has cameras feeding RTSP and wants a warning on a phone when someone walks into a controlled area. It does not fit a quick local trial on a laptop, because the supported Python range is narrow, the model weights have to be fetched, and the production schema comes from a private deployment process rather than the repository.
Can I use it commercially?
Yes, with strict conditions. AGPL-3.0 is a network copyleft licence: if people use a modified version over a network, for example as a hosted service, you must offer them its source code under the same licence.
Is it still maintained?
Yes. The repository last received commits 14 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Eleven detection classes, five warning rules

The hazard list is specific rather than generic. Workers without hard hats, workers without safety vests, and workers standing too close to machinery or vehicles are all separate findings. Two more come from geometry: a worker inside a controlled area derived from safety cone coordinates, and machinery or vehicles too close to utility poles.

The cone areas are the part that is not straight object detection. Cone positions are clustered with HDBSCAN, and the resulting clusters become the monitored zones that the third rule is measured against, so the geometry of the site enters the decision rather than only the detected boxes. A fourth mechanism is post-processing on top of the raw detector output, described as improving accuracy.

The runtime diagram summarises eleven configured detection classes alongside the safety-warning rules and the video and metadata outputs. Labels and notifications are available in Traditional Chinese, Simplified Chinese, English, French, Thai, Vietnamese, Indonesian and Japanese.

Redis holds the keys, MediaMTX holds the video

The architecture draws one line firmly. Redis is not used to store live video frames. Live playback belongs to MediaMTX, which serves RTSP ingest and handles HLS and WebRTC output.

What Redis does hold is small coordination state, and the list is enumerated: an authentication cache, an FCM token cache, compact warning metadata, overlay demand keys and overlay ready keys. Those last two are how a client asks for an overlay on a stream and is then told when that overlay is ready to draw.

Around those two sit the FastAPI services, each with its own port in the example environment: the data management API on 8005, the notification API on 8003, the violation record API on 8002 and the streaming API on 8800. PostgreSQL holds the durable side, namely sites, users, stream configurations and violations, and the project asks for database mode unless you are testing a short-lived local JSON config.

One process per camera, workers sharing the GPU

The supervisor is main.py, and it polls the stream configuration and starts one stream process per active camera. Inference itself is not per process. Local YOLO workers share GPU inference across cameras through shared memory, which is why the capacity numbers are expressed per engine rather than per camera.

The default is three cameras per engine, and it can be overridden by model size:

dotenv
YOLO_WORKER_CAMERAS_PER_ENGINE=3
# Override capacity by model size; unlisted model keys use the default above.
YOLO_WORKER_CAMERAS_PER_ENGINE_BY_MODEL=yolo26n=8,yolo26s=6,yolo26m=4,yolo26l=2

So a nano model carries eight cameras and a large one carries two, and any model key missing from that list falls back to three.

There is also a standalone YOLO API, kept for API testing or separate deployment, but the documented high-throughput local path is the worker mode in src/yolo_worker.py. Choosing the standalone server means paying the inference cost per caller instead of sharing it.

Python 3.14 only, and extras decide which services exist

The supported range is narrow, which is the first thing that breaks an existing environment:

bash
uv sync --locked --all-extras --no-extra yolo-gpu

That installs everything except the GPU extra, because the default dependency set holds only the shared API runtime packages. Docker images install their own extras instead, whichever of streaming, violation, notification or yolo the image needs, and the yolo-gpu extra is reserved for hosts that actually have CUDA and TensorRT.

The optional notification channels for HTTP broadcast, Messenger and WeChat Work come from a separate extra:

bash
uv sync --extra social-notifications

That one is documented for use without the MCP server. Everything in the dependency set is pinned to an exact version rather than a range, from fastapi 0.139.2 and redis 7.1.0 through torch 2.13.0 and ultralytics 8.4.115 in the yolo extra, with tensorrt 11.1.0.106 in the GPU one.

Weights arrive from Hugging Face under a naming convention

The models are not in the repository. They are fetched from the Hugging Face model repository:

bash
hf download yihong1120/Construction-Hazard-Detection \
  --repo-type model \
  --include "models/pt/*.pt" \
  --local-dir .

The filename is not arbitrary. Worker models are expected at models/pt/best_<model_key>.pt, with best_yolo26n.pt given as the example, so the model key in the filename is what selects a weights file, and the capacity overrides are keyed on the same names.

The keys the detection server knows about are listed separately in the environment file as yolo26n, yolo26s, yolo26m, yolo26l and yolo26x. Authentication for the model endpoint is a bearer token that has to belong to a non-guest account allowed to call /get_new_model, and the transfer ceiling is 6 GiB, matched by the upload ceiling for the model path. Ingress is bounded before bytes reach the inference semaphore, with a 20 MiB upload limit and a concurrency of 8.

Compose gates Redis on a password and keeps HLS lazy

Infrastructure is three services on a custom bridge network:

bash
docker compose up -d redis postgres media-server

Redis runs from redis:7.4.10-alpine3.21 and refuses to start without REDIS_PASSWORD set, because compose substitutes the password into redis-server --requirepass. The port is published on 127.0.0.1 only, and a healthcheck pings with redis-cli every ten seconds with three retries.

MediaMTX runs from bluenviron/mediamtx:1.18.2 with restart unless-stopped, so playback survives a Docker daemon restart. RTSP is forced onto tcp rather than UDP. HLS and WebRTC are both on, but HLS multiplexers are created only for active viewers, which the compose file explains as avoiding unnecessary muxers per camera and keeping transient segments out of a shared volume. The defaults that shape playback are a lowLatency variant, 2 second segments, 14 segments held, 200ms parts and a muxer closed after 60 seconds of inactivity.

The image layering is sequential: base builds from scripts/base.Dockerfile, base-gpu from scripts/base-gpu.Dockerfile and waits for base to complete, and the runtime Dockerfile starts from base-gpu because the stream processor is the part that needs TensorRT.

The production schema lives outside the repository

The environment file states the constraint plainly. The production schema is managed through a private deployment process, and AUTO_CREATE_SCHEMA ships as false. Turning it on is described as appropriate only for throw-away local databases created from ORM metadata, which is a fair warning given that enabling it on a real database would derive tables from the models rather than apply a managed migration.

Connection settings are sized for several processes rather than one. A pool size of 2 with an overflow of 1 and a recycle interval of 1800 seconds is shared by all Uvicorn workers, with a 10 second pool timeout, which is a small pool on purpose when several API processes point at the same database.

Two smaller decisions are worth noting. Multi-device login can be forbidden per deployment, and the websocket path to the detection server is configured with a single connection attempt, a single retry and a 30 second failure cooldown, which means a down detector produces retries on a slow cadence rather than a reconnect loop. The package version in pyproject is 0.0.0 even though releases are tagged v5.0.0, so the tags, not the manifest, are the version of record.

Editorial conclusion

This fits a site operator who already has cameras feeding RTSP and wants a warning on a phone when someone walks into a controlled area. It does not fit a quick local trial on a laptop, because the supported Python range is narrow, the model weights have to be fetched, and the production schema comes from a private deployment process rather than the repository. Before trusting a warning, check how the cone zones were derived for your site, since a zone built from misplaced cones turns the rule into noise. Anyone deploying should also set a real JWT secret and a Redis password before starting compose, since the example values are placeholders.

Frequently asked questions

What does Construction-Hazard-Detection detect on a live camera?

Workers without hard hats, workers without safety vests, workers too close to machinery or vehicles, workers inside cone-derived controlled areas, and machinery or vehicles too close to utility poles.

Does Construction-Hazard-Detection keep video frames in Redis?

No. MediaMTX owns live playback, and Redis holds only small coordination state such as the authentication cache, the FCM token cache, compact warning metadata, overlay demand keys and overlay ready keys.

Which Python versions does Construction-Hazard-Detection support?

The project targets Python >=3.14,<3.15, declared in pyproject.toml and repeated in the setup step.

How many cameras can one YOLO worker engine handle?

The default is three per engine. Capacity is overridable per model with YOLO_WORKER_CAMERAS_PER_ENGINE_BY_MODEL, for example yolo26n=8, yolo26s=6, yolo26m=4 and yolo26l=2, and model keys left out of that list use the default.

Where does Construction-Hazard-Detection get its model weights?

From the Hugging Face model repository yihong1120/Construction-Hazard-Detection using hf download with models/pt/*.pt. Worker files are expected at models/pt/best_<model_key>.pt, such as best_yolo26n.pt.

Can Construction-Hazard-Detection create its own database schema?

Not for production. AUTO_CREATE_SCHEMA defaults to false because the production schema is managed through a private deployment process, and the flag is meant for throw-away local databases built from ORM metadata.

Official sources

  1. License: AGPL-3.0
  2. Project website
  3. README
  4. Releases
  5. yihong1120/Construction-Hazard-Detection on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/yihong1120-construction-hazard-detection.svg)](https://hysenlabs.com/projects/yihong1120-construction-hazard-detection)