# ALICE: a YOLO dataset pipeline that reads Frigate's database and trains on CPU if it must

> ALICE annotates bounding boxes, pulls snapshots out of a Frigate NVR, and runs a five-step export, dedup, annotate, train and ONNX export pipeline for camera models. It ships as one assembled alice.py and a Docker Compose file generated from your hardware.

**simoncirstoiu/alice** — Analyse · Learn · Ingest · Curate · Export — AI-powered YOLO dataset management toolkit

- Repository: https://github.com/simoncirstoiu/alice
- Stars: 409 · Forks: 43
- Language: JavaScript
- License: NOASSERTION
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/simoncirstoiu-alice

## builder.py assembles a single alice.py before anything else runs

There is no package to install. Two commands produce a runnable tool:

```bash
# Build and set up virtual environment
python3 builder.py

# Run
./alice.py
```

The builder assembles `alice.py` from the source modules in `src/`, creates a `.venv` with the base dependencies, and patches the shebang so `./alice.py` picks up the virtualenv without an activate step. On Debian and Ubuntu it installs `python3-venv` itself if the package is missing.

First run opens a welcome page at http://localhost:8080 that walks through hardware detection, the model download and dependency installation. Two flags matter afterwards:

```bash
# Options
./alice.py --port 9090                  # custom port
./alice.py --conf /path/to/alice.conf   # custom config path
```

Configuration lives in `alice.conf` and can be edited in the web UI or straight in the file. Everything else, including paths, training parameters and AI model choices, is in that one file.

## The frigate.db mount has to be a file, or Docker creates a directory

ALICE is built around a Frigate NVR. Live mode browses Frigate event snapshots in real time with filters for camera and time window, and video mode steps through Frigate exports frame by frame with a seekbar and adjustable playback FPS. Snapshots move into the training set with automatic WebP to JPG conversion.

The container path is where people lose an afternoon. The generated `docker-compose.yml` mounts six volumes:

```yaml
volumes:
  - ./alice.conf:/app/alice.conf
  - /path/to/datasets:/app/datasets
  - /path/to/models:/app/models
  - /path/to/frigate/clips:/app/clips:ro
  - /path/to/frigate/exports:/app/exports:ro
  - /path/to/frigate/frigate.db:/app/frigate.db:ro
```

The last line must point at the actual `frigate.db` file. If it does not exist on the host at container creation time, Docker creates a directory in its place and ALICE cannot open the database. Two of those mounts are read-only, and the clips and exports directories are where Frigate writes, so the host paths have to be Frigate's real ones.

## Dedup runs pHash, box similarity and NMS, and the split is 90/10

Step 1 of the pipeline extracts the newest snapshots from the Frigate database with round-robin camera distribution, then splits them 90/10 into train and val. Round-robin is the part that matters for a house with eight cameras: it stops one busy driveway from filling the training set.

Step 2 removes duplicates through three mechanisms. Perceptual hashing uses DCT-based 64-bit hashes, computed with multiprocessing, and candidates are compared side by side before you delete anything. Box similarity catches frames that are visually different but carry the same annotations for one camera. NMS cleanup removes overlapping boxes of the same class that survived labelling.

Duplicate detection also runs inside the dataset viewer, not only in the pipeline. The dataset mode filters by split (train, val, empty) or by class, and the stats panel reports total images, train and val distribution, annotation coverage and class distribution, which is how you notice that one class is at 3% coverage before a training run wastes an evening on it.

## A teacher model labels the set, a student model gets fine-tuned

Step 3 auto-labels every image with a teacher model. Step 4 fine-tunes the student model and streams loss, mAP50 and mAP50-95 into the UI while it runs. Step 5 converts the result to ONNX for deployment, offering FP16 and FP32 on GPU and FP32 only on CPU.

Each of the five steps is individually toggleable and runnable on its own or as a sequence, and every run logs to the Logs tab with a COMPLETED, FAILED or STOPPED status. That matters because the steps are not equally slow: annotation over a few thousand images and a training run are different kinds of wait, and a failure in step 3 does not mean step 2 failed.

Inference is available in all three viewer modes. Detected boxes merge into existing annotations with dedup by IoU above 0.5, so running the teacher over a hand-labelled image does not double every box. Preview mode shows detections without saving anything, and live detection auto-runs as you navigate in live mode, which turns the viewer into a labelling assistant rather than a separate tool.

## CPU mode is a complete path, except for the FP16 export

Hardware detection runs at startup and the only setting you normally touch is Settings, System, Device, where the choice is Auto, NVIDIA GPU or CPU. The sidebar and the Device tab show live hardware statistics.

What changes between the two paths is small and worth knowing in advance. Training is GPU accelerated or CPU and slower. Inference follows the same split. ONNX export offers FP16 and FP32 on GPU and FP32 only on CPU, which matters if your deployment runtime expects half precision. PyTorch installs as the CUDA build or the CPU build, and onnxruntime resolves to `onnxruntime-gpu` or `onnxruntime`. ALICE picks the right variant of each, so there is no torch installation step to get wrong.

One boundary is explicit: ALICE does not install NVIDIA drivers or the CUDA toolkit. Those come first, from the system, and only if you want GPU-accelerated training. The docs also ask for Python 3.8 or newer, while the container image is built on `python:3.11-slim`, so the two paths are not tested against the same interpreter.

## The image ships one assembled file and a single preinstalled package

The `Dockerfile` starts from `python:3.11-slim`, installs `libgl1`, `libglib2.0-0` and `wget` for OpenCV headless and model downloads, and installs only Pillow with pip. Every other dependency arrives later, through the welcome page or Settings, System, with a one-click install that adapts to GPU or CPU. Inside the container those dependencies persist in a Docker volume across restarts.

Because the image copies a single `alice.py` into `/app` and runs `python3 alice.py --port 8080`, the assembled file has to exist in the build context before you build. Running the builder with `--no-venv` is what generates `docker-compose.yml`, and it picks the GPU block or the CPU-only service depending on the hardware you have configured.

```bash
python3 builder.py --no-venv
```

```bash
docker compose up --build -d
```

The rest of the dependency list is Pillow for image processing, pHash and format conversion, NumPy for dedup maths, inotify for filesystem watching with a polling fallback, opencv-python-headless for frame extraction, ONNX and onnxslim for the export path, and ultralytics for YOLO training and inference.

## Three patch releases in April 2026, and a license GitHub cannot name

The releases cluster tightly: 0.6.1 on 2026-04-23, 0.6.2 on 2026-04-25 and 0.6.3 on 2026-04-26, with the last push on the same date as 0.6.3. There is nothing after April 2026 in the repository.

Two loose ends are worth flagging before you build on it. A LICENSE file sits at the top level of the tree, but GitHub reports no license identifier for the repository, so the terms are whatever that file says and you have to read it yourself. And the README ends mid-note in the Docker section, so the warning it was carrying, a GPU specific detail by the look of the surrounding text, is not visible on the page.

The code itself is plain enough to audit. `builder.py` sits at the root beside `alice.conf`, with the modules in `src/` and the tests in `tests/`, so the assembly step and the settings file are both readable without running anything.

## Conclusion

ALICE is worth the setup for anyone already running Frigate who wants a per-camera YOLO model without writing an export script, and it is the wrong tool for a generic image classifier, since the pipeline assumes a Frigate database, camera streams and a bounding-box task. Check three things first. Point the frigate.db mount at the file, not at a directory, or Docker creates a directory and the database will not open. Install NVIDIA drivers and CUDA yourself if you expect GPU training, since ALICE does not. And note the project has not pushed since 2026-04-26, with releases 0.6.1 to 0.6.3 all landing in the last week of April 2026, so pin 0.6.3 and expect to read the source if your Frigate schema has moved on.

## FAQ

### What does ALICE do for a YOLO project?

It bundles dataset work and training in one tool: a canvas annotation editor, Frigate snapshot and video browsing, a five-step export, dedup, annotate, train and ONNX export pipeline, and automatic GPU or CPU selection. Configuration lives in one alice.conf file.

### Does ALICE need an NVIDIA GPU to run?

No. Hardware detection at startup falls back to CPU mode when no compatible GPU is found. The differences are training and inference speed, the PyTorch and onnxruntime variants installed, and ONNX export, which offers FP16 and FP32 on GPU and FP32 only on CPU.

### What does ALICE not install for me?

It does not install NVIDIA drivers or the CUDA toolkit. GPU-accelerated training needs both on the system first. Everything else, from ultralytics to PyTorch, is managed by ALICE through a one-click install from the welcome page or Settings, System.

### How does ALICE remove duplicate training images?

Through perceptual hashing with 64-bit DCT-based hashes, box similarity per camera, and NMS cleanup for overlapping boxes of the same class. The pipeline runs all three in its dedup step, and the same duplicate view is available inside the dataset mode.

### Why can ALICE not open the Frigate database in Docker?

The frigate.db volume has to point at the actual file. If it does not exist on the host when the container is created, Docker creates a directory there instead and ALICE cannot open the database.

## Sources

- [Issues](https://github.com/simoncirstoiu/alice/issues)
- [README](https://github.com/simoncirstoiu/alice/blob/main/README.md)
- [Releases](https://github.com/simoncirstoiu/alice/releases)
- [simoncirstoiu/alice on GitHub](https://github.com/simoncirstoiu/alice)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/simoncirstoiu-alice
