# FastestDet: a 250K-parameter anchor-free detector for CPU-only edge boards

> FastestDet is a single-scale, anchor-free object detector from the author of Yolo-Fastest, aimed at ARM CPUs and NPUs. The README reports 25.3% mAP at 0.5 IoU on COCO2017 at 352x352, and the repository ships training, evaluation and export code plus ncnn and ONNX Runtime examples.

**dog-qiuqiu/FastestDet** — :zap: A newly designed ultra lightweight anchor free target detection algorithm， weight only 250K parameters， reduces the time consumption by 10% compared with yolo-fastest, and the post-processing is simpler

- Repository: https://github.com/dog-qiuqiu/FastestDet
- Stars: 861 · Forks: 151
- Language: Python
- License: BSD-3-Clause
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/dog-qiuqiu-fastestdet

## The problem FastestDet targets: detection on boards without a GPU

Most detection projects assume a GPU somewhere in the loop. FastestDet assumes the opposite. The README's benchmark table lists a Radxa Rock3A with an RK3568 ARM Cortex-A55 CPU, locked at 2.0GHz, running through ncnn: 70.62ms single core, 23.51ms with four cores. The same board through the RK3568 NPU with rknn reports 28ms. An i7-8700 on x86 through ncnn reports 4.51ms single core and 4.33ms multi core. Those are the numbers the project publishes, not numbers this article measured.

The intended user is someone shipping a detector onto a fixed-function device: a Raspberry-class board, an Android phone, an industrial camera box. The model is 0.24M parameters according to the README table, and the repository ships a weights directory and a checkpoint directory, so the workflow ends in a small file you can load from C++ or from Python. The pitch is not accuracy. The pitch is that a 352x352 forward pass fits in a CPU budget where YOLOv5s at 640x640 does not: the README lists yolov5s at 395.31ms on four cores on the same board, against 23.51ms for FastestDet.

## Anchor-free, single-scale: how the detector head is actually built

The README lists four design choices under Improvement: anchor-free, a single-scale detector head, cross-grid multiple candidate targets, and dynamic positive and negative sample allocation. The first two are the ones that shape everything else.

An anchor-free head means the network predicts box coordinates directly instead of regressing offsets from a set of pre-defined anchor boxes. That removes the anchor matching step at inference time, which is where the README's claim of simpler post-processing comes from. A single-scale head means there is no feature pyramid: one output grid, one stride. The consequence is visible in the COCO2017 evaluation output pasted into the README, where AP for small objects is 0.021 against 0.129 for medium objects at IoU 0.50:0.95. That is not a bug, it is the arithmetic of a single coarse grid.

Cross-grid candidate assignment and dynamic sample allocation are training-time mechanisms. Instead of assigning each ground-truth box to exactly one grid cell, multiple cells can carry candidates, and the positive/negative split is decided dynamically rather than by a fixed IoU threshold. The README's changelog entry for 2022.7.14 says the loss was changed to IOU aware based on smooth L1 and that AP rose by 0.7. The config exposes a THRESH key under TRAIN, annotated in the README with question marks, so the exact role of that threshold is not documented there.

## Installing FastestDet and running the picture test

The README gives one dependency step and one inference command. Install from the pinned requirements file first. Note that requirements.txt pins torch==1.11.0+cu113 and torchvision==0.12.0+cu113, so a CPU-only machine needs a different torch build than the one listed; the README warns only that the PyTorch CUDA version selection is yours to make.

```bash
pip install -r requirements.txt
```

With dependencies in place, the README's picture test loads a YAML config, a checkpoint and an image. The weight filename in the README encodes the reported score, weight_AP05:0.253207_280-epoch.pth, and the repository has a weights directory for it.

```bash
python3 test.py --yaml configs/coco.yaml --weight weights/weight_AP05:0.253207_280-epoch.pth --img data/3.jpg
```

What you should see is the input image with boxes drawn on it, matching the result.png in the repository root. The command writes no metrics; it is a visual check. To reproduce the numbers in the benchmark table you need the evaluation path instead, which the README documents as eval.py with the same --yaml and --weight flags. The COCO2017 output pasted in the README shows AP@[IoU=0.50:0.95] = 0.130 and AP@[IoU=0.50] = 0.253, and the run itself takes about 30.85s for per-image evaluation plus 4.97s to accumulate results on the machine that produced that log.

## Training on your own data: Darknet labels, a path list, and a YAML

FastestDet reuses the Darknet YOLO dataset convention, which lowers the cost of trying it if you already have YOLO data. Each image gets a .txt file with the same basename in the same directory, one line per object, in the order category cx cy wh, all normalized. The README's example lines are:

```
11 0.344192634561 0.611 0.416430594901 0.262
14 0.509915014164 0.51 0.974504249292 0.972
```

You then write two text files listing absolute image paths, one for train and one for val, plus a .names file with one class name per line. The YAML ties them together and sets the model geometry. The README's reference file is configs/coco.yaml, and the keys are DATASET.TRAIN, DATASET.VAL, DATASET.NAMES, MODEL.NC, MODEL.INPUT_WIDTH, MODEL.INPUT_HEIGHT, and under TRAIN: LR, THRESH, WARMUP, BATCH_SIZE, END_EPOCH and MILESTIONES.

```yaml
DATASET:
  TRAIN: "/home/qiuqiu/Desktop/coco2017/train2017.txt"
  VAL: "/home/qiuqiu/Desktop/coco2017/val2017.txt"
  NAMES: "dataset/coco128/coco.names"
MODEL:
  NC: 80
  INPUT_WIDTH: 352
  INPUT_HEIGHT: 352
```

Training is a single call. The README's example uses 350 epochs with learning-rate drops at 150, 250 and 300, batch size 64 and LR 0.001.

```bash
python3 train.py --yaml configs/coco.yaml
```

Two things to check before you commit to a long run. First, INPUT_WIDTH and INPUT_HEIGHT are part of the model definition, not a runtime resize, so changing them changes the network you train. Second, THRESH appears in the config with a question mark annotation in the README, which means the project itself does not explain what it controls.

## Where FastestDet is the wrong tool

Small objects are the clearest failure mode. The README's own COCO2017 log reports AP 0.021 for the small area bucket at IoU 0.50:0.95, against 0.129 for medium and higher for large. A single-scale head at 352x352 gives a coarse grid, and a pedestrian 20 pixels tall will often fall below it. If your application is counting distant vehicles or reading small defects on a conveyor, a multi-scale detector such as YOLOv5s will beat FastestDet on exactly the cases you care about, at roughly sixteen times the runtime on the same board according to the README table.

The second limitation is maintenance. The last push to the repository was on 2026-05-12, and the only release listed is v1.0 from 2022-07-02. The changelog entry at the top of the README is dated 2022.7.14. The requirements file pins torch 1.11.0+cu113, torchvision 0.12.0+cu113, numpy 1.23.0 and opencv_python 4.2.0.34, which is an old stack; installing it on a current Python is likely to need substitutions the README does not describe, and no upgrade path is documented.

Third, the accuracy ceiling is low by design. 25.3% mAP at 0.5 IoU on COCO means roughly three in four detections at that threshold are wrong in the strict sense. That is acceptable for a coarse trigger that wakes a heavier model, and unacceptable as the only detector in a system that reports counts.

## FastestDet compared with NanoDet-Plus and Yolo-FastestV2

The two projects people reach for in the same slot are NanoDet-Plus and Yolo-FastestV2, the latter by the same author.

NanoDet-Plus is the more conventional anchor-free design: it keeps a feature pyramid, which is why the README's table lists nanodet_m at 20.6% mAP at IoU 0.50:0.95 at 320x320 with 0.95M parameters, against 13.0% for FastestDet at 352x352 with 0.24M. NanoDet-Plus spends four times the parameters to buy multi-scale detection. If your objects vary in size, that is the trade you want, and the extra parameters are still tiny by general detection standards. FastestDet's answer is to give up the pyramid entirely and take the small-object hit.

Yolo-FastestV2 is the closer comparison because it shares an author and a deployment philosophy. The README puts it at 24.10% mAP at 0.5 IoU, 23.8ms on four cores, 0.25M parameters at 352x352, versus FastestDet at 25.3%, 23.51ms and 0.24M. The measured difference is about one point of mAP at the same speed and size, and the README claims post-processing is simpler in FastestDet because the head is anchor-free. That is the actual reason to pick one over the other: not accuracy, but how much matching code you want to maintain in your inference wrapper. The repository ships example/ncnn/ and example/onnx-runtime/ directories, so both deployment routes are demonstrated in-tree.

## Licence, upgrade cost, and what the repository does not tell you

FastestDet is BSD-3-Clause, per the LICENSE.md file in the repository root and the licence badge in the README. That is a permissive licence: it allows commercial use and modification, and it requires retaining the copyright notice and the licence text in redistributions. It contains no patent grant, which matters if you are shipping on hardware where a patent claim is plausible. This is a description of the licence text, not legal advice; check it against your own distribution model.

The upgrade cost is the part to weigh. Because requirements.txt pins exact versions of torch, torchvision, onnx, onnxruntime and opencv_python, moving to a newer CUDA, a newer Python or a newer ONNX Runtime means editing those pins yourself. The README does not document rollback, a supported Python version, or a compatibility matrix. The evaluation log in the README was produced with pycocotools 2.0.4, and that is the only version information the project gives for the eval path.

What the repository does not tell you is equally relevant. There is no documented export command for ONNX or ncnn in the README body, even though example/ncnn/ and example/onnx-runtime/ exist; you will have to read those directories to find the conversion and inference code. The THRESH key is unexplained. And the README makes no statement about training time, dataset size requirements, or how the dynamic sample allocation behaves on small datasets.

## Conclusion

Adopt FastestDet when your target is a fixed 352x352 input on an ARM CPU or an RK3568 NPU and you can accept roughly 25% mAP at 0.5 IoU. Do not adopt it if you need small-object accuracy or a maintained training stack, because the last push was on 2026-05-12 and the pinned requirements still target torch 1.11.0+cu113. Verify the label format and the coco.yaml paths before training, and confirm the exported ONNX graph matches your runtime.

## FAQ

### How many parameters does FastestDet have?

The README's benchmark table lists FastestDet at 0.24M parameters, measured at a 352x352 input resolution. That is the smallest model in the comparison table, below yolo-fastestv2 at 0.25M and yolox-nano at 0.91M.

### What dataset format does FastestDet expect for training?

The README states the dataset is constructed the same way as Darknet YOLO: each image has a .txt label file with the same name in the same directory, and each line is category cx cy wh with normalized values. You also need train.txt and val.txt files listing absolute image paths, and a .names file with one class name per line.

### How do I run inference on a single image with FastestDet?

The README's picture test passes a YAML config, a weight file and an image to test.py, for example with --yaml configs/coco.yaml, --weight weights/weight_AP05:0.253207_280-epoch.pth and --img data/3.jpg. The output is an annotated image, not a metrics report.

### Is Yolo faster than Faster RCNN?

This question is not about FastestDet, and the repository does not benchmark Faster RCNN. The README's comparison table covers yolov5s, yolov6n, yolox-nano, nanodet_m, yolo-fastestv1.1, yolo-fastestv2 and FastestDet itself, all measured on a Radxa Rock3A RK3568 through ncnn.

## Sources

- [dog-qiuqiu/FastestDet on GitHub](https://github.com/dog-qiuqiu/FastestDet)
- [Issues](https://github.com/dog-qiuqiu/FastestDet/issues)
- [License: BSD-3-Clause](https://github.com/dog-qiuqiu/FastestDet/blob/main/LICENSE)
- [README](https://github.com/dog-qiuqiu/FastestDet/blob/main/README.md)
- [Releases](https://github.com/dog-qiuqiu/FastestDet/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/dog-qiuqiu-fastestdet
