# DeepDetect: a C++ deep learning server and CLI for Torch and TensorRT

> DeepDetect packages a C++ training and inference runtime behind a REST server, a Python wheel and a `deepdetect` CLI. It is aimed at teams that want to operate models on their own filesystem rather than inside a notebook, and the trade-off is a narrower set of packaged workflows than a general-purpose framework.

**jolibrain/deepdetect** — Deep Learning Server and CLI for Torch and TensorRT

- Repository: https://github.com/jolibrain/deepdetect
- Website: https://www.deepdetect.com/
- Stars: 2,551 · Forks: 545
- Language: C++
- License: NOASSERTION
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/jolibrain-deepdetect

## The gap DeepDetect fills between a training script and a serving process

Most deep learning code starts as a script that loads a checkpoint and prints a tensor. Turning that into something another team can call means writing a service, a job queue, a model directory convention and a response format. DeepDetect is the project's answer to that second half. The README describes it as a "deep learning runtime, command-line tool, and REST server for training and inference", and the capabilities list is operational rather than architectural: create services, train models, run predictions, monitor jobs, keep model repositories organized on the filesystem.

The intended user is an engineer who already has a model or a known architecture family and needs to run it repeatedly. The README states that the wheel embeds the runtime in the current Python environment and exposes a `deepdetect` CLI for "repeatable model workflows", while the server "remains available for long-running REST services, containerized serving, and integrations that need a dedicated process". That split is the clearest signal of who this is for: people who want the same runtime in two shapes, one for a laptop or CI job and one for a persistent endpoint.

It is not a framework for designing new architectures. Nothing in the README suggests you define custom layers or write training loops. You pick from model families the project has already wired up, supply data and paths, and let the runtime handle the rest.

## Torch as the runtime, TensorRT as the acceleration path

The backend story is stated plainly: Torch is the primary backend for both training and inference, and TensorRT is available for "optimized inference with exported or compatible models". Caffe appears only as a compatibility detail. The README is explicit that "Caffe-format protobufs and prototxt files may appear as compatibility or model-format details, but Caffe is not an active runtime backend". Anyone arriving from an older DeepDetect deployment where Caffe was central should read that sentence twice, because the repository topics still list caffe alongside pytorch and tensorrt.

On top of that runtime sits a template system. The README lists the families the broader Torch API supports: image classification with TorchVision-style classifiers such as ResNet, VGG, DenseNet, MobileNet, ShuffleNet and SqueezeNet; detection with YOLOX, Faster R-CNN and RetinaNet; segmentation with SegFormer; language and traced models such as BERT and GPT-2; time series with recurrent, N-BEATS, transformer and time-transformer templates; and vision transformers such as ViT and Visformer. Those names are the contract. If your architecture is not in that list, the README does not describe a path to adding it without touching the C++ source.

Data flows through a single API surface that the README says covers images, text, CSV and tabular data, time series, sparse and SVM-style data, object-detection boxes and segmentation masks. Outputs are similarly typed: "classes, scores, bounding boxes, segmentation masks, and model-specific metrics". Model storage is filesystem-based with no database dependency, which is the design decision that most shapes day-to-day operation. Backups are directory copies, and a model repository is something you can inspect with `ls`.

## Installing the wheel and running a first YOLOX job

The README gives one install path: a wheel from the project's own index. The two variants are mutually exclusive because both provide `import deepdetect`, so pick one per environment.

```bash
python -m pip install \
  --extra-index-url https://www.deepdetect.com/download/wheels/simple \
  deepdetect-cpu
```

Substitute `deepdetect-gpu` for a CUDA environment. After installing, the CLI can list what it ships with. The README shows these three commands as the way to inspect packaged profiles and per-command options.

```bash
deepdetect inspect models
deepdetect train yolox --help
deepdetect infer segformer --help
```

The first profiles are `yolox`, `segformer`, `torchvision-detector` and `external-pytorch-detector`. A real run needs a dataset and a model repository, and the README provides a preparation script for a tiny object-detection example that reuses the wheel test fixtures.

```bash
python bindings/python/scripts/prepare_cli_yolox_quickstart.py \
  --output /tmp/deepdetect-yolox-quickstart \
  --force
```

The script prints a `Sample image:` path. Training then reads a generated YAML config, and inference reuses the same file, which is the point of the config approach: the run is reproducible from one artifact.

```bash
deepdetect train yolox \
  --config /tmp/deepdetect-yolox-quickstart/yolox-quickstart.yaml \
  --terminal live
```

```bash
deepdetect infer yolox <sample-image> \
  --config /tmp/deepdetect-yolox-quickstart/yolox-quickstart.yaml \
  --visualize \
  --output /tmp/deepdetect-yolox-quickstart/detections.png
```

Add `--gpu` to both commands when using the GPU wheel. The README also suggests starting Visdom on port 8097 in a second terminal if you want live plots. One warning the README makes directly: the shipped YAML examples are starting points, and dataset, weight and repository paths must be replaced before a real run.

## Where DeepDetect is the wrong choice

The packaged CLI covers two workflows, YOLOX and SegFormer. Everything else in the model list is reachable through the broader Torch API and the server, but the README does not present a CLI profile for classification, time series or language models. If your work is a sequence-to-sequence training run with a custom scheduler, this is the wrong layer. You would be fighting the template system to express something a plain PyTorch script does in thirty lines.

The second constraint is the install path itself. Wheels come from `https://www.deepdetect.com/download/wheels/simple`, not from PyPI. That means an environment without access to that host cannot install the wheel at all, and the README's alternative is building from source with `docs/source.md`. For an air-gapped deployment, the practical route is mirroring the index or producing a source build, and neither is described in the README.

A third limitation is documentation depth. Several important behaviours are delegated to other files rather than explained inline: config precedence and output formats live in `bindings/python/deepdetect/cli/CLI_SPEC.md`, service parameters and connectors in `docs/api.md`, container usage in `docs/docker.md`. The README names them but does not summarize them, so the real learning curve starts after the quickstart. Note also that the licence field on the repository is NOASSERTION while the README states LGPL v3.0, which is a discrepancy worth resolving before you rely on either.

## DeepDetect compared with running TorchServe or a plain FastAPI wrapper

The obvious alternative is to skip the server entirely: write a FastAPI or Flask endpoint that loads a PyTorch checkpoint and returns JSON. That gives you exactly the response schema you want and no template constraints. What you give up is everything around the model. DeepDetect's README describes asynchronous training jobs with status inspection, model services backed by local repositories, and a documented request and response format in `docs/api.md`. Reproducing that with a hand-rolled wrapper means writing the job tracking and the repository convention yourself.

TorchServe sits closer to DeepDetect in shape, since both are serving processes with model stores. The difference visible here is the backend range and the CLI. DeepDetect pairs the Torch runtime with TensorRT for optimized inference and ships a Python wheel that embeds the runtime in-process, so the same tool covers a CI job and a long-running service. A serving-only tool does not give you the `deepdetect train yolox --config ...` path at all.

The honest framing is that DeepDetect trades flexibility for operational surface. If your models fit the listed families and you want training and serving behind one API, the trade is favorable. If your models do not fit, the template system becomes the constraint rather than the feature.

## Maintenance, releases and what the LGPL v3.0 licence implies

The repository is not archived, and the last push was on 2026-08-28. Recent releases are v0.30.0 on 2026-08-26, v0.29.0 on 2026-06-25 and v0.28.0 on 2026-06-12, with `package.json` reporting version 0.30.0 and a `standard-version` setup that regenerates `CHANGELOG.md`. The release cadence visible in that list is roughly every two to three months, and the gap between the v0.30.0 tag and the last push is two days, so the release line and the branch are close together.

Upgrade cost depends on which surface you use. The CLI and wheel are versioned together, and the README warns that `deepdetect-cpu` and `deepdetect-gpu` are mutually exclusive in one environment, so an upgrade means reinstalling the right variant. Config files are YAML and the README treats them as the unit of reproducibility, which means a config written against one release is the artifact most likely to need attention after a bump. The README does not document rollback or a config migration path, and `CHANGELOG.md` is the only place release-to-release changes are recorded.

The licence situation needs care. The README states DeepDetect is distributed under the GNU Lesser General Public License v3.0 and points at `COPYING`, while the repository metadata reports NOASSERTION. LGPL v3.0 is a copyleft licence with conditions that attach to distribution and to linking, and those conditions differ depending on whether you ship the runtime, link against it or only call the server over HTTP. Whether your deployment triggers them is a question for your own counsel; what matters here is that the two sources disagree on the identifier, so read `COPYING` directly rather than trusting the metadata field.

## Conclusion

Adopt DeepDetect if you need a long-running inference or training service with JSON payloads and filesystem model storage, or if you want the `deepdetect` CLI around YOLOX and SegFormer workflows. Do not adopt it if you expect the full breadth of a general PyTorch training stack, or if you cannot run your own wheel index and model repository. Before committing, verify the wheel variant you need (`deepdetect-cpu` or `deepdetect-gpu`), confirm that the default YAML configs have been pointed at your own dataset and weight paths, and read `bindings/python/deepdetect/cli/CLI_SPEC.md` for config precedence, because the README only points at it rather than restating it.

## FAQ

### What is DeepDetect?

It is a deep learning runtime, command-line tool and REST server for training and inference, written in C++ and maintained by Jolibrain. The Python wheel embeds the runtime and provides the `deepdetect` CLI, while the server handles long-running REST services and containerized serving.

### How do I install DeepDetect?

Install one wheel variant with pip using the project's extra index: `deepdetect-cpu` or `deepdetect-gpu`. The two are mutually exclusive in the same environment because both provide `import deepdetect`.

### Does DeepDetect still use Caffe?

No. The README states that Torch is the primary backend for training and inference, and that Caffe-format protobufs and prototxt files may appear as compatibility or model-format details, but Caffe is not an active runtime backend.

## Sources

- [Issues](https://github.com/jolibrain/deepdetect/issues)
- [jolibrain/deepdetect on GitHub](https://github.com/jolibrain/deepdetect)
- [Project website](https://www.deepdetect.com/)
- [README](https://github.com/jolibrain/deepdetect/blob/master/README.md)
- [Releases](https://github.com/jolibrain/deepdetect/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/jolibrain-deepdetect
