depth-anything.cpp: Depth Anything V2 and V3 in C++ without Python
A from-scratch C++17/ggml port of Depth Anything 2 and 3 (ByteDance)
At a glance
- What is it?
- depth-anything.cpp is a MIT-licensed C++17/ggml port of Depth Anything V2 and V3 by the LocalAI team that runs metric depth estimation, camera pose recovery, and 3D point cloud export on CPU or GPU with no Python, PyTorch, or CUDA toolkit required at inference time.
- Who is it for?
- depth-anything.cpp is the right choice when Depth Anything inference must run in an environment without Python or PyTorch, when model file size matters (q4_k at 99 MB versus 516 MB for the PyTorch f32 checkpoint), or when load latency is a constraint. The C API makes it embeddable in Go, Rust, and C projects without any Python interop layer.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 24 days ago.
- What is it written in?
- Mainly C++, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What depth-anything.cpp provides and who needs it
Depth Anything is a monocular depth estimation model family from ByteDance. The V2 and V3 generations, particularly Depth Anything V3 (DA3), produce metric depth maps, per-pixel confidence scores, camera extrinsics (a 3x4 matrix), camera intrinsics (a 3x3 matrix), a sky mask, and a back-projected 3D point cloud from a single input image. The official reference implementation runs in Python with PyTorch, which means every deployment carries a Python environment, the PyTorch library, and a CUDA toolkit.
depth-anything.cpp removes that dependency stack from the inference path. It is a complete port of the model written in C++17 using the ggml tensor library, packaged as a single self-contained GGUF file and a small CLI binary. A machine running depth-anything.cpp at inference time does not need Python, PyTorch, or the CUDA toolkit installed. The ggml library itself is bundled as a submodule.
The project targets three kinds of users: embedded and systems developers who need depth estimation in a C or Rust application without pulling in a Python runtime, teams deploying on hardware where Python installation is impractical or prohibited, and anyone who wants to trade the full Python model footprint for a 99 MB quantized file. The repository does not have GitHub releases; the code on master corresponds to the state at the last push on 2026-09-07. The project is MIT-licensed.
How the ggml port achieves bit-exact parity with PyTorch
The README states that every component of the DA3 forward pass is gated against PyTorch-dumped reference tensors, and that end-to-end depth output matches the real PyTorch model at correlation 1.0. The port does not approximate the model or skip components; it reconstructs each layer with equivalent ggml operations and verifies the result numerically.
The GGUF file format is the key to making this self-contained. Every dimension, hyperparameter, and preprocessing constant lives inside the GGUF file itself. The loader reads them out at startup; nothing is hardcoded in the library. This means switching from a small model to a large one requires loading a different GGUF file, not recompiling or changing code.
Parity holds across quantization levels down to f16. The q8_0 quantization is described as near-lossless. The q4_k format at 99 MB is the smallest and trades a small quality reduction for a 5.2x reduction in model size relative to the PyTorch f32 checkpoint at 516 MB.
The explanation for why the C++ path is faster than PyTorch on CPU comes down to caching. Two positional embeddings in the DPT head and backbone were recomputed every forward pass in PyTorch using single-threaded scalar operations. They depend only on input geometry and are identical across calls. The C++ port caches them after the first call, removing roughly 95 ms of per-forward host overhead.
The README publishes a full benchmark table measured on an AMD Ryzen 9 9950X3D (16-core, 32-thread x86) at 504x336 pixels with threads=16 and repeat=25. PyTorch f32 loads in 749 ms and runs each forward in 416.9 ms, using 1328 MB of peak RAM. The C++/ggml f32 build loads in 112 ms and runs in 346.4 ms with 614 MB of peak RAM, a 1.20x speedup over PyTorch. At q8_0 the load time drops to 40 ms and the forward pass to 319.4 ms at 363 MB, a 1.31x speedup. The q4_k build loads in 25 ms and uses only 320 MB of peak RAM. The full methodology, including the CPU optimization history, is documented in benchmarks/BENCHMARK.md.
Building the CLI and enabling GPU backends
The build requires cloning with submodules so that the ggml dependency is fetched:
git clone --recursive https://github.com/mudler/depth-anything.cpp
cd depth-anything.cpp
cmake -B build -DDA_BUILD_CLI=ON
cmake --build build -jThe resulting binary is at build/examples/cli/da3-cli. By default, the build enables the da3-cli tool (DA_BUILD_CLI is ON) and the tinyBLAS AVX-512/AVX2 matrix multiplication path via DA_GGML_LLAMAFILE (also ON by default). The parity test suite is off by default; to build it, pass -DDA_BUILD_TESTS=ON.
To enable CUDA acceleration, add -DDA_GGML_CUDA=ON to the cmake configuration line. The README also documents Metal (for Apple Silicon) and Vulkan as alternative ggml GPU backends. On an NVIDIA GB10 GPU with CUDA enabled, the README reports that depth-anything.cpp ties PyTorch's tuned cuDNN at roughly 47 ms per forward pass across all quantizations, while loading 1.75x to 2.9x faster on cold start (roughly 548 ms versus 926 ms). Full GPU benchmark details including the Grace Blackwell configuration are in benchmarks/BENCHMARK.md.
Building the shared library instead of the CLI uses -DDA_SHARED=ON. This produces libdepthanything.so and is the path for embedding the library in another application. The export capabilities available through the CLI include glb (binary glTF), COLMAP sparse reconstruction format, and PLY point cloud files.
Model variants and the GGUF conversion path
The repository supports seven DA3 model variants, ranging from DA3-SMALL (ViT-S backbone, fastest) to DA3-NESTED-GIANT-LARGE (ViT-g plus ViT-L, two-branch alignment for metric depth). All share the same metadata-driven engine; the GGUF file encodes which architecture to use. The DA3-GIANT variant adds 3D Gaussian output on top of the standard depth, confidence, and pose outputs.
For Depth Anything V2, the supported variants include relative-depth and metric-depth models for indoor (max_depth=20 metres with the Hypersim-trained variants) and outdoor (max_depth=80 metres with the VKITTI-trained variants) scenes. V2 produces depth only, without confidence, pose, or sky outputs.
Converting a checkpoint from the official HuggingFace repository to GGUF requires the Python dependencies listed in requirements.txt: torch>=2.4, numpy>=1.26, gguf>=0.19.0, huggingface_hub>=0.24, safetensors>=0.4, einops>=0.8, pillow>=10, matplotlib>=3.8, and opencv-python-headless>=4.9. The requirements.txt comment explicitly states these are for conversion and parity checks only; the inference binary (da3-cli or libdepthanything.so) needs none of them. A download helper script at scripts/download_model.py fetches checkpoints from HuggingFace using the huggingface_hub library.
Pre-converted GGUF files for several DA3 variants are available on HuggingFace at mudler/depth-anything.cpp-gguf, which the README links directly. These allow skipping the Python conversion step entirely.
The flat C API and embedding in other applications
The public interface for embedding depth-anything.cpp in another project is declared in include/da_capi.h. The API is a flat C interface with no C++ exceptions or templates exposed, which makes it callable from C, C++, Go, and Rust via the standard foreign function interface without C++ name-mangling issues.
The README states that this C API powers the LocalAI backend. LocalAI is the project's own open-source AI engine that runs multiple model types (LLMs, vision, voice, image, video) without requiring a GPU. The depth-anything.cpp library slot into LocalAI the same way it would slot into any application that loads a C shared library at runtime.
The practical constraint is that the shared library build (-DDA_SHARED=ON) produces a statically linked ggml with PIC, so linking against it in a larger application requires care around symbol conflicts if that application already bundles another ggml version. The flat C API header is at include/da_capi.h and has no C++ exceptions or templates in its public interface.
Limitations and comparison with PyTorch-based inference
depth-anything.cpp runs a conversion step to produce a GGUF file from the official checkpoints. That conversion requires Python, torch, and safetensors. Teams who already have a Python environment for model development will find the conversion step reasonable; teams who want a completely Python-free pipeline from source model to inference must use pre-converted GGUF files from the HuggingFace repository.
The V2 Giant checkpoint is listed as gated and unreleased on HuggingFace. The README notes this specifically, and it means the largest V2 model is not available for conversion or inference through this tool.
Multi-view depth and pose, which requires providing multiple input images rather than a single frame, is only supported for the DA3 model family, not for V2 checkpoints. The output surface for V2 is depth only, without confidence, pose, sky mask, or point cloud.
For teams running a Python computer vision stack, the reference PyTorch implementation of Depth Anything 3 from bytedance-seed is the more direct path: it uses the same GPU kernels that the GGUF was verified against, it integrates naturally with other torchvision transformations, and it does not require a conversion step. The C++ port's advantage is size (99 MB quantized), load time (6.7x faster than PyTorch for the f32 model on the benchmark machine), and the elimination of Python from the deployment environment entirely.
Editorial conclusion
depth-anything.cpp is the right choice when Depth Anything inference must run in an environment without Python or PyTorch, when model file size matters (q4_k at 99 MB versus 516 MB for the PyTorch f32 checkpoint), or when load latency is a constraint. The C API makes it embeddable in Go, Rust, and C projects without any Python interop layer. Teams who are already running a Python-based computer vision stack and want the original Depth Anything 3 reference implementation should use the official PyTorch weights directly. The project is MIT-licensed and was last updated on 2026-09-07.
Frequently asked questions
What is the Depth Anything model?
Depth Anything is a monocular depth estimation model family from ByteDance that estimates per-pixel depth from a single image. Version 2 added metric depth variants trained for indoor and outdoor scenes. Version 3 (DA3) extended the output surface to include camera extrinsics, camera intrinsics, per-pixel confidence, a sky mask, and a back-projected 3D point cloud.
What are the key differences between Depth Anything V2 and Depth Anything V3?
According to the README, Depth Anything V2 produces depth only with no confidence, pose, or sky outputs, and the V2 Giant checkpoint is gated and unavailable. Depth Anything V3 adds per-pixel confidence, camera extrinsics and intrinsics, a sky mask, and a back-projected 3D point cloud, and all seven V3 model sizes are available for conversion to GGUF.
How does Depth Anything V3 work in the C++ port?
depth-anything.cpp loads a self-contained GGUF file that has all model dimensions, hyperparameters, and preprocessing constants baked in. At inference it runs the full DA3 forward pass using ggml tensor operations, producing depth, confidence, camera pose, and optionally 3D Gaussians, verified to match the PyTorch reference at correlation 1.0.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/localai-org-depth-anything-cpp)