# MoGe: Monocular 3D Geometry from a Single Image

> MoGe is a Microsoft Research model that estimates metric depth maps, point maps, normal maps, and camera field of view from a single image. Three published versions cover different accuracy and parameter-count trade-offs, with the newest MoGe-3 adding fine-grained volumetric refinement and requiring a Linux or Windows GPU environment because its Triton dependency publishes no macOS wheels.

**microsoft/MoGe** — [CVPR'25 Oral] MoGe: Unlocking Accurate Monocular Geometry Estimation for Open-Domain Images with Optimal Training Supervision

- Repository: https://github.com/microsoft/MoGe
- Website: https://wangrc.site/MoGePage/
- Stars: 2,992 · Forks: 241
- Language: JavaScript
- License: NOASSERTION
- Published: 2026-09-24 · Updated: 2026-09-24 · Language: en
- Canonical page: https://hysenlabs.com/projects/microsoft-moge

## What MoGe Recovers from a Single Image

Standard depth estimation returns relative depth: a scaled version of the scene geometry that does not correspond to real-world distances. MoGe, starting with MoGe-2, returns metric depth, meaning the distances are in actual units. Beyond depth, a single forward pass produces a metric point map (the full 3D position of each pixel), a normal map, and an estimate of the camera field of view. Providing the true FOV as input improves accuracy further.

The README states that MoGe handles resolutions from 2:1 to 1:2 aspect ratios. The ViT-L variant of MoGe-2 achieves 60ms latency per image on an A100 or RTX 3090 in FP16, according to the README. This makes it suitable for preprocessing pipelines that need 3D-aware representations from large image collections without a calibrated camera setup.

## Three Generations: MoGe-1, MoGe-2, and MoGe-3

The repository covers three published model generations. MoGe-1 is a 314-parameter ViT-L model that estimates point maps but not metric scale or normal maps. MoGe-2 adds metric scale and normals. It is available in four variants: a 326M ViT-L without normals, a 331M ViT-L with normals, a 104M ViT-B with normals, and a 35M ViT-S with normals. The smallest variant is a practical choice for constrained environments.

MoGe-3, released on 2026-08-18 according to the README, targets fine-grained point map geometry using self-guided sparse volumetric refinement. It is available in two sizes: a 1.25B ViT-G model and a 370M ViT-L model. Both include metric scale and normal estimation. MoGe-3 is the current recommended version for the best geometric detail, but it requires more GPU memory than the MoGe-2 variants and does not run on macOS.

All pretrained weights are on Hugging Face Hub under the `Ruicheng` namespace. The model IDs as given in the README are `Ruicheng/moge-vitl` for MoGe-1, `Ruicheng/moge-2-vitl` through `Ruicheng/moge-2-vits-normal` for MoGe-2, and `Ruicheng/moge-3-vitg` and `Ruicheng/moge-3-vitl` for MoGe-3. Weights download automatically when you call `MoGeModel.from_pretrained`.

## Installing MoGe with uv or pip

The README recommends uv for installation and states that Python 3.10 or newer is required. The simplest setup clones the repository and syncs:

```bash
git clone https://github.com/microsoft/MoGe.git
cd MoGe
uv sync                             # inference, all model versions
```

This creates a `.venv/` directory and installs MoGe in editable mode. After syncing, prefix every command with `uv run`, for example `uv run moge infer ...`, or activate the environment with `source .venv/bin/activate`. If you intend to edit the code, you can also install with pip from a clone:

```bash
git clone https://github.com/microsoft/MoGe.git
cd MoGe
pip install -e .
```

The pyproject.toml pins the default PyTorch wheel to the CUDA 13.0 index. To target CUDA 12.8 instead, the README gives this example:

```bash
uv pip install --torch-backend=cu128 torch torchvision --reinstall
```

For pip, the index must be passed explicitly:

```bash
pip install -e . --index-url https://download.pytorch.org/whl/cu130
```

For a pip install from the repository root without cloning:

```bash
pip install git+https://github.com/microsoft/MoGe.git
```

The README notes that macOS is not supported. MoGe-3 depends on FlexGEMM, which builds on Triton, and Triton publishes no macOS wheels. MoGe-1 and MoGe-2 may still be usable on macOS in theory, but the project does not test or document that path.

## Loading a Model and Running Inference

The README provides a minimal Python example for loading the model and running it on an image. To use MoGe-3, the import is:

```python
from moge.model.v3 import MoGeModel
```

The model loads from a local checkpoint or a Hugging Face model ID:

```python
model = MoGeModel.from_pretrained("PATH_TO_CKPT.pt").to(device)
```

The input image is read with OpenCV, converted to RGB, and normalized to the [0, 1] range before being passed to the model. The model object and the Hugging Face loading pattern are the same across all three versions; only the import path changes (`v1`, `v2`, or `v3`). Optional ONNX export is documented in `docs/onnx.md` for deployments that cannot use PyTorch at inference time.

## Where MoGe Is the Wrong Choice

MoGe estimates geometry from a single image, which means it has no ground-truth depth to anchor its predictions. For scenes with unusual geometry not well represented in its training data, the metric scale estimate may be inaccurate. The README does not describe the training distribution in detail beyond calling the task open-domain, so the practical boundary of where the metric scale holds is not documented.

MoGe-3 uses FlexGEMM, which means it is tied to Triton and therefore to Linux or Windows with CUDA. Teams using Apple Silicon for local development cannot run MoGe-3, not as a configuration choice but as a hard dependency failure. The README states this directly. This creates a gap between local development and production environments that teams with macOS developers and Linux servers will need to manage.

For video geometry estimation, where temporal consistency matters, MoGe operates frame by frame and does not model motion or inter-frame consistency. Applications that need temporally coherent 3D reconstruction would need to handle that externally.

## MoGe Compared to Depth Estimation Approaches Like Depth Anything

Depth Anything is a monocular depth estimation model trained on a large mixed dataset that returns relative (affine-invariant) depth. MoGe differs in two key ways documented in the project. First, MoGe-2 and MoGe-3 produce metric depth rather than relative depth. Second, a single MoGe forward pass also returns point maps and normal maps alongside depth, which Depth Anything does not provide.

The architectural difference is that MoGe uses a ViT backbone with a geometry-specific training objective, not the same affine-invariant scale-and-shift depth supervision used by relative depth models. The README references a publication at CVPR 2025 (Oral) for the underlying methodology. Engineers who need relative depth for tasks like monocular novel view synthesis, where metric scale is not required, may find Depth Anything simpler to integrate. Engineers who need metric 3D point clouds from in-the-wild images are the primary target for MoGe.

## Maintenance, License, and PyTorch Version Coupling

The repository is under the MIT license (as declared in pyproject.toml) with an active push history: the last push was on 2026-09-09 and MoGe-3 was released on 2026-08-18. The project is actively developed with three model generations published. The CHANGELOG.md tracks changes and the repository has no GitHub release tags, meaning version tracking is through commit history and pyproject.toml version fields.

One maintenance consideration is the tight coupling to specific CUDA wheel indices. The pyproject.toml pins to CUDA 13.0 by default, and the uv environment locks to Linux and Windows platforms explicitly. Upgrading to a new CUDA version requires editing the index URL in pyproject.toml or running reinstall commands. The dependency list also includes several git-pinned packages, including `utils3d_moge`, `pipeline`, and `flex-gemm`, each locked to specific commit hashes. These pins prevent dependency drift but mean that any bug fix in those libraries requires an update to the MoGe pyproject.toml to pick them up.

## Conclusion

MoGe is the right tool for engineers who need metric 3D geometry from arbitrary in-the-wild images without a stereo rig or depth sensor, and who are working on Linux or Windows with CUDA hardware. The progression from MoGe-1 to MoGe-3 lets teams choose between a smaller 314M-parameter model for throughput and a 1.25B-parameter ViT-G model for fine detail. macOS users cannot run MoGe-3 because FlexGEMM depends on Triton, which has no macOS wheels. Before starting, confirm your CUDA version against the index URL in pyproject.toml and choose the matching PyTorch wheel, since the default pins CUDA 13.0.

## FAQ

### What is MoGe and what does it produce from an image?

MoGe is a model from Microsoft that estimates metric 3D point maps, metric depth maps, surface normal maps, and camera field of view from a single image in one forward pass. It is designed for open-domain images and does not require a stereo setup or depth sensor.

### Does MoGe run on macOS?

MoGe-3 does not run on macOS. The README states that MoGe-3 depends on FlexGEMM, which builds on Triton, and Triton publishes no macOS wheels. Earlier versions (MoGe-1, MoGe-2) do not have this stated restriction, but macOS support is not documented or tested.

### How do I choose between MoGe-1, MoGe-2, and MoGe-3?

MoGe-1 (314M parameters) does not provide metric scale or normals. MoGe-2 adds metric scale and normals and is available in sizes from 35M to 331M parameters, making it the right choice for constrained environments or when macOS support is needed. MoGe-3 (370M or 1.25B parameters) provides finer point map detail but requires Linux or Windows with CUDA and more GPU memory.

## Sources

- [Issues](https://github.com/microsoft/MoGe/issues)
- [microsoft/MoGe on GitHub](https://github.com/microsoft/MoGe)
- [Project website](https://wangrc.site/MoGePage/)
- [README](https://github.com/microsoft/MoGe/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/microsoft-moge
