LHM: single-image animatable human reconstruction you can run yourself
[ICCV2025] LHM: Large Animatable Human Reconstruction Model from a Single Image in Seconds
At a glance
- What is it?
- LHM turns one photo into an animatable 3D human, and the repository ships several model sizes with different GPU and input requirements. Here is what the code actually asks of you, and where it stops.
- Who is it for?
- LHM is worth adopting if you have a CUDA 11.8 or 12.1 machine with a recent NVIDIA GPU, need an animatable human avatar rather than a static mesh, and can live with the repository's own model download behaviour. It is the wrong tool if you have no NVIDIA hardware, if you need training code or the promised training and testing data, or if you need a documented rollback path when a model download changes under you.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Activity is slowing. The repository last received commits 6 months ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What LHM reconstructs, and who the repository is aimed at
LHM stands for Large Animatable Human Reconstruction Model, and the repository describes it as producing an animatable human from a single image in seconds. The distinction that matters is the word animatable. A photogrammetry or mesh-generation pipeline gives you geometry; LHM's stated output is a human that can then be driven, which is why the repository ships motion extraction and animation inference nodes alongside the core reconstruction pipeline. That places it in the digital human and AIGC tooling space rather than in general 3D scanning.
The intended user is a developer or technical artist with an NVIDIA GPU who wants a human avatar without a multi-camera rig or a capture stage. The README points to a HuggingFace Space and a ModelScope studio, so a first look needs no installation at all. The repository itself is the path for anyone who wants the model locally, wants to batch inputs, or wants to integrate the pipeline into a larger application. The training data and testing data are listed as an unchecked TODO item, and the training codes are also unreleased, so this is an inference repository, not a research reproduction package.
How the pipeline is put together: checkpoints, layers and the two apps
The architecture visible from the repository is a set of pretrained checkpoints with different capacities, plus a Gradio front end. The model table lists five weights: LHM-MINI, LHM-500M, LHM-500M-HF, LHM-1.0B and LHM-1B-HF. The column that separates them most clearly is BH-T Layers, which runs 2, 5, 5, 15 and 15 respectively. The README also gives an inference time per model, from 1.41 s for LHM-MINI up to 6.57 s for the 1B variants. Those figures come from the project's own table; treat them as the authors' numbers on their hardware, not as a promise for yours.
The input requirement column is the practical constraint. LHM-500M and LHM-1.0B are marked full body only. LHM-MINI, LHM-500M-HF and LHM-1B-HF accept half and full body. The HF suffix therefore is not just a different hosting location; it changes what the model will accept. If you are building a photo-booth style application where users frame themselves from the chest up, the non-HF 500M and 1B checkpoints are the wrong choice regardless of quality.
At the top level the repository holds two entry points, app.py and app_motion.py, with app_motion_ms.py as a ModelScope variant. There is an engine/ directory, a configs/ directory, and separate shell scripts for inference and mesh export (inference.sh and inference_mesh.sh). The presence of both an app and a motion app matches the README's description of motion extraction and animation inference as distinct stages: extract motion once, then reuse it. The README states that with an extracted offline motion, a 10s animation clip can be generated in 20s, again as the project's own figure.
Installing LHM on Linux or Windows and running the first reconstruction
The README gives three installation routes. The quickest on Linux is the prebuilt Docker image, which is documented for Linux only and pinned to CUDA 121. It downloads a tar archive, loads it, and runs a container that publishes port 7860.
wget -P ./lhm_cuda_dockers https://virutalbuy-public.oss-cn-hangzhou.aliyuncs.com/share/aigc3d/data/for_lingteng/LHM/LHM_Docker/lhm_cuda121.tar
sudo docker load -i ./lhm_cuda_dockers/lhm_cuda121.tar
sudo docker run -p 7860:7860 -v PATH/FOLDER:DOCKER_WORKSPACES -it lhm:cuda_121 /bin/bashThe README notes that nvidia-docker must already be installed. Replace PATH/FOLDER with a host directory you want visible inside the container. Once the container is up, the Gradio interface is reachable on port 7860.
For a native install, clone the repository and pick the script that matches your CUDA version. The README states the installation has been tested with python3.10, CUDA 11.8 or CUDA 12.1.
git clone [email protected]:aigc3d/LHM.git
cd LHM
sh ./install_cu121.sh
python ./app.pySwap in `sh ./install_cu118.sh` for the CUDA 11.8 path. The Windows route uses a batch file instead, and the README shows it inside a virtual environment created with `python -m venv lhm_env` and activated with `lhm_env\Scripts\activate`, followed by `install_cu121.bat` and then `python ./app.py`. The README also points to INSTALL.md for a step-by-step dependency install, and to community videos on YouTube and bilibili that walk through both LHM and the ComfyUI branch.
Weights are the part to watch. The README states in red text that the model will be downloaded automatically if you do not download it yourself. If you prefer to control the location, the documented HuggingFace route uses the huggingface_hub client:
from huggingface_hub import snapshot_download
model_dir = snapshot_download(repo_id='3DAIGC/LHM-MINI', cache_dir='./pretrained_models/huggingface')The same call works with repo_id `3DAIGC/LHM-500M-HF` and `3DAIGC/LHM-1B-HF`, which are the two variants the README shows for the HF models. ModelScope mirrors exist for all five checkpoints under the Damo_XR_Lab namespace.
GPU memory, model size and the LHM-MINI trade-off
The README's update log is the clearest statement of the hardware floor, and it moved over time. In April 2025 the project released LHM-MINI, described as allowing LHM to run on 16 GB GPUs. Days later it released a memory-saving version of motion and LHM that it says lets the entire pipeline run on 14 GB GPUs. The repository does not give a single consolidated VRAM table, so the honest reading is that the floor depends on which checkpoint and which stage you run, and that 14 GB is the lowest figure the project claims for the full pipeline.
LHM-MINI is where the trade-off bites. It has 2 BH-T layers against 5 for the 500M models and 15 for the 1B models, and it is the only non-HF checkpoint that accepts half body. So the smallest model is also the most flexible on input framing. That combination is unusual and it is the one I would reach for first on a constrained machine: you get half-body support and the lowest memory footprint, at the cost of capacity. Whether the quality drop matters depends entirely on your output resolution and how close the camera gets.
A second, less obvious constraint is dependency weight. The requirements file pins torch==2.3.0, torchvision==0.18.0 and xformers==0.0.26.post1, along with gsplat==1.4.0, diffusers==0.32.0, transformers==4.41.2 and a long tail of graphics libraries including open3d, pyrender, trimesh and xatlas. Several of these are compiled extensions. The install scripts exist precisely because installing that set by hand is unpleasant, and the CUDA 11.8 and 12.1 split exists because the compiled wheels differ between them. This is not a project you can drop into an arbitrary Python environment and expect to work.
Where LHM is the wrong tool
The most concrete limitation is in the repository's own TODO list. Training data, testing data and training codes are all unreleased. If your goal is to fine-tune on your own subjects, or to reproduce the paper's numbers, LHM does not give you the material to do it. The README says the data release is "License Available", which suggests a separate arrangement rather than an open download, but the repository does not document that process.
The second limitation is the automatic weight download. It is convenient and it is also a moving part you do not control: the README does not document how to pin a specific revision, and it does not describe a rollback path if an upstream checkpoint changes. For a research demo that is fine. For a production service where a silent model change alters output, you want to pre-download with snapshot_download into a directory you own and treat that directory as the artifact.
The third is scope. LHM reconstructs a human. It is not a general object reconstruction tool, it does not do scene capture, and the input requirement column means a cropped or partial-body photo may simply be rejected by the larger checkpoints. If your images are not full-body, you are choosing between LHM-MINI and the two HF variants, not the whole table.
Finally, there is no NVIDIA-free path. The Docker image is built for CUDA 121, the install scripts are named after CUDA versions, and the dependency list includes xformers and gsplat. CPU-only or AMD users should stop at the HuggingFace Space.
LHM against the ComfyUI route and the LHM++ successor
The most useful comparison is not against another reconstruction project but against the project's own branches. The README documents a ComfyUI integration on a separate branch, with dedicated motion extraction and animation inference nodes. The difference in approach is workflow versus application. The main branch gives you app.py, a Gradio interface and shell scripts, which suits batch jobs and scripted pipelines. The ComfyUI branch gives you node graphs, which suits iterative art direction where you want to swap inputs and re-run a stage without touching code. The README states that with an extracted offline motion you can generate a 10s animation clip in 20s on that branch, and links a step-by-step Windows 11 install guide for it.
There is also a successor. The March 2026 update announces that LHM++ is open-sourced, described as supporting arbitrary view inputs with higher efficiency, with 8-view input running on just 8GB GPU memory, and improved rendering quality. It lives in a separate repository, aigc3d/LHM-plusplus, with its own arXiv paper. This matters for anyone starting today: the 8GB figure in the LHM++ announcement is lower than anything LHM itself claims, and arbitrary view input is a capability LHM's single-image framing does not have. LHM remains the reference implementation for the ICCV 2025 paper, but if your constraint is memory or multi-view input, the newer repository is where the project's own numbers point.
Licence, maintenance and what an upgrade costs
LHM is released under Apache-2.0, and the repository carries a LICENSE file at the top level. That is a permissive licence, which generally means you can use, modify and redistribute the code, including commercially, provided you keep the licence and notices intact. It says nothing about the model weights, which are distributed separately through HuggingFace and ModelScope under the Damo_XR_Lab and 3DAIGC namespaces, and the README does not state their terms. If you plan to ship something built on these checkpoints, read the model cards on those hosting pages rather than assuming the code licence covers the weights. This is a description of what the repository states, not legal advice.
On maintenance: the repository is not archived, and the last push was on 2026-03-17. The update log shows a burst of activity in April 2025 (LHM-MINI, the memory-saving release, the ComfyUI nodes, the Windows tutorial, the LHM_Track video pipeline) and then the March 2026 entry pointing at LHM++. The pattern suggests the main branch is in a stable state while new work happens in the successor repository.
Upgrade cost is dominated by the pinned dependency set. torch 2.3.0, xformers 0.0.26.post1, gsplat 1.4.0 and the rest are version-locked, and moving any of them means rebuilding compiled extensions. The CUDA 11.8 and 12.1 scripts are separate for that reason. The realistic upgrade path is not to bump individual pins but to move to a newer checkpoint or to LHM++, and that decision is about VRAM and input framing rather than about code changes.
Editorial conclusion
LHM is worth adopting if you have a CUDA 11.8 or 12.1 machine with a recent NVIDIA GPU, need an animatable human avatar rather than a static mesh, and can live with the repository's own model download behaviour. It is the wrong tool if you have no NVIDIA hardware, if you need training code or the promised training and testing data, or if you need a documented rollback path when a model download changes under you. Verify first that your image satisfies the input requirement of the checkpoint you pick: LHM-500M and LHM-1.0B are full body only, while LHM-MINI, LHM-500M-HF and LHM-1B-HF accept half and full body. Then run the Gradio app on port 7860 and confirm the automatic weight download lands where you expect before wiring it into anything.
Frequently asked questions
What is LHM++?
LHM++ is a separate open-source successor to LHM, announced in the LHM repository's March 2026 update. The announcement states it supports arbitrary view inputs with higher efficiency and improved rendering quality, with 8-view input running on 8GB of GPU memory, and points to the aigc3d/LHM-plusplus repository and its own arXiv paper.
Which LHM model should I download for a half-body photo?
The README's model table marks LHM-500M and LHM-1.0B as full body only, while LHM-MINI, LHM-500M-HF and LHM-1B-HF are listed as accepting half and full body. So a half-body input points at one of those three, and LHM-MINI is also the smallest at 2 BH-T layers.
How much GPU memory does LHM need?
The README does not give a single consolidated VRAM figure. Its update log states that LHM-MINI allows running LHM on 16 GB GPUs, and that a later memory-saving release of motion and LHM lets the entire pipeline run on 14 GB GPUs.
Does LHM come with training code or training data?
No. The TODO list in the README shows the core inference pipeline, HuggingFace demo, ModelScope deployment and motion processing scripts as done, while releasing the training data and testing data and releasing the training codes remain unchecked.
What licence does LHM use?
The repository is released under Apache-2.0 and includes a LICENSE file at the top level. The model weights are hosted separately on HuggingFace and ModelScope, and the README does not state the terms for those weights.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/aigc3d-lhm)