ml-sharp: Apple's SHARP Model Turns One Photo Into 3D Gaussians in Under a Second
Sharp Monocular View Synthesis in Less Than a Second
At a glance
- What is it?
- SHARP is a research inference release from Apple that regresses a metric 3D Gaussian representation from a single image, then renders nearby views in real time. The CLI installs in two commands, but video rendering is CUDA-only and the licence is not a standard OSI one.
- Who is it for?
- Adopt ml-sharp if you want a feedforward single-image to 3DGS step inside a Python pipeline and you have a CUDA GPU for the video path. Do not adopt it if you need a permissively licensed model, a Windows-tested install, or a training recipe, because the repository ships inference code only and points to LICENSE_MODEL for the weights.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 19 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What SHARP actually solves, and who is standing in line for it
Most view synthesis pipelines need more than one photograph. You capture a set of images, run structure-from-motion, optimise a radiance field or a Gaussian set per scene, and wait. SHARP collapses that into a single feedforward pass: one photograph in, the parameters of a 3D Gaussian representation out, in less than a second on a standard GPU, according to the README. The representation is metric, with absolute scale, so camera moves can be expressed in real units rather than in arbitrary scene coordinates.
The people this is aimed at are not photographers. They are engineers building viewer features, robotics or AR prototypes that need a 3D scene from whatever image is already on hand, and researchers who want a baseline that runs fast enough to iterate on. The repository is the inference side of a paper, not a product. There is no training code in the top-level layout, no dataset loader for reproducing the paper's numbers, and no hosted demo. If your job is to reproduce the LPIPS and DISTS improvements the README quotes, this repository will not do it for you.
The mechanism: a single network pass that emits Gaussian parameters
The README describes the architecture in one sentence: given a single photograph, SHARP regresses the parameters of a 3D Gaussian representation of the depicted scene. That is the whole data flow. There is no per-scene optimisation loop, no multi-view bundle adjustment, no depth pre-pass that the documentation exposes. The neural network is the scene reconstruction.
What comes out is a set of 3D Gaussian splats written as .ply files into the output folder. Those files are the handoff point. The README states they are compatible with various public 3DGS renderers, which matters because it means SHARP does not lock you into its own viewer. The coordinate convention is OpenCV: x right, y down, z forward. The scene centre sits roughly at (0, 0, +z). Third-party renderers will need scaling and rotation to re-centre the scene, and the README says so plainly rather than hiding it.
Rendering is a separate stage from prediction. The `--render` option produces videos along a camera trajectory, and it uses the gsplat renderer, which the README notes takes a while to initialise on first launch. That initialisation cost is worth planning around if you are running short jobs.
Installing ml-sharp and running your first prediction
The README recommends creating a Python environment first, pinned to Python 3.13. That version is not incidental: the pyproject.toml sets pythonVersion to 3.13 under the pyright configuration, and the conda command in the README matches it.
conda create -n sharp python=3.13Activate that environment, then install the pinned dependency set. The requirements.txt is a compiled lockfile generated by uv from requirements.in, so it carries exact versions for torch, gsplat, timm and the rest rather than floating ranges.
pip install -r requirements.txtThe project installs a console script named `sharp`, declared in pyproject.toml as sharp.cli:main_cli. The README's installation check is simply to ask it for help. If the entry point is wired correctly you get a command listing rather than an import error.
sharp --helpTo run a prediction, point the CLI at a directory of input images and a directory for the output gaussians. The model checkpoint downloads on first run and caches at ~/.cache/torch/hub/checkpoints/, so the first invocation is slower than later ones.
sharp predict -i /path/to/input/images -o /path/to/output/gaussiansIf you would rather fetch the weights yourself, the README gives a direct URL, and a `-c` flag to point at the local file. Both forms appear in the documentation.
wget https://ml-site.cdn-apple.com/models/sharp/sharp_2572gikvuh.pt
sharp predict -i /path/to/input/images -o /path/to/output/gaussians -c sharp_2572gikvuh.ptAfter a successful run the output folder holds 3DGS .ply files. To go further and produce a video along a trajectory, add `--render` to the same predict call, or run the render subcommand against the intermediate gaussians. Both require a CUDA GPU.
sharp predict -i /path/to/input/images -o /path/to/output/gaussians --render
sharp render -i /path/to/output/gaussians -o /path/to/output/renderingsWhere ml-sharp stops: CUDA-only video, no training path, no Windows story
The clearest constraint is in the README itself. Gaussian prediction works on CPU, CUDA and MPS. Video rendering through `--render` currently requires a CUDA GPU. So an Apple Silicon user can produce splats but cannot render the trajectory video that the project page uses to demonstrate results. That is a real split in capability, not a footnote.
The second gap is training. The pyproject description calls this the inference, network and model code for the SHARP model. Nothing in the top-level layout suggests a training entry point, and the README's evaluation section does not offer scripts either. It directs you to the paper for quantitative and qualitative evaluation and to a qualitative examples page for video comparisons. If you want to fine-tune on your own captures, this repository does not document how.
The third is platform. Nothing in the README addresses Windows. The compiled requirements.txt does carry a colorama marker for sys_platform == 'win32' pulled in by click and tqdm, and CUDA wheels are marked for x86_64 Linux, but that is dependency metadata, not a support statement. Treat Windows as unverified.
Finally, the licence. The repository root contains both LICENSE and LICENSE_MODEL, and the README tells you to check both before using the code and the released models. The GitHub licence field reads NOASSERTION, which means GitHub could not classify it automatically. Read the files; do not assume Apache or MIT because the repository is published by Apple.
How SHARP differs from per-scene 3DGS optimisation and from multi-view pipelines
The obvious alternative is the standard 3D Gaussian splatting workflow: capture many views, run COLMAP or a similar reconstruction front end, then optimise Gaussian parameters per scene with gradient descent. That approach is far more general. It handles wide baselines, large scenes and arbitrary capture rigs, and it is the basis for most public 3DGS tooling and renderers. It is also slow per scene and needs the multi-view capture in the first place.
SHARP trades that generality for a single pass. The README frames the payoff as zero-shot generalisation across datasets, with reported reductions of 25 to 34 percent in LPIPS and 21 to 43 percent in DISTS versus the best prior model, and a synthesis time lower by three orders of magnitude. Those are the paper's claims, not measurements you can reproduce from this repository, since the evaluation scripts are not shipped. What you can verify locally is the speed shape: one forward pass instead of an optimisation loop.
A second alternative is the broader family of monocular depth and novel-view tools that estimate geometry and warp pixels. SHARP does not warp pixels. It produces an explicit Gaussian representation you can hand to a renderer, which is a different contract. The .ply output is the point of the design.
Maintenance, upgrade cost and what the licence files mean for you
The last push to the default branch was on 2026-09-11, which is recent, and the repository is not archived. There are no retrieved releases, so the version string in pyproject.toml, 0.1, is the only version marker available. Expect to track the main branch rather than a tagged artifact.
Upgrade cost lives mostly in the dependency set. The requirements.txt is a compiled lockfile with exact pins, including gsplat 1.5.3, torch-linked CUDA packages such as nvidia-cublas-cu12 12.8.4.1 for x86_64 Linux, and numpy 2.3.3. gsplat is the piece most likely to move under you: it pulls in ninja and jaxtyping and is the renderer used by `--render`. A gsplat upgrade is the change most likely to break video rendering, and the README already warns that the renderer is slow to initialise on first launch.
The licence situation is the one to resolve before anything else. Two files, LICENSE for the code and LICENSE_MODEL for the released models, and a NOASSERTION classification on the repository page. Model weights often carry terms that differ from the code, and the README's instruction to check both is the only guidance given. This is not legal advice; it is a reason to open both files before you ship anything that depends on the checkpoint at https://ml-site.cdn-apple.com/models/sharp/sharp_2572gikvuh.pt.
Editorial conclusion
Adopt ml-sharp if you want a feedforward single-image to 3DGS step inside a Python pipeline and you have a CUDA GPU for the video path. Do not adopt it if you need a permissively licensed model, a Windows-tested install, or a training recipe, because the repository ships inference code only and points to LICENSE_MODEL for the weights. Before you build on it, read LICENSE and LICENSE_MODEL, confirm your Python 3.13 environment can build gsplat with ninja, and check that the .ply output orients correctly in your renderer under the OpenCV convention.
Frequently asked questions
How do I install ml-sharp?
Create a Python 3.13 environment with conda, then run pip install -r requirements.txt from the repository root. The README's installation check is sharp --help, which exercises the console script declared in pyproject.toml.
How do I use ml-sharp to make a Gaussian splat?
Run sharp predict with an input image directory and an output directory, for example sharp predict -i /path/to/input/images -o /path/to/output/gaussians. The checkpoint downloads automatically on first run and caches at ~/.cache/torch/hub/checkpoints/, and the output folder receives 3DGS .ply files.
What is ml-sharp?
It is the inference code for SHARP, an Apple research model that regresses a metric 3D Gaussian representation of a scene from a single photograph in less than a second on a standard GPU. The resulting splats can be rendered in real time for nearby views.
What is Apple ml-sharp?
It is Apple's published repository for SHARP, accompanying the paper Sharp Monocular View Synthesis in Less Than a Second. It contains the inference, network and model code, and the README directs readers to the paper and a qualitative examples page for evaluation.
How can I convert a video to a Gaussian splat with ml-sharp?
The README does not document a video input path. The CLI takes a directory of input images via sharp predict -i, and the rendering options produce video output from gaussians rather than the reverse.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/apple-aiml-research-ml-sharp)