3d-ken-burns needs CuPy and a CUDA toolchain, and its depth script stops one step short of the pipeline
an implementation of 3D Ken Burns Effect from a Single Image using PyTorch
At a glance
- What is it?
- sniklaus/3d-ken-burns is the PyTorch reference implementation of the 3D Ken Burns camera effect on a single still image, with an automatic script, a Flask interface and two depth benchmarks. Its dataset is CC BY-NC-SA and non-commercial, the tree carries no release and no CI directory, and the license field asserts nothing.
- Who is it for?
- Read this as a paper reproduction rather than a tool to install. The code is a reference implementation whose depth adjustment step is deliberately left to the reader, whose dataset carries a non-commercial clause, and whose commit history stops at 2026-06-01 with no tagged release, so if you need a maintained camera animation library this is not it, and if you do reuse the code or the data, settle the license question first.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 130 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 9, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Several functions live in CUDA, so CUDA_HOME is part of setup
The setup section is three sentences and they are all constraints. Several functions are implemented in CUDA using CuPy, which is why CuPy is a required dependency. It can be installed with `pip install cupy` or through one of the binary packages the CuPy project provides. Then the environment variable `CUDA_HOME` has to be configured, and for anything that writes video output, moviepy has to be installed as well.
So there are three separate things to get right: a Python array library that compiles and runs GPU kernels, a CUDA toolkit the runtime can find, and a video writer for the final file. A machine with a GPU and a working PyTorch install is not by itself enough.
requirements.txt sets floors and no ceilings for all of it: cupy>=5.0.0, flask>=1.0.0, gevent>=1.3.0, moviepy>=1.0.0, numpy>=1.15.0, opencv-contrib-python>=3.4.0, scipy>=1.1.0, torch>=1.6.0, torchvision>=0.7.0. The oldest of those floors is Flask 1.0.0, a release from 2018, so nothing in the file records which version the code was last exercised against, and nothing stops a modern resolution from installing something the interface code predates. The tree also carries common.cuda next to common.py, and the setup text never says how that file is compiled, only that CuPy and CUDA_HOME must be present.
Three scripts, three outputs, and one page served on port 8080
The usage section names four things you can run and they are all one-liners. The automatic path animates an image and writes a video file:
python autozoom.py --in ./images/doublestrike.jpg --out ./autozoom.mp4The depth estimate on its own takes the same shape of arguments and writes a NumPy array instead of a video:
python depthestim.py --in ./images/doublestrike.jpg --out ./depthestim.npyThe interactive path starts a web server rather than processing anything from the command line:
python interface.pyThat one opens http://localhost:8080/ in the browser, and the image is loaded with a button in the bottom right corner of the page. The interface is the only documented way to adjust a camera path by hand, and the two other scripts expose no camera parameters on the command line at all, so there is no documented way to set a focal length, a zoom range or a motion path without going through the page or writing your own arguments.
The two benchmark scripts are separate again: `python benchmark-ibims.py` and `python benchmark-nyu.py`. The first argument pair across all of these is in and out, which makes the scripts easy to chain, and the images directory at the root supplies the doublestrike example.
The depth script deliberately stops before the adjustment
The depth command carries a warning in the same paragraph that introduces it: the script does not perform the depth adjustment, and issue 22 in the repository holds the information on how to add it. That single sentence explains the shape of the whole project.
A 3D Ken Burns effect needs a depth map so the virtual camera can move in front of and behind parts of the scene, producing parallax. The raw depth estimate is the input to that, not the output. What you get from depthestim.py is a .npy file containing the raw estimate, and the step that turns a raw estimate into the adjusted depth the camera math expects is the part the author left to the reader and documented in an issue rather than in the code path of the script.
That is a reasonable line to draw in a reference implementation that accompanies a paper, and the paper is cited at the top of the file with a request to cite it in turn. It also means the two entry points are not interchangeable. autozoom.py is the complete pipeline and hides the adjustment inside; depthestim.py is a diagnostic tool that produces the raw tensor and stops. Anyone wanting to reproduce the paper's depth numbers from the raw file needs that issue's instructions, and anyone wanting to change how depth is estimated has to add the adjustment themselves before the output is comparable.
The interface is a Flask page and the file documents no access control
The interface is two files at the root, interface.py and interface.html, with flask and gevent in requirements.txt. The page is reached on localhost:8080, the image is chosen with a button in the bottom right corner, and then the file asks for patience twice: be patient when loading an image and be patient when saving the result, because there is a bit of background processing going on.
That is the whole description of the failure surface. Nothing states what the background processing is, how long it takes, whether a request has a timeout, or what happens if the image is large or if the GPU is busy, so the practical answer is that the interface gives no progress signal and no error text. Nothing describes an authentication step, a permission model, or how much the server accepts from disk, and the file does not say which address the port is bound to. The URL given is localhost, which is what a local development server binds by default, but that is a default rather than a documented restriction.
The security surface is worth naming as a fact rather than a finding. This is a page that accepts an image path, runs a CUDA pipeline over it and returns a rendered result, driven by whichever address the socket takes. If you expose port 8080 beyond your own machine, the only control described anywhere in the repository is the port number itself.
Two benchmark scripts named after depth datasets, with no expected numbers
The benchmarking line is short: run `python benchmark-ibims.py` or `python benchmark-nyu.py`, and use them to easily verify that the provided implementation runs as expected. The two script names point at the two depth datasets the estimation code is usually measured against, and their presence tells you the author expected people to check the depth estimation against external references.
What the repository does not contain is any number. There is no table of results, no expected error threshold, no reference metrics for either dataset, and no statement of which model checkpoint the scripts load. So the scripts can tell you whether the pipeline executes and roughly what shape its output is, and they cannot tell you whether the depth estimate is any good. A passing run and a correct run are indistinguishable from the outside.
That gap matters more here than in a project with a model card, because the paper this implements is a 2019 preprint, identifier 1909.05483, and the reference values most readers would compare against are the ones in that paper. The README asks you to cite the paper if you make use of the work, and it points at a related implementation by another author and a Hacker News discussion thread, but it does not restate any of the reported figures, so there is nothing in the repository to diff your output against without going back to the paper.
The dataset is 110 GB of non-commercial material on an object store
The dataset section states the terms before the table: it is licensed under Creative Commons Attribution-NonCommercial-ShareAlike 4.0, may only be used for non-commercial purposes, and the LICENSE file holds the detail. That is a research dataset with a share-alike condition and a commercial prohibition attached, which is the single most consequential line in the README for anyone planning to build on it.
The table itself is organised as scene, mode, color, depth, normal, with two modes per scene, flying and walking, and three archives per row. The scenes visible are asdf, blank, chill, city, environment, fort and grass. One of those names, asdf, is a keyboard mash rather than a place, which suggests a placeholder that was never replaced. The sizes vary by an order of magnitude: city is the smallest at 0.7 to 0.8 GB per mode, while chill and fort carry normal maps of 9.2 GB and 10.8 GB, since a normal map has three channels per texel against one for depth. The fourteen visible rows add up to a little over 110 GB of downloads.
Everything is hosted on a Netbright object storage bucket under the name sniklaus-kenburns, as direct zip links with no checksums, no manifest and no instructions for verifying an archive. The table ends partway through the grass-walking row, with the link for its normal archive cut off after the host name, so the last entry visible has no usable address.
No tagged release, no CI directory, and a license field that says nothing
The repository has no GitHub releases, and nothing in the tree carries a version number, so there is no artifact to install and no tag to pin a fork to. The last commit is dated 2026-06-01. The root is a flat list of scripts and support files: LICENSE, README.md, autozoom.py, benchmark-ibims.py, benchmark-nyu.py, common.cuda, common.py, depthestim.py, images/, interface.html, interface.py, models/ and requirements.txt. The models directory holds the network weights the estimation and generation scripts load, so a clone is not enough to run anything without it.
There is no .github directory at the root, which means no workflow configuration is visible in the tree, and no test runner is named anywhere in the setup or usage text. The closest thing to a verification step is the pair of benchmark scripts.
The license position needs stating without resolving it. The license field on the repository asserts nothing, a root LICENSE file exists, and the dataset section separately places the data under CC BY-NC-SA 4.0 with a non-commercial restriction and points at that same LICENSE file for detail. Nothing says whether the code carries those same terms or a different set. That question has to be answered by reading the LICENSE file and asking the author, not by assuming the dataset clause reaches the Python, because the non-commercial term in particular changes whether any of this can be used in a product.
Editorial conclusion
Read this as a paper reproduction rather than a tool to install. The code is a reference implementation whose depth adjustment step is deliberately left to the reader, whose dataset carries a non-commercial clause, and whose commit history stops at 2026-06-01 with no tagged release, so if you need a maintained camera animation library this is not it, and if you do reuse the code or the data, settle the license question first.
Frequently asked questions
What does the depthestim.py script in 3d-ken-burns output?
It writes a raw depth estimate as a .npy file and explicitly does not perform the depth adjustment; issue 22 holds the information on how to add that step. The adjustment is inside the automatic path instead.
What does 3d-ken-burns need installed before it will run?
CuPy, because several functions are implemented in CUDA, a configured CUDA_HOME environment variable, and moviepy for video output. requirements.txt also floors torch at 1.6.0, torchvision at 0.7.0, opencv-contrib-python at 3.4.0, numpy at 1.15.0, scipy at 1.1.0, flask at 1.0.0 and gevent at 1.3.0.
Can I use the 3d-ken-burns dataset commercially?
No. The dataset is licensed under Creative Commons Attribution-NonCommercial-ShareAlike 4.0 and the README states it may only be used for non-commercial purposes, with the LICENSE file holding the detail.
How do I run the 3d-ken-burns web interface?
Start it with `python interface.py` and open http://localhost:8080/, then load an image with the button in the bottom right corner. The file asks for patience while loading and saving because there is background processing, and describes no access control on the port.
Are there published results for the 3d-ken-burns depth estimation?
Not in the repository. There are two scripts, benchmark-ibims.py and benchmark-nyu.py, described as a way to verify that the implementation runs as expected, but no expected metrics, thresholds or reference numbers are given anywhere in the file.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/sniklaus-3d-ken-burns)