TerraTorch: fine-tuning geospatial foundation models, and the licence you still own
A Python toolkit for fine-tuning Geospatial Foundation Models (GFMs).
At a glance
- What is it?
- A PyTorch Lightning and TorchGeo library with trainers for segmentation, classification and pixel-wise regression, wired to a dozen backbones and a dozen optional extras. It hosts no models at all, which is both its cleanest design decision and the one that leaves licence checking entirely with you.
- Who is it for?
- TerraTorch fits a geospatial team that has a pretrained backbone from the Hugging Face collection and needs the boring half of fine-tuning, namely trainers, data modules and configuration, without writing them. Skip it if you want hosted weights or a single supported stack, because the backbone list is deliberately wide and the optional extras are where the version conflicts live.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 4 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.
Editorial analysis
A framework, not a model hub
The disclaimer is the most clarifying paragraph in this readme, and it is worth reading before anything else.
TerraTorch hosts no models. It provides the training and inference framework only. Model weights come from elsewhere, from the Hugging Face organisation the project links, from the general timp model collection, or from your own checkpoints.
That choice pushes one responsibility onto the user, and the project says so directly: it is up to you to check whether the licence of any model you download, fine-tune or deploy permits your intended use. The maintainers add that they give no legal advice and are not liable for misuse of third-party models.
For a foundation-model toolkit in 2026 that is an unusually clean position to take, because most of them bundle a curated model zoo and inherit a licensing surface they cannot control. The trade is real: you get flexibility and you get the paperwork.
One related tool lives elsewhere too. Hyperparameter optimisation and architecture search are pointed at a separate repository rather than folded into this one, so the split between framework and search tooling is deliberate rather than accidental.
One version string, three releases in four seconds, and a v prefix
Small packaging details here are worth a minute, because they are the kind of thing that costs an afternoon.
The release history shows three patch releases on the same day, seconds apart: 1.2.11, then 1.2.12, then 1.2.13. That pattern suggests a batch publish, probably a dependency bump replayed through a release workflow, rather than three separate fixes. The last push to the default branch was later than all of them, so development continues well past the newest tag.
In the project metadata the version field is written with a leading letter v, matching the tag naming. Most tooling expects a plain version number in that field, so anything that reads the installed distribution metadata will see an unusual string.
The Python requirement is another one to note. The metadata asks for 3.10 or newer, while the classifier list advertises only 3.11, 3.12 and 3.13, and there are commented-out classifier lines in the file where a 3.10 entry used to be. So the floor and the advertised support disagree by one release.
None of this stops you installing, and all of it matters if you are pinning in a reproducible environment where a metadata parser is in the loop.
Trainers for three tasks and factories that mix backbones with decoders
The library's scope is best read as two lists: what it does for you, and what it will let you plug in.
On the first side there are flexible trainers for image segmentation, classification and pixel-wise regression, so three families of task rather than one. There are model factories for combining a backbone with a decoder. There are ready-to-go datasets and data modules whose selling point is that you point them at your data instead of writing a custom class, which is the single largest amount of glue code in a geospatial project. And tasks launch from the command line with configuration files, or from notebooks, so the same run can be debugged interactively and scheduled properly.
On the plug-in side, the dedicated geospatial backbones are named individually, including Prithvi, TerraMind, SatMAE, ScaleMAE and Clay. Several more come through TorchGeo, including Satlas, DOFA and the self-supervised SSL4EO models. Any backbone available in the general timp collection works too, and decoders come from the PyTorch segmentation models package or from MMSegmentation.
On datasets, the claim is that every benchmark in the GEO-Bench suite and every TorchGeo dataset and data module is available. That is the part of the stack you are least likely to write yourself.
Eleven optional extras, and one of them needs compiled CUDA operators
Optional dependencies are the shape of this project's complexity, and there are more of them than the headline suggests.
There is support for a local inference server, parameter-efficient fine-tuning, weights and biases logging, MMSegmentation, Surya, Tortilla-format file loading for datasets, visualization tooling, a second version of the benchmark suite, and weather foundation models, which are restricted to Python 3.11 and newer. Several can be requested in one command with a comma-separated list.
Two extras behave differently from the rest. The object-detection integration depends on an upstream package and is installed with a named extra, and the deformable variant of DETR needs work that pip cannot do for you: it requires compiling CUDA operators for the multi-scale deformable attention module, which means cloning a third-party repository, entering its operators directory and running a build script.
git clone https://github.com/fundamentalvision/Deformable-DETR.git
cd Deformable-DETR/models/ops
sh make.shThe documented requirements for that build are Linux, a CUDA toolkit of 9.2 or newer and GCC 5.4 or newer.
What matters here is the graceful degradation. Core DETR functionality works without that extension, and the tests that need it are skipped rather than failing when the operators are absent. That is the right default: the fast path stays fast and the expensive path is opt-in.
One more asymmetry to know about: the contributor install path is an editable install with the test and object-detection extras together, so a developer environment is not quite the same shape as a user environment.
GDAL is the install's real difficulty, not pip
Installation is offered four ways and each has a catch.
Pip requires a reasonably recent pip before the project file will be honoured, with an upgrade command given if yours is older. For a stable release you install an exact version; for the moving branch you install from the git URL. Conda-forge carries the package too. And pipx is offered for people who want the command line tool in an isolated environment without activating a virtual environment, which is the cleanest option if you only intend to launch fine-tuning runs from a terminal.
The stated hard part is GDAL. The project requires it, calls the installation process complex, and recommends a conda environment with GDAL from conda-forge, adding that installing from that channel will probably avoid the problem.
That recommendation is worth taking literally. Geospatial raster stacks depend on GDAL's native libraries for reading and writing, and the version alignment between GDAL, rasterio, fiona and the geospatial packages is exactly the kind of thing that pip cannot resolve and conda can. The dependency list shows the neighbours: rasterio for raster I/O, an array library with a raster extension, geopandas for vector work, and a data augmentation stack.
So the ordering matters: get the geospatial stack resolving first, then install TerraTorch into it.
The container runs as uid 10000 in group 0 with a New York clock
The published container image is worth reading as a configuration document, because three choices in it tell you how the maintainers expect it to be deployed.
First, the base image is the NVIDIA CUDA runtime image on Ubuntu 24.04 with cuDNN. That is a runtime image rather than a development one, so there is no compiler in it, and the file installs build tools and then purges them again at the end of the build. That is a deliberate slimming step rather than an oversight.
Second, the process runs as an unprivileged user with id 10000, and the application root is given to group 0 with group write and execute permission before ownership is changed. That is the group-zero convention used for OpenShift-compatible and Kubernetes deployments with restricted security contexts, where arbitrary high user ids are assigned and group zero is the writable group. If you deploy this on a cluster that enforces a restricted policy, the pattern will save you a fight with the admission controller.
Third, the image hardcodes a timezone by symlinking the New York zone file and installing timezone data non-interactively. That is convenient for the maintainers and mildly surprising for you, since a container that reports New York time while your pipeline assumes UTC will produce offset bugs in logs and in any timestamped output.
The image also ships a command of a bare shell, so it is a development environment rather than an application image, which is what you want for a fine-tuning toolkit.
A dual licence expression and a duplicated author name
A few housekeeping details are visible in the project metadata and worth knowing if you are auditing compliance.
The licence field is an expression rather than a single licence: the project's own code is Apache-2.0 and there is a second MIT-licensed file group, tracked in a file at the repository root that lists which files fall under which terms. The packaging metadata points at both the licence file and that list, so vendoring the library means reading the list rather than assuming the whole tree is Apache.
There is a vulnerability policy document at the root alongside a constraints directory, which together suggest pinned dependency versions for builds and a documented reporting path. Test infrastructure is thorough for a research library: a tox configuration, a dedicated test directory, a separate integration test directory, and three Dockerfiles, one for a normal image, one for builds and one for running tests, plus several shell scripts that drive those images.
In the author list one name appears twice, which is the kind of thing that happens when a contributor list is edited by hand.
The examples directory is the best map of coverage, with thirteen directories covering classification, custom modules, datasets and benchmarks, embeddings, models, multimodal data, multitemporal data, object detection, pixel-wise regression, segmentation, utilities, a local inference server folder, weather and climate, and web applications.
Editorial conclusion
TerraTorch fits a geospatial team that has a pretrained backbone from the Hugging Face collection and needs the boring half of fine-tuning, namely trainers, data modules and configuration, without writing them. Skip it if you want hosted weights or a single supported stack, because the backbone list is deliberately wide and the optional extras are where the version conflicts live. Before you start, install GDAL through conda rather than fighting it, and read the licence of every checkpoint you plan to deploy, because the project deliberately hands you that responsibility.
Frequently asked questions
What is TerraTorch?
A PyTorch domain library built on PyTorch Lightning and TorchGeo for geospatial data, providing a flexible fine-tuning framework for geospatial foundation models, with trainers for segmentation, classification and pixel-wise regression launched from the CLI, configuration files or notebooks.
Does TerraTorch host the models?
No. It provides only the training and inference framework, and it states that verifying whether a model's licence allows your intended use is the sole responsibility of the user, with no legal advice offered by the maintainers.
How do I install TerraTorch?
With pip at an exact version for a stable release or from the git URL for the main branch, from conda-forge, or with pipx for an isolated command line install. GDAL is also required, and conda-forge is the recommended way to get a working geospatial stack.
Which backbones can TerraTorch use?
Prithvi, TerraMind, SatMAE, ScaleMAE and Clay, plus Satlas, DOFA and the SSL4EO models through TorchGeo, and anything available in the timm collection. Decoders come from segmentation-models-pytorch or from MMSegmentation.
What do the TerraTorch optional extras provide?
They gate separate integrations: local inference serving, parameter-efficient fine-tuning, weights and biases logging, MMSegmentation, Surya, Tortilla dataset files, visualization, a second benchmark version, weather foundation models on Python 3.11+, and an object-detection integration.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/torchgeo-terratorch)