SMT (Surrogate Modeling Toolbox): a Python library built around derivatives
SMT: The Surrogate Modeling Toolbox
At a glance
- What is it?
- SMT is a BSD-licensed Python package of surrogate models, sampling methods and benchmark functions, with derivatives as its organising idea. It installs with pip and requires Python 3.10 or newer.
- Who is it for?
- Adopt SMT if your surrogate work needs gradients, mixed or hierarchical input variables, or multi-fidelity data, and you are comfortable in Python. Skip it if you want a drop-in AutoML regressor or a GUI, or if you cannot build C++ extensions on your target machine.
- Can I use it commercially?
- Yes. BSD-3-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 13 days ago.
- What is it written in?
- Mainly Jupyter Notebook, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What SMT is for, and who ends up using it
SMT targets a specific workflow: you have an expensive simulation, you have already run it a few hundred times, and you want a cheap approximation you can call inside an optimiser. The README describes the package as a collection of surrogate modeling methods, sampling techniques, and benchmarking functions, and states that it differs from existing surrogate modeling libraries because of its emphasis on derivatives. That emphasis is the whole pitch. Training derivatives used for gradient-enhanced modeling, prediction derivatives, and derivatives with respect to the training data are all first-class concerns in the API rather than an afterthought bolted onto a prediction method.
The audience follows from that. SMT is aimed at research and engineering users: the classifiers in pyproject.toml list Intended Audience :: Science/Research and Intended Audience :: Developers, and the project is maintained by Rémi Lafage at ONERA. The associated publications are aerospace and multidisciplinary-design papers on gradient-enhanced kriging, airfoil shape design and sparse Gaussian processes. If you are fitting a surrogate to a CFD campaign and you want the adjoint derivatives your solver already produces to improve the fit, this is the library's home ground. If you just want to regress a CSV of tabular data, the derivative machinery is weight you will not use.
The mechanism: a model registry, a sampling layer and compiled kernels
The repository layout tells you most of the architecture. Everything lives under smt/, with the surrogate models in smt/surrogate_models, and the top level of the repository also carries doc/, tutorial/ and environment.yml. The package exposes a family of model classes rather than a single estimator, so the user picks a model by importing it. The search phrases people use around this project, Smt krg, Smt kpls and Smt surrogate_models, map onto that structure: KRG is the kriging model, KPLS is the kriging by partial-least-squares reduction that the README names as one of the models not available elsewhere, and surrogate_models is the module that holds them.
Three of those models are not pure Python. setup.py builds Cython and C++ extensions for smt.surrogate_models.rbfclib, smt.surrogate_models.idwclib and smt.surrogate_models.rmtsclib, compiling sources from smt/src/rbf, smt/src/idw and smt/src/rmts. The rmts extension alone pulls in utils.cpp, rmts.cpp, rmtb.cpp and rmtc.cpp. The build passes -std=c++11 on non-Windows platforms and needs numpy headers at compile time. This is why the build-system table in pyproject.toml requires numpy and Cython in addition to setuptools: a wheel has to compile before it can be installed. The payoff is that the radial basis function, inverse distance weighting and RMT implementations run as native code.
Sampling and benchmarking sit alongside the models. The README lists sampling techniques and benchmarking functions as part of the package, and the topics tag sampling and predictive-modeling. In practice you generate a design of experiments, evaluate your expensive function on it, fit a model, then query the model and its derivatives. Optional extras change what is available: the numba extra pulls in numba, gpx pulls in egobox, and rbfgen pulls in torch.
Installing SMT and fitting a first kriging model
The README gives two installation commands. The first installs the latest release from PyPI, which is what most users want, and the second installs from the current master branch when you need something that has not been released yet. Note the Python floor: pyproject.toml sets requires-python to >=3.10, so an older interpreter will refuse the install.
pip install smtBecause the package ships Cython extensions, this is not a pure-Python download. If pip cannot find a matching wheel for your platform and Python version, it falls back to building from source, which means you need a working C++ toolchain and numpy headers. On a machine where the build fails, the error appears during installation rather than at import time.
The README points at two places for usage: the tutorial notebooks under tutorial/ and the smt/examples folder. It does not print a minimal code sample in the README itself, so the reliable route to a first fit is to open one of those notebooks and run it against your own data. The pattern you will find there is the same across models: build the training arrays, instantiate the model class, call fit, then call predict. The documentation site at smt.readthedocs.io is where the API reference and the generated plots live, and the README notes that the docs are built with custom tooling that embeds automatically tested code, so the snippets there are executed as part of the build rather than pasted by hand.
If you want the development version instead of the release, the README gives this command:
pip install git+https://github.com/SMTOrg/smt.git@masterThat route clones master and builds it locally, so it inherits the same compiler requirement. For contributors, requirements.txt lists the full working set: Cython, packaging, numpy, scipy, scikit-learn, pydoe between 1.0.0 and 2.0, numba, matplotlib, pytest, pytest-xdist, pytest-cov, ruff, setuptools-scm, jenn between 2.0.0 and 3.0, egobox between 0.36 and 1.0, and torch 2.4.1 or newer. The pytest-xdist entry is there so tests can run in parallel with pytest -n <num_workers>.
Where SMT is the wrong tool
The compiled extensions are the sharpest limitation. A pure-Python surrogate library installs anywhere; SMT needs a C++ compiler when no wheel is available, and setup.py explicitly branches on the platform, adding -std=c++11 only when sys.platform does not start with win. That branch is a hint about how much platform-specific care the build needs. If your deployment target is a locked-down environment with no compiler and no prebuilt wheel, SMT is a poor fit regardless of how good the models are.
The second limitation is scope. SMT is a toolbox of model classes, not a pipeline. There is no mention in the README of automatic hyperparameter search across model families, no preprocessing chain, no persistence format, and no serving layer. You assemble those yourself. If your team's need is a single estimator that ingests a dataframe and reports cross-validated accuracy, a general-purpose regression library covers that with less surface area. SMT assumes you already know which surrogate family suits your problem and that you care about gradients.
The third is documentation shape. The README is short by design and defers to the notebooks and the ReadTheDocs site. The README does not document rollback between versions, and it does not give a compatibility matrix for the optional extras. The extras are pinned loosely (egobox ~=0.36, numba ~=0.64, torch>=2.4.1), so a fresh environment can resolve to versions that differ from what a tutorial notebook was written against. Treat the notebook as an example of the API shape, not as a frozen environment.
How SMT differs from scikit-learn's Gaussian process regressor
The obvious comparison is scikit-learn, which SMT depends on. sklearn.gaussian_process.GaussianProcessRegressor is a solid single-model implementation with a clean estimator interface, and if your inputs are continuous, your sample count is modest, and you do not have derivatives, it will do the job. The difference in approach is what each library treats as the interesting part of the problem.
Scikit-learn treats the surrogate as one estimator among many, so it optimises for a uniform fit/predict interface and for composability with pipelines and model selection utilities. SMT treats the surrogate as an engineering artefact with structure: gradients in, gradients out, mixed and hierarchical variables, multi-fidelity data, and model families that scikit-learn does not offer. The README names two models it says are not available elsewhere, kriging by partial-least-squares reduction and energy-minimizing spline interpolation, and the topics list includes mixture-of-experts and multi-fidelity. Scikit-learn has none of those as built-in regressors.
The practical consequence: with scikit-learn you get less code to write and a broader ecosystem; with SMT you get gradient-enhanced fitting and model families tailored to design optimisation, at the cost of a compiled dependency and a smaller API surface. The two are not exclusive. SMT lists scikit-learn as a runtime dependency, so installing SMT does not remove it.
Release cadence, licence and the cost of upgrading
SMT is released often. The recent tags are 2.15.0 on 2026-09-08, 2.14.1 on 2026-06-22 and 2.14.0 on 2026-05-11, and the last push to master was on 2026-09-08. The repository is not archived. That cadence is a maintenance signal in both directions: fixes arrive quickly, and so do version bumps you may need to track. Version numbers are derived from git tags through setuptools-scm, with version_scheme set to guess-next-dev and local_scheme set to node-and-date, so an install from a non-tagged commit carries a local version suffix rather than a clean release number. If your build system pins exact versions, that suffix can surprise you.
Upgrade cost concentrates in the optional extras and the compiled extensions. The extras are version-ranged rather than pinned to exact releases, so resolving them fresh can move numba, egobox or torch without you changing anything. The extensions are rebuilt on every source install, so a compiler that worked for one release is a prerequisite for the next. There is no migration guide in the README; the release notes are the place to look before bumping.
The licence is BSD-3-Clause, declared both in pyproject.toml and in the LICENSE.txt file at the repository root, and the README says the package is distributed under the New BSD license. For most users that means permissive reuse with attribution and no copyleft obligation on your own code. It does not mean the dependencies are equally permissive; scipy, scikit-learn, torch and egobox carry their own terms, and if you redistribute a bundled environment you need to check each one. That is a licensing question for your own counsel, not something the SMT repository answers.
Editorial conclusion
Adopt SMT if your surrogate work needs gradients, mixed or hierarchical input variables, or multi-fidelity data, and you are comfortable in Python. Skip it if you want a drop-in AutoML regressor or a GUI, or if you cannot build C++ extensions on your target machine. Before committing, install it in the Python version you actually deploy on, confirm the compiled rbf, idw and rmts extensions import, and check that the model class you plan to use appears in the stable documentation rather than only in the tutorial notebooks.
Frequently asked questions
What is surrogate modeling, and what does SMT have to do with it?
Surrogate modeling fits a cheap approximation to an expensive function so you can evaluate it many times, for example inside an optimiser. SMT is a Python package that collects surrogate modeling methods, sampling techniques and benchmarking functions for exactly that job.
How do I install SMT with pip?
The README gives pip install smt for the latest release, or pip install git+https://github.com/SMTOrg/smt.git@master for the current master branch. The package requires Python 3.10 or newer and builds Cython and C++ extensions, so a source install needs a working compiler.
Which model classes does SMT provide under smt.surrogate_models?
The README does not enumerate them, but it names kriging by partial-least-squares reduction (KPLS) and energy-minimizing spline interpolation as models not available elsewhere, and the associated publications cover gradient-enhanced kriging, mixture of experts and sparse co-kriging. The full list is in the API documentation at smt.readthedocs.io.
Does SMT work with derivatives?
Yes, and that is its stated differentiator. The README says SMT emphasises training derivatives used for gradient-enhanced modeling, prediction derivatives, and derivatives with respect to the training data.
What Python version and dependencies does SMT need?
pyproject.toml sets requires-python to >=3.10 and lists packaging, scikit-learn, pydoe between 1.0.0 and 2.0, scipy and jenn between 2.0.0 and 3.0 as runtime dependencies. Optional extras add numba, egobox and torch.
What licence is SMT released under?
BSD-3-Clause, declared in pyproject.toml and in LICENSE.txt at the repository root, and described in the README as the New BSD license. Your dependencies carry their own licences, which you should check separately.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/smtorg-smt)