BayesianOptimization: a small pure Python library for functions that are too expensive to sample
A Python implementation of global optimization with gaussian processes.
At a glance
- What is it?
- Eight thousand stars for what is, in the end, one class and one method: fit a Gaussian process to what you have measured so far, pick the next point with an acquisition function, and repeat.
- Who is it for?
- BayesianOptimization earns its place when each evaluation of your function costs real money, in compute, lab time, or human patience, and the parameter space is small enough that a grid search would waste evaluations. The API surface is deliberately tiny: construct with a function and pbounds, call maximize, read optimizer.max or optimizer.res.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 46 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 21, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Installing the package from PyPI or conda-forge
The README offers two routes and nothing else, which for a library of this size is a good sign.
$ pip install bayesian-optimizationThe conda route goes through conda-forge.
$ conda install -c conda-forge bayesian-optimizationWhat that install actually pulls in is worth a look, because the README does not tell you and the packaging metadata is more informative than most projects of this age. The pyproject.toml declares requires-python >=3.9, classifiers for Python 3.9 through 3.14 plus a free-threading classifier, and a dependency list of colorama, packaging, numpy, scipy and scikit-learn. It is MIT licensed, authored by Fernando Nogueira, and the version in the file is 3.3.0.
The dependency specifiers are conditional on the Python version, which tells you the project is being kept current rather than abandoned. numpy is pinned at >=1.25 below Python 3.13, >=2.1.3 on 3.13, and >=2.3.0 on 3.14. scipy moves from >=1.0.0 to >=1.14.1 to >=1.17.0 across the same ranges, and scikit-learn from >=1.0.0 to >=1.8.0 on 3.14. Matching the new NumPy major versions per interpreter is exactly the maintenance work that separates a maintained numerical package from one that quietly breaks on a new Python, and the fact that it is written down in the file means you can see it without reading the issue tracker.
The repository also carries a ruff.toml and a .pre-commit-config.yaml, so style is enforced by tooling rather than by review comments.
The whole tutorial is one class and one method
The basic tour is unusually short, which is the point. First you define the function to optimize. The README includes a disclaimer that in a real scenario you would not know exactly how the output depends on its parameters, and that you do not need to: all the package requires is a function that takes a known set of parameters and returns a real number.
def black_box_function(x, y):
"""Function with unknown internals we wish to maximize.
This is just serving as an example, for all intents and
purposes think of the internals of this function, i.e.: the process
which generates its output values, as unknown.
"""
return -x ** 2 - (y - 1) ** 2 + 1Then you construct the optimizer with that function and a dictionary of bounds, one min and max pair per parameter. The README is clear that bounds are not optional, because this is a constrained optimization technique and it cannot work without them.
from bayes_opt import BayesianOptimization
# Bounded region of parameter space
pbounds = {'x': (2, 4), 'y': (-3, 3)}
optimizer = BayesianOptimization(
f=black_box_function,
pbounds=pbounds,
random_state=1,
)Then you call maximize, which the README describes as doing exactly what you think it does. Two parameters matter more than the rest: n_iter, how many steps of Bayesian optimization to perform, and init_points, how many steps of random exploration to perform before the model takes over, offered because random exploration can help by diversifying the explored space.
optimizer.maximize(
init_points=2,
n_iter=3,
)Reading the results through max and res
Results come out through two properties. `optimizer.max` gives the single best combination of parameters and target value found, and `optimizer.res` gives the full list of every parameter probed with its target value.
print(optimizer.max)
>>> {'target': -4.441293113411222, 'params': {'y': -0.005822117636089974, 'x': 2.104665051994087}}Iterating over `res` is how you see the whole trajectory, which matters more for this technique than for a normal optimizer because the path is the information.
for i, res in enumerate(optimizer.res):
print("Iteration {}: \n\t{}".format(i, res))The five-row table the README prints for that run is instructive. The first two rows are the random initialization points and include one at the exact corner of the box, x=2.0 and y=-1.186, which is a reminder that random sampling does hit bounds. Then the acquisition function takes over and the target climbs from -19.0 to -16.3 to -4.441, landing near the true maximum of the example function at x=2, y=1.
That progression is the algorithm doing what the README describes it as doing. At each step a Gaussian process is fitted to the known samples, the posterior distribution is combined with an exploration strategy such as UCB, Upper Confidence Bound, or EI, Expected Improvement, and the result picks the next point to explore. The README is careful to name the proxy problem underneath: finding the maximum of the acquisition function is itself hard, but it is cheaper in computational terms, so common tools can be used for it.
The examples directory shows the real feature set
The README sends you to the hosted documentation for anything beyond the tour, so the repository's examples directory is the better map of what the library actually does. It contains notebooks and scripts covering acquisition functions, an advanced tour, async optimization in both `async_optimization.py` and `async_optimization_dummies.py`, constraints, domain reduction, duplicate points, exploitation versus exploration, parameter types, visualization, and two hyperparameter-oriented examples: `sklearn_example.py` and `typed_hyperparameter_tuning.py`.
That list tells you a lot about the intended use. Constraints means the package supports telling the optimizer about invalid regions rather than only about bounds. Domain reduction means you can restrict the search to a promising subspace as earlier iterations narrow things down. Parameter types covers categorical and other non-continuous dimensions, which the release history confirms was a source of real bugs: v3.2.1 included a fix for `.max()` for categorical, and v3.3.0 fixed a pbounds type mismatch.
Async optimization matters more than it sounds. A Bayesian loop is inherently sequential, because each point depends on everything measured before it, so making it asynchronous means running candidate evaluations concurrently and folding results in as they land. The two async examples, one real and one with dummies, suggest the project wants that to be usable rather than aspirational.
The tree is small and disciplined: `bayes_opt/` for the library, `tests/`, `examples/`, `docsrc/` for the Sphinx sources behind the hosted documentation, `scripts/`, and the two config files. There is no vendored dependency directory and no compiled extension, which is what a pure Python implementation should look like.
Where this loses to the alternatives
The honest comparison is with three other things, and this library only wins one of them.
Against a grid search, Bayesian optimization wins when each evaluation is expensive. If you can evaluate the function a thousand times in a minute, exhaustive search is better in expectation, simpler to reason about, and easier to parallelize trivially. The README says this itself in its own terms: the method is most adequate for situations where sampling the function to be optimized is a very expensive endeavor.
Against Optuna, which appears in the related search terms for this project, the trade is different. Optuna defines a search space declaratively and prunes trials that are clearly bad, which scales better when you have thousands of trials and a real training loop to kill. This package asks for a Python callable and optimizes whatever you return, which is a smaller surface and easier to debug but gives up pruning. `typed_hyperparameter_tuning.py` exists to bridge that gap, so the intended workflow is to wrap a scikit-learn estimator rather than to hand the library a training loop.
Against scikit-learn's own Bayesian optimization, which the dependency list implies is available in the same environment, the case for a standalone package is control over the acquisition function and the loop. The examples directory shows acquisition functions as a first-class topic, which suggests that if you want to substitute your own criterion, this is the library that makes it easy.
The other limitation is dimensional. Gaussian process regression scales badly in the number of parameters, so this is a tool for tuning a handful of knobs, not thirty.
Maintenance signals and what the README leaves out
The repository is not archived, has 8,712 stars and 1,604 forks, and only 7 open issues, which for a project this widely used is a strikingly small queue. The last push was 2026-08-21 and the README states plainly that the project is under active development and invites issues.
The recent releases are small and specific, which fits a mature numerical library. v3.3.0 on 2026-05-30 fixed a pbounds type mismatch and refactored the state handling by extracting `_state_to_dict` and `_load_state_dict` methods, which is groundwork for saving and restoring an optimizer. v3.2.2 on 2026-05-11 removed stale ndarray sorting warnings, and v3.2.1 on 2026-03-16 fixed type hints and the categorical `.max()` case. Dependency bumps are handled by a maintainer in separate commits, which is a sign of a release process that exists.
Those state methods are worth noticing for a practical reason. An optimization run can be long and expensive, and being able to serialize and reload an optimizer is what lets you checkpoint a search across process restarts instead of losing the evaluations already paid for.
What the README does not cover is everything about the hosted documentation: the API reference, the acquisition function catalogue, constraint syntax, and the async API surface. The badge for the docs is marked stable and hosted at bayesian-optimization.github.io, so that documentation exists and is versioned separately from the README. Read the constraints and domain reduction notebooks in the examples directory before you use the library on something expensive; they are where the real behaviour lives.
Editorial conclusion
BayesianOptimization earns its place when each evaluation of your function costs real money, in compute, lab time, or human patience, and the parameter space is small enough that a grid search would waste evaluations. The API surface is deliberately tiny: construct with a function and pbounds, call maximize, read optimizer.max or optimizer.res. What it is not is a replacement for grid search on a cheap function, where exhaustive evaluation wins anyway, nor a general hyperparameter tuner, though examples/typed_hyperparameter_tuning.py shows the intended shape of that work. The repository has 8,712 stars, 1,604 forks and 7 open issues, with the last push on 2026-08-21 and v3.3.0 published 2026-05-30, so the dependency pinning in pyproject.toml against numpy, scipy and scikit-learn is being maintained rather than left to rot. Start with the basic tour, then read the constraints and domain reduction notebooks before you trust it on a real budget.
Frequently asked questions
Can you explain Bayesian optimization in a simple way?
It fits a Gaussian process to the points you have already evaluated, which gives you a posterior belief about the function everywhere, then picks the next point using an exploration strategy such as UCB or Expected Improvement. Each iteration makes the model more certain about which regions of parameter space are worth visiting, and the process is designed to minimise the number of steps needed to find a near-optimal combination.
Can you explain Bayesian inference in a simple way?
Bayesian inference updates a belief about unknowns as evidence arrives, combining what you assumed beforehand with what you have just observed. In this package that machinery is the Gaussian process: it constructs a posterior distribution over functions that best describes the function you want to optimize, and the posterior improves as the number of observations grows.
Is Bayesian optimization machine learning or AI?
It is better described as a global optimization technique that borrows from machine learning. The README calls it a constrained global optimization package built upon Bayesian inference and Gaussian processes, and it optimizes your function rather than predicting labels, although the Gaussian process surrogate it fits is itself a machine learning model.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/bayesian-optimization-bayesianoptimization)