# Nevergrad: gradient-free optimization for Python functions you cannot differentiate

> Nevergrad is a Python 3.8+ library that searches an input space with derivative-free optimizers. It fits black-box tuning and simulation loops; it does not replace gradient-based training, and its documentation is still a work in progress.

**facebookresearch/nevergrad** — A Python toolbox for performing gradient-free optimization

- Repository: https://github.com/facebookresearch/nevergrad
- Website: https://facebookresearch.github.io/nevergrad/
- Stars: 4,211 · Forks: 371
- Language: Python
- License: MIT
- Published: 2026-09-23 · Updated: 2026-09-23 · Language: en
- Canonical page: https://hysenlabs.com/projects/facebookresearch-nevergrad

## The problem Nevergrad targets: objectives with no usable gradient

Some objective functions cannot be differentiated. A simulator, a rendering pipeline, a hardware sweep or a training script that returns one scalar has no analytic gradient, and finite differences cost one extra evaluation per dimension. Nevergrad exists for that case: it is described in the README as a gradient-free optimization platform, and it is a Python 3.8+ library. The audience is engineers and researchers who can write a function from parameters to a float and who control the input space, not the internals of the function. The README's own example is a quadratic, but the more representative one is a fake_training function that takes a learning rate, a batch size and an architecture name and returns a loss. That shape, mixed parameter types with no derivative, is what the library is built around. The Facebook users group linked from the README is the stated place for user questions.

## How parametrization, optimizers and budget fit together

Nevergrad separates the search space from the search algorithm. The search space is a parametrization object. The README shows ng.p.Log for a log-distributed scalar between 0.001 and 1.0, ng.p.Scalar(lower=1, upper=12).set_integer_casting() for an integer, and ng.p.Choice(["conv", "fc"]) for a categorical value. When the function takes several positional or keyword arguments, ng.p.Instrumentation wraps them into one parametrization. The optimizer is constructed with that parametrization and a budget, the number of evaluations it is allowed. NGOpt is the default recommender shown in both README examples. Calling optimizer.minimize(fn) returns a recommendation object that exposes .value for a plain array parametrization and .kwargs for an Instrumentation parametrization. The repository also lists examples such as examples/disc.py and examples/advbinmatrix.py, which suggests discrete and adversarial binary matrix cases are covered beyond the README snippets. The optimizer families named in the project's own search vocabulary include oneplusone, differential evolution and CMA, so the choice of algorithm is a real decision rather than an implementation detail.

## Installing Nevergrad and running a first optimization

The README gives a single install command and points to the Getting started section of the documentation for more options, including Windows installation. The README does not list a conda package, so pip is the documented path.

```bash
pip install nevergrad
```

After install, the smallest useful program defines a function, builds an optimizer with a budget, and calls minimize. The README's first example uses a two-dimensional parametrization, which means the function receives an array.

```python
import nevergrad as ng

def square(x):
    return sum((x - .5)**2)

optimizer = ng.optimizers.NGOpt(parametrization=2, budget=100)
recommendation = optimizer.minimize(square)
print(recommendation.value)
```

The printed value is an array near 0.5 in both coordinates, matching the README output [0.49971112 0.5002944]. For a function with named arguments, use ng.p.Instrumentation and read recommendation.kwargs instead of recommendation.value.

```python
import nevergrad as ng

def fake_training(learning_rate: float, batch_size: int, architecture: str) -> float:
    return (learning_rate - 0.2)**2 + (batch_size - 4)**2 + (0 if architecture == "conv" else 10)

parametrization = ng.p.Instrumentation(
    learning_rate=ng.p.Log(lower=0.001, upper=1.0),
    batch_size=ng.p.Scalar(lower=1, upper=12).set_integer_casting(),
    architecture=ng.p.Choice(["conv", "fc"])
)
optimizer = ng.optimizers.NGOpt(parametrization=parametrization, budget=100)
recommendation = optimizer.minimize(fake_training)
print(recommendation.kwargs)
```

The README reports {'learning_rate': 0.1998, 'batch_size': 4, 'architecture': 'conv'} for this run. Note that the budget is 100 evaluations, so a real objective that takes minutes per call needs a plan for that cost before you start.

## Where Nevergrad is the wrong tool

If your function is differentiable and you have the gradient, Nevergrad is the wrong choice. Gradient descent and its variants use derivative information that a derivative-free method throws away, and the README does not claim otherwise: it presents the library as gradient-free by design. The same applies when the objective is cheap to evaluate and the dimension is low, where a grid search is easier to reason about and easier to explain to a reviewer. Budget is the other hard boundary. optimizer.minimize runs the number of evaluations you declared, and each evaluation is a call to your function. A budget of 100 on a training script that takes ten minutes per run is a seventeen-hour experiment before any overhead. The README does not document checkpointing, resumption or a rollback mechanism, so an interrupted long run is a real operational risk. Finally, the README states that the documentation is still a work in progress and invites issues and pull requests to make it clearer. That is an honest signal that some behaviour will need to be read from the source or the examples directory rather than the docs.

## Nevergrad versus Optuna and the tuning-library comparison

The most common comparison for this library is with Optuna, and the difference is in the starting point. Optuna is built around hyperparameter search for machine learning frameworks, with study objects, trial pruning and a storage backend for distributed runs. Nevergrad starts lower: you define the parametrization and the optimizer directly, and the library treats the objective as an arbitrary black box, not necessarily a model training loop. That makes Nevergrad a better fit when the thing being optimized is a physics simulation, a compiler flag set or a control problem, and a worse fit when you want pruning of bad trials inside a training loop. The repository shows the boundary is not absolute: examples/raytune_nevergrad_biobjective.py integrates Nevergrad with Ray Tune, and examples/overfit.py and examples/check_metamodel.py show it being used on learning problems. The practical rule is that if your objective is a model training run with epochs, a tuning library gives you more scaffolding; if it is any function returning a float, Nevergrad's parametrization model is the more direct abstraction.

## Maintenance, licence and the cost of upgrading

The repository is not archived, and its last push was on 2026-07-24. Releases are tagged in the 1.0.x line, with 1.0.12 published on 2025-04-23 and the 1.0.10 notes titled Preparing Retrofitting and Github Actions, which indicates packaging and CI work rather than a change in the optimizer interface. The installable version comes from nevergrad/__init__.py, which setup.py reads with a regex, so a version bump is a source edit rather than a tag-only operation. Nevergrad is released under the MIT license, per the README and the LICENSE file, which permits commercial use and modification; the README also links Facebook's Terms of Use and Privacy Policy, and if those matter to your organisation, read them rather than treating the MIT label as the whole story. This is a description of the licence text, not legal advice. On upgrade cost, the README shows a stable surface (ng.optimizers.NGOpt, parametrization objects, minimize, .value and .kwargs), but the CHANGELOG.md at the repository root is where a breaking change would be recorded, and the README does not promise API stability across 1.0.x releases.

## Conclusion

Adopt Nevergrad when the objective is a black box you can call but cannot differentiate, and when the search space mixes continuous, integer and categorical variables. Skip it if you already have usable gradients, since first-order methods are the wrong comparison class. Before committing, run your own objective through NGOpt at a budget you can afford and check whether the recommended value is stable across seeds; the README does not document rollback or reproducibility guarantees.

## FAQ

### What are the downsides of using gradient descent?

Gradient descent needs a usable derivative, and many objectives do not have one: a simulator or a hardware sweep returns a scalar with no analytic gradient, and finite differences cost one extra evaluation per dimension. Nevergrad is described in the README as a gradient-free optimization platform for exactly that situation.

### Is gradient descent still used?

The README does not discuss gradient-based training at all, so it says nothing about whether gradient descent is still used. What it does state is that Nevergrad is gradient-free by design, which makes it a complement for objectives where gradients are unavailable rather than a replacement for first-order methods.

### What is Nesterov Accelerated gradient (NAG)?

NAG is not mentioned in the README, the repository layout or the release notes, so this material cannot answer the question. Nevergrad's documented surface is the parametrization classes and optimizers such as NGOpt, not first-order momentum methods.

### What are the alternatives to Nevergrad?

The README does not name alternatives. The repository does show one integration point: examples/raytune_nevergrad_biobjective.py combines Nevergrad with Ray Tune, whose tuning layer provides study and trial scaffolding that Nevergrad itself does not.

## Sources

- [facebookresearch/nevergrad on GitHub](https://github.com/facebookresearch/nevergrad)
- [License: MIT](https://github.com/facebookresearch/nevergrad/blob/main/LICENSE)
- [Project website](https://facebookresearch.github.io/nevergrad/)
- [README](https://github.com/facebookresearch/nevergrad/blob/main/README.md)
- [Releases](https://github.com/facebookresearch/nevergrad/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/facebookresearch-nevergrad
