Library / SDK
google/uncertainty-baselines avatar
google/uncertainty-baselines

google/uncertainty-baselines: A Forkable Template for Uncertainty Estimation in Deep Learning

High-quality implementations of standard and SOTA methods on a variety of tasks.

1,596 stars225 forksPythonApache-2.0

At a glance

What is it?
Google's uncertainty-baselines repository collects reference implementations of standard and state-of-the-art uncertainty methods on standard datasets, designed to be copied rather than imported. It is a researcher's starting point, not a production library.
Who is it for?
Adopt it if you are a researcher who needs a working reference for a method such as BatchEnsemble or a deterministic Wide ResNet, and you are willing to fork a baseline script and accept that the install pulls tfds-nightly pinned to 4.4.0.dev202111160106. Do not adopt it if you need a stable, versioned dependency for a production training pipeline, because the README states there is not yet a stable version and all APIs are subject to change.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 2 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The comparison problem uncertainty-baselines was built to fix

The README opens with a complaint that will be familiar to anyone who has tried to reproduce an uncertainty paper. Implementations are scattered across GitHub, most of them are one-off experiments tied to a single paper, and many papers ship no code at all. Even when code exists, architectures, hyperparameters and data preprocessing differ slightly from project to project, so a reported gap between two methods may be an artifact of the setup rather than a property of the method.

The repository's stated goal is to be a template for researchers to build on, in three ways: high-quality implementations of standard and state-of-the-art methods on standard tasks, minimal dependencies on other files in the codebase so a baseline can be forked without dragging along the rest of the project, and prescribed best practices for benchmarking. That third point is the interesting one. The first two are engineering hygiene; the third is an attempt to standardize how uncertainty and robustness results are reported.

The intended user is a researcher who wants to prototype a new idea and needs something credible to compare against without rebuilding a Wide ResNet training loop from scratch. It is explicitly not aimed at application developers who want to add calibrated predictions to a product.

How the baselines, datasets and models modules fit together

The repository layout separates three concerns. The baselines/ directory holds training scripts organized by dataset, so baselines/cifar/deterministic.py is a Wide ResNet 28-10 reaching 96.0% test accuracy on CIFAR-10, and baselines/cifar/batchensemble.py is a separate script for the BatchEnsemble method. Each script is meant to stand alone.

The uncertainty_baselines package holds the shared pieces. ub.datasets follows the TensorFlow Datasets API and adds default preprocessing, so a dataset builder is constructed with a split and an optional validation_percent, then loaded with a batch size. ub.models follows the tf.keras.Model API, so a model is constructed by calling a function such as ub.models.wide_resnet with explicit depth, width_multiplier, num_classes and l2 arguments.

Data flow is conventional: a dataset builder yields batches, a model consumes them, and metrics are computed over the test set. The README defines the reported metrics, including number of parameters, test accuracy, and test calibration error as expected calibration error over the test set. Results are reported to roughly three significant digits and averaged over 10 runs, which is the part that makes cross-paper comparison meaningful.

The README also notes that for Jax and PyTorch you can convert batches with tfds.as_numpy, but points out that this calls tensor.numpy() and invokes an unnecessary copy compared to tensor._numpy(). The suggested alternative iterates the dataset directly and maps jax.tree.map over the batch. That is a small detail, and a good signal that the authors care about the training loop rather than just the model definitions.

Installing uncertainty_baselines and running a first baseline

There is no PyPI release. The README gives a single install command that pulls the development version straight from GitHub, and it warns that installing the package does not install any backend. For TensorFlow you also need TensorFlow, TensorFlow Addons and TensorBoard, with the nightly variants listed as alternatives. The setup.py extras are the place to look for the optional dependency groups.

bash
pip install "git+https://github.com/google/uncertainty-baselines.git#egg=uncertainty_baselines"

After installing, the README's usage example for the CIFAR-10 BatchEnsemble baseline assumes a cloud TPU and a few environment variables pointing at a GCS bucket. The script takes the TPU name, a data directory and an output directory as flags.

bash
export BUCKET=gs://bucket-name
export TPU_NAME=ub-cifar-batchensemble
export DATA_DIR=$BUCKET/tensorflow_datasets
export OUTPUT_DIR=$BUCKET/model

python baselines/cifar/batchensemble.py \
    --tpu=$TPU_NAME \
    --data_dir=$DATA_DIR \
    --output_dir=$OUTPUT_DIR

The README states that the TPU accelerator type must align with the baseline's number of cores, and that BatchEnsemble defaults to num_cores=8, so the TPU must be set up with accelerator_type=v3-8. If you do not have a TPU, the documented fallback is to change the flags and run on GPUs, which the README shows as a separate invocation with --use_gpu=True and --num_cores=8, using local paths for the data and output directories. The README's own caveat is blunt: results may be similar, but ultimately all bets are off, and changing the number of cores matters a lot because total batch size during each training step is often determined by num_cores.

To use the dataset layer on its own, the README shows constructing a builder and loading batches. In an ipython or Colab notebook it notes that you may need to activate eager execution first.

python
import uncertainty_baselines as ub

dataset_builder = ub.datasets.Cifar10Dataset(split='train',
                                             validation_percent=0.1)
train_dataset = dataset_builder.load(batch_size=FLAGS.batch_size)
for batch in train_dataset:
  # Apply code over batches of the data.

Models are constructed the same way, by calling a factory function with explicit shape and size arguments rather than by instantiating a class.

The dependency pin is the sharpest edge in setup.py

The install_requires list in setup.py contains one entry that deserves attention before you build anything on top of this repository: tfds-nightly pinned to exactly 4.4.0.dev202111160106. That is a nightly build of TensorFlow Datasets locked to a single dated version, and it sits alongside tensorflow_probability, which the comment explains is required because robustness_metrics does not do lazy loading. Another dependency, robustness_metrics, is installed directly from a git URL rather than from a package index.

Pinning a nightly is a deliberate trade. It guarantees that the dataset preprocessing in ub.datasets matches what the baselines were validated against, which is exactly the reproducibility property the project is selling. It also means the package cannot be installed alongside a different tfds version without conflict, and that the pinned nightly is the kind of artifact that can disappear from an index. Anyone integrating this into a larger environment should expect to resolve that conflict by hand.

The commit history is current: the last push to the default branch was on 2026-09-10. The repository is not archived. There are no retrieved releases, which is consistent with the README's statement that there is not yet a stable version nor an official release of the library.

Where uncertainty-baselines is the wrong tool

The README is unusually direct about the project's limits, and the most important one is the API stability statement: there is not yet a stable version, and all APIs are subject to change. If you are building a service that depends on ub.models or ub.datasets, an upstream commit can change a function signature and break you. Forking is the intended mitigation, and the README's second goal, minimal dependencies between baselines, exists precisely to make forking cheap.

The second limitation is hardware. The README says you often need TPUs to reproduce baselines. Colab offers free TPUs, which the README calls the most convenient and budget-friendly option, but it also warns against relying on Colab long-term because TPU access is not guaranteed and Colab does not manage multiple long experiments well. Google Cloud is described as the most flexible option, which is another way of saying it costs money and setup time. Running on GPUs is possible by changing flags, but the README explicitly declines to promise that the numbers will match.

The third limitation is scope. This is a benchmarking template for standard tasks. It does not ship a serving path, a monitoring story or a calibration layer you can drop into an existing inference stack. If what you actually need is calibrated predictions in production, you are looking at the wrong repository.

How it differs from a general-purpose uncertainty library

The closest comparison in the same space is TensorFlow Probability, which is a dependency here rather than a competitor. The difference in approach is structural. TensorFlow Probability provides layers, distributions and inference machinery that you compose into your own model. uncertainty-baselines provides complete training scripts for named methods on named datasets, with the hyperparameters already chosen and the metrics already defined.

That means the two solve different halves of the problem. If you want to build a variational layer into a model you already have, TensorFlow Probability is the lower-level tool and uncertainty-baselines will not help you. If you want to know what test calibration error a BatchEnsemble reaches on CIFAR-10 under a specific setup, and you want to modify that setup, the baseline script is the artifact you want, and TensorFlow Probability alone gives you no reference point.

The trade-off is flexibility. A baseline script is opinionated about architecture, batch size and core count, and the README's warning that changing num_cores changes the effective batch size shows how tightly those choices are coupled. You inherit the opinions along with the reproducibility.

Licence and the cost of staying current

The repository is Apache-2.0, and setup.py declares the same licence. Apache-2.0 permits commercial use and modification and includes a patent grant, which is more permissive than a copyleft licence would be. One practical consequence of forking under this licence is that you must retain the licence and attribution notices in the copies you distribute. This is a description of the licence text, not legal advice; check the LICENSE file in the repository and your own counsel for anything that matters.

The upgrade cost is dominated by the pinned nightly. Because tfds-nightly is fixed at 4.4.0.dev202111160106, moving to a newer TensorFlow Datasets means either accepting that the dataset layer may behave differently or updating the pin yourself and re-validating the baselines you depend on. Since the project has no releases, there is no version number to pin against on your side; you either track the main branch or fork at a commit. Forking at a commit is the cheaper option for a research project that needs to stay reproducible, and the README's forkability goal is written with that in mind.

Editorial conclusion

Adopt it if you are a researcher who needs a working reference for a method such as BatchEnsemble or a deterministic Wide ResNet, and you are willing to fork a baseline script and accept that the install pulls tfds-nightly pinned to 4.4.0.dev202111160106. Do not adopt it if you need a stable, versioned dependency for a production training pipeline, because the README states there is not yet a stable version and all APIs are subject to change. Before committing, verify that the accelerator type matches the baseline's default num_cores, since the README warns that total batch size is often determined by num_cores.

Frequently asked questions

What is uncertainty in machine learning, and how does google/uncertainty-baselines address it?

The repository is a template for uncertainty and robustness research, providing implementations of standard and state-of-the-art methods on standard tasks along with prescribed benchmarking practices. It reports metrics such as test accuracy and test calibration error, the latter defined as expected calibration error over the test set.

How can uncertainty be measured in artificial intelligence?

According to the README, one reported metric is test calibration error, computed as expected calibration error over the test set following Naeini et al., 2015. Results across the baselines are reported to roughly three significant digits and averaged over 10 runs.

Does google/uncertainty-baselines have a stable release I can install from PyPI?

No. The README states there is not yet a stable version nor an official release of the library, and that all APIs are subject to change. The documented install pulls the development version from the GitHub repository.

Can I run google/uncertainty-baselines on GPUs instead of TPUs?

Yes, by changing the flags. The README shows running a baseline with --use_gpu=True and --num_cores=8, but warns that results may be similar while ultimately all bets are off, and that changing the number of cores matters a lot because total batch size is often determined by num_cores.

What do I need to install besides uncertainty_baselines itself?

Installing the package does not automatically install any backend. For TensorFlow the README lists TensorFlow, TensorFlow Addons and TensorBoard, with nightly variants as alternatives, and notes that setup.py lists the extra dependencies one can install.

Official sources

  1. google/uncertainty-baselines on GitHub
  2. Issues
  3. License: Apache-2.0
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/google-uncertainty-baselines.svg)](https://hysenlabs.com/projects/google-uncertainty-baselines)