TensorFlow Probability: probabilistic layers, bijectors and inference on top of TensorFlow
Probabilistic reasoning and statistical analysis in TensorFlow
At a glance
- What is it?
- TensorFlow Probability is a library for probabilistic reasoning and statistical analysis inside the TensorFlow ecosystem, with a JAX substrate as well. The useful question is not what it is, but which layer of it you actually need.
- Who is it for?
- Adopt TensorFlow Probability when your model needs uncertainty over functions or parameters and you already train in TensorFlow: tfp.layers and tfp.distributions slot into an existing graph, and the MCMC and variational inference modules cover the two standard routes to a posterior. Do not adopt it if you want a standalone sampler with a modelling language, or if you only need descriptive statistics, because the library assumes a TensorFlow (or JAX) computation throughout.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 6 days ago.
- What is it written in?
- Mainly Jupyter Notebook, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What TensorFlow Probability adds that plain TensorFlow does not
TensorFlow gives you tensors, gradients and hardware acceleration. It does not give you a probability distribution as a first-class object, and it does not give you a way to express uncertainty over a model's parameters or over the function a neural network represents. TensorFlow Probability fills that gap. The README describes the library as providing integration of probabilistic methods with deep networks, gradient-based inference via automatic differentiation, and scalability to large datasets and models via hardware acceleration and distributed computation. The audience is therefore narrower than the word probability suggests: this is for people building statistical models or Bayesian neural networks inside TensorFlow, not for anyone who wants to compute the odds of a dice roll. The repository's own topics list bayesian-methods, probabilistic-programming and statistics alongside deep-learning and neural-networks, which matches that reading.
The four-layer structure, from LinearOperator to transition kernels
The README lays the library out as four layers, and the split is the most useful thing to understand before writing code. Layer 0 is TensorFlow itself, and specifically the LinearOperator class, which the README says enables matrix-free implementations that can exploit special structure such as diagonal or low-rank matrices; it now lives in tf.linalg in core TensorFlow and is maintained by the TensorFlow Probability team. Layer 1 is statistical building blocks: tfp.distributions, a collection of distributions with batch and broadcasting semantics, and tfp.bijectors, reversible and composable transformations of random variables, which the README describes as the route from classical cases like the log-normal distribution to masked autoregressive flows. Layer 2 is model building: joint distributions such as tfp.distributions.JointDistributionSequential for one or more possibly interdependent distributions, and tfp.layers, neural network layers with uncertainty over the functions they represent. Layer 3 is inference: tfp.mcmc for approximating integrals via sampling (Hamiltonian Monte Carlo, random-walk Metropolis-Hastings, and custom transition kernels), tfp.vi for approximating integrals via optimization, tfp.optimizer for stochastic optimization methods including Stochastic Gradient Langevin Dynamics, and tfp.monte_carlo for Monte Carlo expectations. The layering is not decoration. A bijector in layer 1 becomes the support transformation for a variational posterior in layer 3, and a LinearOperator from layer 0 is what makes a large covariance matrix tractable in between.
Installing TensorFlow Probability and sampling from a distribution
The repository ships a setup.py whose docstring is "Install tensorflow_probability." That file distinguishes two packages: with the --release flag it builds project_name tensorflow-probability, and without it the default is tfp-nightly. The release path pins TF_PACKAGE to tensorflow >= 2.16 and KERAS_PACKAGE to tf-keras >= 2.16, and TFDS_PACKAGE to tensorflow-datasets >= 2.2.0; the nightly path uses tfds-nightly instead. The README's homepage points at https://www.tensorflow.org/probability/, and the README does not spell out a pip command, so treat the package names above as the authoritative identifiers rather than guessing an install line. Once installed, the smallest real use is drawing samples and evaluating a density. The README's own description of the distributions module is the guide here: distributions carry batch and broadcasting semantics, so the shape of the sample you get back depends on the batch shape you asked for.
A first model: a joint distribution and a sampler
The README points to a colab called Modeling with JointDistribution for an introduction to joint distributions, and to the Eight Schools notebook as a hierarchical normal model for exchangeable treatment effects. The pattern those examples follow is: declare a joint distribution over parameters and data, condition on observed data, then hand the resulting log-probability function to an inference algorithm. In tfp.mcmc terms that means a transition kernel plus a sampling loop; the README lists Hamiltonian Monte Carlo and random-walk Metropolis-Hastings as the built-in kernels and notes that custom transition kernels can be written. Two practical constraints follow from that structure. First, every parameter you want to infer must be a TensorFlow variable or a tensor the sampler can trace, because gradients come from automatic differentiation. Second, the shape conventions of tfp.distributions propagate into the target log-probability, so a batch-shape mistake surfaces as a sampling error rather than a clear shape warning. The repository includes a notebook titled Understanding TensorFlow Distributions Shapes precisely because that distinction between sample, batch and event dimensions trips people up.
The JAX substrate and what it does not share
The README states that TFP also works as "Tensor-friendly Probability" in pure JAX, with the import from tensorflow_probability.substrates import jax as tfp, and links to a page on TensorFlow Probability on JAX. There is a SUBSTRATES.md file at the repository root, so the substrate mechanism is a documented part of the project rather than an experiment. The important caveat is that a substrate is not the whole library. Code written against the TensorFlow substrate uses TensorFlow ops, TensorFlow variables and TensorFlow's gradient tape; the JAX substrate replaces that substrate, not the API surface above it. If you are choosing between the two, decide before you write model code, because mixing tensor types from both frameworks in one computation is not something the README describes as supported.
Where TensorFlow Probability is the wrong tool
The README carries an explicit warning: TensorFlow Probability is under active development and interfaces may change at any time. That is a real cost for production code that pins a version, and it is the reason the release history matters. The most recent release listed is v0.25.0 from 2024-11-08, preceded by v0.24.0 in March 2024 and v0.23.0 in November 2023, so the tagged release cadence is roughly two per year even though the default branch sees commits. If your team needs a stable API with a long deprecation window, that cadence is a constraint to weigh. There are also cases where the library is simply the wrong shape of tool. If you want to write a model in a probabilistic programming language and get a sampler without touching tensors, this is not that: you will be writing TensorFlow. If your problem is a small conjugate model, the full machinery of tfp.mcmc and tfp.vi is more than you need. And if your data pipeline is already in another array framework, adopting TensorFlow Probability means adopting a second numerical stack.
Alternatives and the difference in approach
The README itself contains the most honest comparison available: a notebook named HLM_TFP_R_Stan, described as comparing hierarchical linear models among TensorFlow Probability, R, and Stan. That is the right frame. Stan and its R interface express a model in a dedicated modelling language and compile it to a sampler; you describe the posterior, and the toolchain handles the inference. TensorFlow Probability inverts that. You build the model out of distributions, bijectors and layers as ordinary TensorFlow objects, and you assemble the inference loop yourself from a transition kernel and a sampling function. The payoff is composability: a tfp.layers layer can sit inside a Keras model and be trained by gradient descent like any other layer, and the same distribution objects can be reused across a variational fit and an MCMC fit. The cost is that you own more of the wiring, including shape conventions and the choice of kernel. For a standalone Bayesian regression with no neural network, Stan is less code. For a model where the probabilistic part is one component of a larger differentiable system, TensorFlow Probability is the better fit, and the JAX substrate gives you an exit if you later move off TensorFlow.
Licence, maintenance and the cost of upgrading
The repository is licensed Apache-2.0, with the LICENSE file at the root and the standard Apache header reproduced at the top of setup.py. Apache-2.0 permits commercial use and modification and includes an explicit patent grant; it also requires that you preserve the licence and notice files. That is a description of the licence text, not legal advice, and if you redistribute a modified copy you should read the LICENSE file rather than this paragraph. On maintenance, the default branch last received a push on 2026-09-09 and the repository is not archived, so work is ongoing, but the README's own statement that interfaces may change at any time is the operative fact for planning. Upgrading between releases is not a drop-in operation if you depend on inference internals. The release path in setup.py pins tensorflow >= 2.16 and tf-keras >= 2.16, which means a TensorFlow major upgrade can force a TensorFlow Probability upgrade in turn, and the tf-keras dependency is a separate package you will need to keep aligned. Budget for reading release notes before bumping the pin, and prefer the tagged tensorflow-probability package over tfp-nightly for anything you intend to keep running.
Editorial conclusion
Adopt TensorFlow Probability when your model needs uncertainty over functions or parameters and you already train in TensorFlow: tfp.layers and tfp.distributions slot into an existing graph, and the MCMC and variational inference modules cover the two standard routes to a posterior. Do not adopt it if you want a standalone sampler with a modelling language, or if you only need descriptive statistics, because the library assumes a TensorFlow (or JAX) computation throughout. Before committing, verify three things: that your TensorFlow version satisfies the tensorflow >= 2.16 requirement in setup.py, that the interfaces your code touches are acceptable given the README's statement that they may change at any time, and that the joint distribution class you pick (JointDistributionSequential or another) matches how your variables depend on each other.
Frequently asked questions
What is TensorFlow Probability in simple terms?
It is a library for probabilistic reasoning and statistical analysis inside TensorFlow, providing probability distributions, reversible transformations of random variables, probabilistic neural network layers, and inference algorithms such as Markov chain Monte Carlo and variational inference. The README also notes it works as a JAX substrate.
What are examples of what TensorFlow Probability can model?
The examples directory includes notebooks for linear mixed effects models, the Eight Schools hierarchical normal model, Bayesian Gaussian mixture models for clustering, probabilistic principal components analysis, and Gaussian copulas for dependence across random variables. There are also example scripts such as a vector-quantized autoencoder and a disentangled sequential variational autoencoder.
How do I use probability distributions in TensorFlow Probability?
The tfp.distributions module holds a large collection of probability distributions and related statistics with batch and broadcasting semantics. The README links a Distributions Tutorial and a separate notebook on distinguishing sample, batch and event shapes, which is where most shape errors originate.
How do I calculate probability with TensorFlow Probability?
You build the computation out of tfp.distributions objects rather than closed-form arithmetic: declare a distribution, then call its methods to evaluate the density or draw samples. The README points to a Distributions Tutorial and to a notebook on distribution shapes for the conventions involved.
What is the basic formula for probability in a TensorFlow Probability model?
There is no single formula; the library expresses probability through distribution objects and joint distributions such as tfp.distributions.JointDistributionSequential, which the README describes as joint distributions over one or more possibly interdependent distributions. Inference then approximates integrals, either via sampling in tfp.mcmc or via optimization in tfp.vi.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/tensorflow-probability)