bayesian-machine-learning: thirteen notebooks, seven dependency sets, and a Zenodo DOI instead of a release
Notebooks about Bayesian methods for machine learning
At a glance
- What is it?
- This is a teaching collection rather than a library. It holds thirteen notebooks on Bayesian methods across seven topic directories, and its most interesting structural decision is that each directory carries its own requirements file, which is what lets one repository hold notebooks written against three different probabilistic programming stacks. Its citable identity is a Zenodo DOI, and its releases stopped in December 2020 while the branch kept moving.
- Who is it for?
- Use this for what it is: worked examples you read and rerun, with the arithmetic visible and the framework optional. The pattern of implementing each idea twice, once in plain NumPy or SciPy and once with a library, is the part worth copying, because the first version shows you what the library is doing for you.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 84 days ago.
- What is it written in?
- Mainly Jupyter Notebook, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The first thing on the page is a Zenodo DOI, not a badge
The very first element under the title is a link to a Zenodo record, and it is the only citation apparatus the page carries. There is no paper reference, no author list, and no bibtex entry.
For a collection of notebooks this is the correct mechanism rather than an odd one. Software gets versioned by releases and installed by a package manager. A set of teaching notebooks gets neither, and the thing it actually needs is a stable, citable identity that survives the repository moving or disappearing. A concept DOI on an archive record is exactly that, and it also gives the collection a version history independent of the three GitHub releases.
So the project has two version lines pointing at different things: an archive deposit that is meant to be cited, and a repository that is meant to be read. The page does not explain what the deposit contains or which commit it corresponds to, so if you are citing rather than reading, that mapping is something you have to establish yourself.
The license is the permissive one, Apache-2.0, with a license file at the top of the repository. That matters more than usual for teaching material, since the likely uses are teaching from it and reusing the code in a course or a workshop.
Three releases from 2020, a branch called dev, and every link hardcodes it
The releases are v-0.1, v-0.2 and v-0.3, published in August, September and December 2020. The tag names carry a hyphen between the letter and the number, which is a small unusual detail, and the last of them is close to six years old.
The branch is a different matter. The default branch is called dev, and every single link in the documentation is written against that name. Each notebook URL spells out the repository, the owner, the dev branch, the topic directory and the notebook file, with no shorthand anywhere. That makes the page maximally explicit and maximally fragile at the same time: rename the branch and all of them break at once, and there are more than a dozen of them to fix by hand.
Commits did not stop where the releases did. The default branch was last pushed on 2026-07-12, so the notebooks have been touched within the last few months even though nothing has been published since 2020. For a notebook collection that is a reasonable shape of change, because a notebook edit is usually a fix to an explanation rather than a change to an interface, and it explains why the archive deposit and the GitHub releases are both stale relative to the branch.
The practical consequence is that there is no release to check out for a class, and the Zenodo deposit is the only frozen artefact. If you are teaching from this, pin the branch you teach and record the commit.
Seven directories, thirteen notebooks, eleven topics, seven requirements files
The structure is one directory per topic area, and the counts do not line up in the way you would expect, which makes them worth stating once.
There are seven topic directories: autoencoder applications, Bayesian linear regression, Bayesian neural networks, Bayesian optimization, Gaussian processes, latent variable models, and noise contrastive priors. Inside them sit thirteen notebooks covering eleven distinct topics, because two topics have a second implementation and one directory holds three topics.
That gives seven requirements files, one per directory, and the page states the reason plainly: dependencies are specified in requirements files in the subdirectories. This is the structural decision that makes the collection possible at all. The notebooks are not written against one stack. They span NumPy and SciPy for the from-scratch versions, scikit-learn and GPy for Gaussian processes, scikit-optimize and GPyOpt for optimization, PyMC3 for two probabilistic programming implementations, TensorFlow 2 and TensorFlow Probability for variational inference, JAX for sparse variational Gaussian processes, and Keras for the neural network and autoencoder material.
Eleven libraries, three generations of tooling, one repository. A single top-level requirements file could not express that without forcing a version negotiation across stacks that were never meant to coexist, so the collection pushes the problem down into the directory and lets a reader install only the one topic they are reading.
The contents of those seven files are not shown on the page, and no Python version is stated anywhere, so the resolution work is left to the reader.
Most topics are implemented twice, once from scratch and once with a library
The pattern that runs through the collection is deliberate duplication. Bayesian linear regression appears as a plain NumPy and scikit-learn implementation and again as a PyMC3 implementation. Gaussian processes for regression are done with plain NumPy and SciPy and also with scikit-learn and GPy. The Gaussian process classification notebook repeats the same pairing. Latent variable models part one, on Gaussian mixture models and the expectation maximization algorithm, is done with plain NumPy and SciPy plus scikit-learn, and again with PyMC3.
So the from-scratch version comes first in the reading order every time, and the library version is the second link. For a reader, that ordering is the lesson: you see the model as arithmetic before you see it as a call. For a maintainer, it is a maintenance cost, because two implementations of the same model can and will drift apart, and nothing in the repository checks that they still agree.
Two topics break the pattern by using a newer style of library. Sparse Gaussian processes use a variational approach with JAX, and Bayesian optimization is implemented with plain NumPy and SciPy as well as with scikit-optimize and GPyOpt, with hyper-parameter tuning as the stated application example.
The naming also carries history. The probabilistic programming library is called PyMC3 rather than the shorter later name, and the TensorFlow notebooks are labelled with the major version only, one as Tensorflow 2 and another as Tensorflow 2.x. Since nothing on the page pins a version, which of these stacks still installs cleanly for a new reader is the first practical question the documentation leaves open.
One notebook is about a Bayesian network that is confidently wrong
Most entries in a Bayesian machine learning collection demonstrate that the method works. One of these is about the case where it does not, and it is the most substantial single piece of content on the page.
The notebook is titled as being about reliable uncertainty estimates for neural network predictions, and its description says it uses noise contrastive priors for Bayesian neural networks to get more reliable uncertainty estimates for out-of-distribution data. It is implemented with TensorFlow 2 and TensorFlow Probability.
The framing is the interesting part. A Bayesian neural network gives you a predictive distribution rather than a point estimate, which is normally sold as the solution to knowing when the model does not know. The premise here is that the uncertainty this produces can still be badly wrong on data outside the training distribution, and that a prior choice is what fixes it. Noise contrastive priors are the mechanism, and TensorFlow Probability is what supplies the variational machinery.
That makes this a notebook about a failure mode rather than a technique, and it is worth reading next to the variational inference notebook in the Bayesian neural networks directory, which uses Keras to demonstrate the machinery that the other one then questions.
It is also the one notebook whose stated purpose is a negative result, which is what makes it the most useful thing in the collection for anyone who has watched a Bayesian model produce a confident answer on an input it had never seen.
Bayesian optimization appears twice and closes the collection
The last entry in the list uses Bayesian optimization in a way that only makes sense if you have read the optimization notebook earlier, which gives the collection a small arc rather than being an unordered pile.
The optimization notebook is introduced as an introduction to the method with an application example of hyper-parameter tuning. That is the standard framing: treat the search over hyper-parameters as the function you are optimising. The final notebook, conditional generation via Bayesian optimization in latent space, takes a variational autoencoder that has already learned a latent space and searches inside it, which is described as an approach for conditionally generating outputs with desired properties.
The variational autoencoder also appears twice more, so it is the most revisited idea in the collection and the one it uses as a component. It is introduced in latent variable models part two, on stochastic variational inference with a variational autoencoder as the application, implemented with TensorFlow 2.x. It returns in the autoencoder applications directory for a deep feature consistent version, which describes how a perceptual loss improves the quality of generated images, implemented with Keras. And it returns a third time as the search space for the final notebook, with Keras and GPyOpt.
So the autoencoder appears in both TensorFlow and Keras flavours across two directories, and the last notebook depends on both an autoencoder and an optimizer. If you read the collection in order rather than by topic, that is the thread: probabilistic models first, variational inference second, and then the autoencoder as a space you can search.
nbviewer for the formulas, a hosted service for six topics, and one cache buster
The page solves two presentation problems and states one of them. It says the links display the notebooks through a viewer to ensure a proper rendering of formulas, which is the honest reason notebooks with mathematics get linked to a rendering service instead of to the repository: the raw file does not show you the equations.
The second problem is compute. Six of the eleven topics also carry a link to a hosted notebook service, which is the route for the ones where running the code needs more than a laptop: the three Gaussian process notebooks, the optimization notebook, and both latent variable model notebooks. The five remaining topics, Bayesian linear regression, the Bayesian neural network notebook, the noise contrastive priors notebook and the two autoencoder notebooks, have no such link and are read-only.
That pairing is done with an unlabelled image link next to a labelled text link for each hosted route, so in the rendered page you get two badges and in a text reader you get one. That is a small thing, and it means the compute route for the heavier notebooks is only discoverable if your reader renders images.
One detail is a genuine loose end: one of the viewer links carries a cache-busting query parameter and the other ten do not. Which notebook it is, is not explained, so a reader who hits a stale rendering on that one file learns nothing about why it happens there and not elsewhere.
Underneath all of this, the only guidance for running anything is the sentence about dependencies in the subdirectories. There is no install command on the page, no environment specification, and no claim that any notebook has been executed as part of a check.
Editorial conclusion
Use this for what it is: worked examples you read and rerun, with the arithmetic visible and the framework optional. The pattern of implementing each idea twice, once in plain NumPy or SciPy and once with a library, is the part worth copying, because the first version shows you what the library is doing for you. Three things to check before you plan an afternoon with it. First, nothing here is installed for you: dependencies live in requirements files inside the seven topic directories, their contents are not shown on the page, and there is no Python version or framework version stated anywhere in the documentation, so expect to resolve a stack that spans three eras of tooling. The Zenodo DOI is what you cite, since the three releases stopped in December 2020 and the branch has been pushed as recently as July 2026. And there is no automated check that any notebook still runs, so a notebook that fails halfway through is a documentation gap rather than a regression signal. For learning Gaussian processes, variational inference or latent variable models from code you can read end to end, that trade is clearly worth it. For a dependency you want to add to a project, it is the wrong shape entirely, because there is no library here to install and nothing pinning what the notebooks expect.
Frequently asked questions
What is in the krasserm/bayesian-machine-learning collection?
Thirteen notebooks across seven topic directories and eleven topics, covering Bayesian linear regression, Gaussian processes for regression and classification, sparse Gaussian processes, Bayesian optimization, variational inference in Bayesian neural networks, latent variable models, and three variational autoencoder applications.
How do I run the notebooks in bayesian-machine-learning?
Dependencies are specified in requirements files inside the subdirectories, one per topic directory, and the page gives no install command and states no Python version. Six of the eleven topics also link to a hosted notebook service for the heavier examples.
Which libraries do the Bayesian machine learning notebooks use?
Eleven across the collection: NumPy, SciPy, scikit-learn, GPy, scikit-optimize, GPyOpt, PyMC3, TensorFlow 2, TensorFlow Probability, JAX and Keras. Several topics are implemented twice, once in plain NumPy or SciPy and once with a library.
What does the noise contrastive priors notebook demonstrate?
It applies noise contrastive priors to Bayesian neural networks to get more reliable uncertainty estimates for out-of-distribution data, implemented with TensorFlow 2 and TensorFlow Probability. It is the one entry framed around a failure mode rather than a technique.
How current is the bayesian-machine-learning collection?
The default branch was last pushed on 2026-07-12, while the three releases date from August, September and December 2020. A Zenodo DOI is the versioned citation for the collection, and the code is licensed Apache-2.0.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/krasserm-bayesian-machine-learning)