# ngboost: the README says Boston housing, the example loads California

> stanfordmlgroup/ngboost turns scikit-learn boosting into probabilistic prediction by learning a proper scoring rule, with distributions and base learners as pluggable pieces. Its own example, version ranges and release tooling disagree with each other in small, checkable ways.

**stanfordmlgroup/ngboost** — Natural Gradient Boosting for Probabilistic Prediction

- Repository: https://github.com/stanfordmlgroup/ngboost
- Stars: 1,891 · Forks: 256
- Language: Jupyter Notebook
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/stanfordmlgroup-ngboost

## The example is labelled Boston and loads California

The usage section introduces itself as a probabilistic regression example on the Boston housing dataset. The code underneath imports a California housing fetcher, assigns the data and target, splits with a test size of two tenths, fits a regressor, then predicts and asks for the predictive distribution:

```python
from ngboost import NGBRegressor

from sklearn.datasets import fetch_california_housing
from sklearn.model_selection import train_test_split
from sklearn.metrics import mean_squared_error

# Load California housing dataset
cal = fetch_california_housing()
X, Y = cal.data, cal.target

X_train, X_test, Y_train, Y_test = train_test_split(X, Y, test_size=0.2)

ngb = NGBRegressor().fit(X_train, Y_train)
Y_preds = ngb.predict(X_test)
Y_dists = ngb.pred_dist(X_test)

# test Mean Squared Error
test_MSE = mean_squared_error(Y_preds, Y_test)
print('Test MSE', test_MSE)

# test Negative Log Lik
```

Two more loose ends in the same block. The predictive distribution is assigned to a variable nothing then reads, and the last line is a comment announcing a negative log likelihood check with no code under it. The snippet prints a metric but shows no output, so nothing in the repository demonstrates what the distribution object contains.

## Two installers in one shell block

Installation is a single fenced block that mixes labels with commands:

```sh
via pip

pip install --upgrade ngboost

via conda-forge

conda install -c conda-forge ngboost
```

The two routes are the obvious ones and both are upgrade-or-install rather than pinned, so the version you get is whatever the index holds at that moment. Nothing in the block mentions a virtual environment, and the plain-language route labels sit inside the fence where a shell would try to execute them, which suggests the block was written as prose and formatted as a script. The package is published under the name ngboost, and the README links its download counts from a package statistics page, so at least the intent to publish is unambiguous.

## Four dependencies change version range on the Python version

The manifest supports Python from 3.9 up to but not including 3.15, and four of its dependencies are declared twice with different ranges on either side of a Python version marker. Scikit-learn splits at 3.14, taking 1.6 to 1.7 below it and 1.8 or newer above. NumPy splits at 3.13, taking 1.21.2 or newer below and 2.1 or newer above. SciPy splits at the same 3.13 mark, taking 1.7.2 below and 1.14.1 above. Matplotlib splits earliest, at 3.10, taking 3.0 up to 3.10 below and 3.10 or newer above. Three different cut points across four libraries in one dependency block. The rest are flat: a progress bar from 4.3, lifelines from 0.25 for survival analysis, sympy from 1.12, and joblib from 1.2. The practical effect is that the resolved numerical stack for a 3.9 interpreter and a 3.14 interpreter differ in more than the interpreter.

## The manifest is a dev version and the tag is not

The manifest on the default branch declares the version as 0.5.11dev while the newest release is tagged v0.5.11, so the branch you clone is a pre-release of the version already published. The three most recent release names describe what each one was for: bug fixes in the APIs in June 2026, backwards compatibility in March, and a Sympy factory alongside Python 3.14 support in February. That last one lines up with the dependency work in the manifest, since a Sympy floor of 1.12 and the 3.13 splits are the kind of change that arrives with a new interpreter. Two smaller things sit alongside. The license field is written as a long-form string rather than the short SPDX identifier, and the classifier list holds exactly one entry, the operating system one, so nothing advertises a Python version or a supported framework even though the Python constraint is precise.

## Four linters are installed and the lint target runs one wrapper

The development dependency group lists a test runner and five style tools: pytest in the 8 line, pre-commit in the 4 line, and then black, isort, pylint and flake8. Black and isort overlap almost completely by design, and pylint and flake8 are the older pairing that black and isort were meant to replace, so four checkers are configured against one codebase. The lint target does not call any of them directly. It runs pre-commit across all files at the manual hook stage, which means what actually executes is whatever the pre-commit configuration wires in, and the four tools are reachable only if that configuration names them. The black settings set a line length of 88 and list target versions from 3.9 through 3.13, one short of the newest interpreter the package allows.

## The publish target sources an env file through shell syntax

Five targets, all of them poetry. Install pins the tool itself to an exact version before installing the project. Package builds the distributions. Publish depends on package, then sources a local env file, sets a package index token from an environment variable through the poetry config command, and publishes. Lint runs the pre-commit wrapper described above, test runs pytest in verbose mode with a slow flag, and clean empties the distribution directory. The publish recipe is the one to note. Sourcing a file is shell syntax that make does not guarantee, since make runs recipes through its own default shell, and the env file it reads is not in the repository, so a publish on a fresh clone fails before reaching the index. That is a safe failure, but it is undocumented in the file itself.

## Six distribution examples, and committed results

The examples directory is the widest evidence of what the library actually covers, and it is broader than the single regression snippet in the README: plain regression, multiclass classification, multivariate normal output, Poisson output, survival analysis, a classifier example, a cross-validation example, and subdirectories for experiments, simulations, model interpretation, tuning and the user guide. Survival is the reason a lifelines dependency sits in the manifest. Around those sit directories for data, results and figures at the repository root, which means experiment output is committed alongside the code, and a notebooks directory, which is also why the repository's recorded primary language is a notebook format rather than Python. The README itself defers the substance: distributions, scoring rules, learners, tuning and interpretation are all pointed at an external user guide, with the note that it also explains how to add a new distribution or score.

## Conclusion

ngboost is the right shape of library for anyone who needs calibrated distributions rather than a point prediction, because the distribution, the scoring rule and the base learner are separate seams, and the examples directory shows the range the seams cover, from regression and multiclass through Poisson, multivariate normal and survival. Three things to check before you pin it. Four of its dependencies switch version ranges on the Python version, at three different split points, so the resolved stack differs by interpreter. The publish target reads a token out of a local env file through shell syntax the make default shell does not supply. And the headline example is labelled with one dataset and loads another, so treat the prose in this README as approximate and the code as the specification.

## FAQ

### What is ngboost?

A Python library implementing Natural Gradient Boosting for probabilistic prediction, built on top of scikit-learn and designed to be modular across the choice of proper scoring rule, distribution and base learner.

### How do you install ngboost?

Two routes, shown in one block: pip install --upgrade ngboost, or conda install -c conda-forge ngboost. Neither pins a version, and the block includes plain-language route labels inside the shell fence.

### Which Python versions does ngboost support?

From 3.9 up to but not including 3.15. Four dependencies switch ranges on the interpreter: scikit-learn at 3.14, numpy and scipy at 3.13, and matplotlib at 3.10, so the resolved numerical stack differs by Python version.

### How does the ngboost example evaluate its predictions?

It fits a regressor, calls predict and pred_dist, then prints a test mean squared error. The block ends on a comment announcing a negative log likelihood check with no code beneath it, and the predictive distribution it computed is assigned to a variable nothing reads.

### Is ngboost compared with LightGBM or XGBoost anywhere in the repository?

The README does not compare it with either. It states what ngboost is, a library for probabilistic prediction built on scikit-learn and modular across scoring rule, distribution and base learner, and then points readers to an external user guide for distributions, scoring rules, learners, tuning and interpretation.

## Sources

- [Issues](https://github.com/stanfordmlgroup/ngboost/issues)
- [License: Apache-2.0](https://github.com/stanfordmlgroup/ngboost/blob/master/LICENSE)
- [README](https://github.com/stanfordmlgroup/ngboost/blob/master/README.md)
- [Releases](https://github.com/stanfordmlgroup/ngboost/releases)
- [stanfordmlgroup/ngboost on GitHub](https://github.com/stanfordmlgroup/ngboost)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/stanfordmlgroup-ngboost
