# YDF (Yggdrasil Decision Forests): a C++ and Python library for training and serving decision forests

> Google's Yggdrasil Decision Forests trains, evaluates and exports Random Forest, Gradient Boosted Trees, CART and Isolation Forest models from Python or C++. It is a strong fit when you need one model format that runs in both a notebook and a production binary.

**google/yggdrasil-decision-forests** — A library to train, evaluate, interpret, and productionize decision forest models such as Random Forest and Gradient Boosted Decision Trees. 

- Repository: https://github.com/google/yggdrasil-decision-forests
- Website: https://ydf.readthedocs.io/
- Stars: 676 · Forks: 85
- Language: C++
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/google-yggdrasil-decision-forests

## What YDF is for, and who ends up using it

YDF is a library for decision forest models: Random Forest, Gradient Boosted Decision Trees, CART and Isolation Forest. The README states that it covers training, evaluation, interpretation and serving. Those four verbs describe four different jobs, and the library is aimed at teams that want all of them behind one model format rather than gluing a training library to a separate serving runtime.

The C++ core is the reason the project exists in this shape. A model trained from Python is saved with model.save("/tmp/my_model") and the same artifact is what the C++ API loads through SaveModel and the learner classes. For an engineer shipping a latency-sensitive service, that removes a translation step: no reimplementation of tree traversal in the serving language, no divergence between the notebook model and the deployed one.

It is a poor fit for someone who only wants a quick classifier on a small CSV and already has scikit-learn in the environment. The install footprint and the C++ build machinery are not free, and the project does not pretend to be a general-purpose ML framework. It does trees, and it does the surrounding work of describing, evaluating, analyzing and benchmarking them.

## How training, evaluation and export hang together

The Python surface is a thin layer over learners. ydf.GradientBoostedTreesLearner(label="income").train(train_ds) returns a model object, and every subsequent operation is a method on that object: describe, evaluate, predict, analyze, benchmark, save. The training configuration is expressed as constructor arguments plus a label column, which maps onto the C++ TrainingConfig, where the learner name is a string such as "RANDOM_FOREST" and the task is an enum such as Task::CLASSIFICATION set alongside set_label.

Data enters through a data specification. The C++ example calls CreateDataSpec(dataset_path, false, {}, &spec) on a path prefixed with csv:, which means the library reads the file, infers column types and returns a spec that the learner then consumes. In Python the same role is played by a Pandas DataFrame, so the type inference is already done by Pandas and YDF takes the frame directly.

The interpretability methods are part of the same object rather than a separate package. model.analyze(test_ds) is documented as covering partial dependence plots and variable importance, and model.benchmark(test_ds) measures inference speed. That combination matters because tree ensembles are often chosen precisely when someone has to justify the model to a reviewer, and having the explanation and the latency number come from the same code path as the prediction avoids arguments about whether the analysis reflects the deployed model.

## Installing YDF and training a first model

The README gives one install command for the Python package, published on PyPI as ydf. The -U flag upgrades an existing installation.

```bash
pip install ydf -U
```

The usage example loads the adult dataset from a URL in the repository's test_data directory, trains a Gradient Boosted Trees model on the income label, and then walks through describing, evaluating, predicting, analyzing, benchmarking and saving. Running it end to end is the fastest way to see what each method returns.

```python
import ydf
import pandas as pd

ds_path = "https://raw.githubusercontent.com/google/yggdrasil-decision-forests/main/yggdrasil_decision_forests/test_data/dataset/"
train_ds = pd.read_csv(ds_path + "adult_train.csv")
test_ds = pd.read_csv(ds_path + "adult_test.csv")

model = ydf.GradientBoostedTreesLearner(label="income").train(train_ds)
model.evaluate(test_ds)
model.save("/tmp/my_model")
```

model.evaluate(test_ds) is documented as producing metrics such as roc, accuracy, confusion matrix and confidence intervals. The save call writes the model to a directory, which is the artifact the C++ side reads back.

For the C++ path the repository ships examples/beginner.cc with a shell script and a batch file next to it, and the README points at examples/beginner.cc as the source of its C++ snippet. The build system is Bazel, visible in MODULE.bazel and .bazelversion at the repository root, with a configure/ directory for setup. The README itself does not walk through the Bazel invocation, so treat the example scripts as the entry point rather than guessing at targets.

## Where YDF gets in the way

The C++ distribution is the main friction point. There is a PyPI wheel for Python, and the repository carries Bazel files, a configure directory and third_party entries, but the README does not document a CMake build or a prebuilt C++ package. If your build environment cannot run Bazel, the C++ API is effectively out of reach, and you are left with the Python package and whatever export path you build around it.

Model portability has a boundary the README does not discuss. Saving to a directory and loading it from C++ is shown, but the README does not document a conversion path to ONNX or any other interchange format, nor does it describe rollback if a newly saved model turns out to be worse than the one it replaces. Version skew between the Python release and the C++ core is also unaddressed: the recent releases are tagged pydf_0.16.1, pydf_0.16.0 and pydf_0.15.0, and the README does not state which C++ revision pairs with which Python wheel.

The third limitation is scope. Isolation Forest is listed among the supported model types, but if your problem is sequence data, images or text embeddings, a tree ensemble is the wrong family regardless of how good the library is. The repository does include a text_classification.sh example, which suggests tokenized text features are workable, but that is feature engineering feeding trees, not a learned representation.

## YDF against TensorFlow Decision Forests and scikit-learn

TensorFlow Decision Forests is the closest relative, and the README is explicit that both projects share developers. The difference is the layer they attach to. TF-DF sits inside the TensorFlow graph and its ecosystem, which is the natural choice if the rest of your pipeline is already TensorFlow and you want the forest to participate in that graph. YDF is standalone: the Python API depends on Pandas for data and the model object carries its own evaluation and analysis methods, with no TensorFlow requirement stated in the install instructions. Choosing between them is mostly a question of what already surrounds the model.

scikit-learn's RandomForestClassifier and GradientBoostingClassifier are the other obvious comparison, and the difference is deployment rather than training. scikit-learn trains in Python and serves in Python, typically behind a pickle and a process boundary. YDF's C++ core means the same saved model can be embedded in a native binary. The trade-off runs the other way too: scikit-learn's ecosystem, its pipeline composition, its cross-validation helpers and its sheer volume of third-party documentation have no equivalent here. If your model never leaves a Python service, scikit-learn is the lower-friction option and YDF's C++ core buys you nothing.

## Maintenance, licensing and what an upgrade costs

The repository is not archived, and the last push was on 2026-09-10, ten days before this writing. The Python API has moved through three releases in the first half of 2026: pydf_0.15.0 on 2026-02-04, pydf_0.16.0 on 2026-03-17 and pydf_0.16.1 on 2026-03-26. A 0.x version series with a minor bump between 0.15.0 and 0.16.0 is worth reading the CHANGELOG.md for before you upgrade, because pre-1.0 libraries are free to change signatures. The CHANGELOG file is at the repository root, so the cost of checking is one file read.

The licence is Apache-2.0, stated in the README badge and in the LICENSE file, and the repository also carries a CITATION.cff. Apache-2.0 permits commercial use and modification and includes an explicit patent grant, which matters if you are embedding the library in a product. It also carries attribution and notice obligations that you should have a lawyer read rather than an article. The README asks that scientific publications cite the KDD 2023 paper by Guillame-Bert et al., which is a request rather than a licence term, but it is easy to honour.

Upgrade cost is concentrated in two places: the Python API surface between minor versions, and the Bazel toolchain pin in .bazelversion if you build the C++ side yourself.

## Conclusion

Adopt YDF if you need decision forest models that train in a Python notebook and then load in a C++ service, and you accept the Apache-2.0 terms and the Bazel build path for the C++ side. Do not adopt it if you need a Python-only workflow with deep integration into the scikit-learn ecosystem, or if you cannot run Bazel in your build environment. Before committing, check the CHANGELOG for the API surface that changed between the 0.15.0 and 0.16.1 Python releases, and confirm that examples/beginner.sh still builds against your toolchain.

## FAQ

### What are decision forests?

They are the model family YDF covers: the README lists Random Forest, Gradient Boosted Decision Trees, CART and Isolation Forest, all of which the library can train, evaluate, interpret and serve. A random forest is an ensemble of decision trees, while a single decision tree is one tree.

### What is Yggdrasil?

In this context Yggdrasil Decision Forests, abbreviated YDF, is the Google library for training and serving decision forest models. It is distributed on PyPI as ydf and documents a Python and a C++ API.

### What is the difference between a random forest and a decision tree?

YDF treats them as separate learner choices: the C++ example sets the learner to "RANDOM_FOREST" with a classification task and a label column, and the README lists Random Forest, Gradient Boosted Decision Trees, CART and Isolation Forest as the model types the library supports.

### Is random forest ml or ai?

The README does not frame the question that way. It describes YDF as a library to train, evaluate, interpret and serve decision forest models, and lists Random Forest among the supported model types.

## Sources

- [google/yggdrasil-decision-forests on GitHub](https://github.com/google/yggdrasil-decision-forests)
- [License: Apache-2.0](https://github.com/google/yggdrasil-decision-forests/blob/main/LICENSE)
- [Project website](https://ydf.readthedocs.io/)
- [README](https://github.com/google/yggdrasil-decision-forests/blob/main/README.md)
- [Releases](https://github.com/google/yggdrasil-decision-forests/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/google-yggdrasil-decision-forests
