CLI tool
tensorflow/decision-forests avatar
tensorflow/decision-forests

TensorFlow Decision Forests: Keras-Native Random Forests and Gradient Boosted Trees

A collection of state-of-the-art algorithms for the training, serving and interpretation of Decision Forest models in Keras.

693 stars115 forksPythonApache-2.0

At a glance

What is it?
TF-DF trains decision forest models inside Keras and exports them as SavedModels, but its own README now sends new users to Yggdrasil Decision Forests instead. Here is what it still does well and where it stops being the right choice.
Who is it for?
Adopt TF-DF if you already have a Keras or TensorFlow pipeline, need a forest model that can be saved as a SavedModel alongside neural networks, or want a tree-based baseline without leaving Python. Do not adopt it for a new project if you have no TensorFlow dependency: the README itself recommends migrating to Yggdrasil Decision Forests, which trains the same models and is described as faster with more functionality.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 133 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What TF-DF Solves That scikit-learn Does Not

The pitch is narrower than "decision forests in Python", because scikit-learn already covers that. What TF-DF adds is placement inside the TensorFlow graph. A forest trained with tfdf.keras.RandomForestModel is a Keras model, so it takes tf.data datasets, participates in the same fit and evaluate calls as a neural network, and saves as a TensorFlow SavedModel. The README states that TF-DF supports classification, regression and ranking, and that the models are compatible with Yggdrasil Decision Forests in both directions.

That compatibility matters more than it first appears. A team can train in Python, hand the exported model to a serving stack, and convert it to a YDF model if they later want the C++, JavaScript, CLI or Go runtimes. The audience is therefore engineers who already live in TensorFlow and want a tree baseline that ships through the same artifact pipeline as the rest of their models, not people shopping for a general-purpose gradient boosting library.

How TF-DF Sits on Top of the Yggdrasil C++ Core

TF-DF is a Python and Keras layer over the Yggdrasil Decision Forests library, which is written in C++ and also exposes JavaScript, CLI and Go interfaces. The training algorithms live in that core; the Keras classes are the front end. This explains several design details at once: why the pip wheel is platform-specific, why the model format is portable to YDF, and why the README frames TF-DF and YDF as two doors into the same engine.

The data flow in the README example is short. A Pandas dataframe goes through tfdf.keras.pd_dataframe_to_tf_dataset with a label argument, producing a tf.data dataset. The model is constructed, fit on that dataset, inspected with model.summary(), evaluated on a held-out dataset, and written to disk with model.save(). The repository also ships examples/minimal.py, examples/hyperparameter_optimization.py, examples/distributed_training.py and examples/distributed_hyperparameter_optimization.py, so distributed training and hyperparameter search are part of the intended surface rather than afterthoughts.

One consequence of the Keras wrapper is worth stating plainly: the forest is a Keras model, so anything you would do to a Keras model (callbacks, subclassing, saving hooks) is on the table, and so are the constraints that come with Keras version churn.

Installing tensorflow_decision_forests and Training a First Forest

The README gives a single install command. It upgrades an existing installation if one is present.

bash
pip3 install tensorflow_decision_forests --upgrade

The README points to documentation/installation.md for troubleshooting and alternative installation paths, which is where to look if the wheel does not match your platform. Platform support is stated explicitly: Linux and macOS are supported, and Windows users are directed to WSL plus Linux.

A first real run follows the README example closely. Two CSV files are loaded into Pandas, converted to tf.data datasets with the label column named, then a random forest is trained and evaluated.

python
import tensorflow_decision_forests as tfdf
import pandas as pd

train_df = pd.read_csv("project/train.csv")
test_df = pd.read_csv("project/test.csv")

train_ds = tfdf.keras.pd_dataframe_to_tf_dataset(train_df, label="my_label")
test_ds = tfdf.keras.pd_dataframe_to_tf_dataset(test_df, label="my_label")

model = tfdf.keras.RandomForestModel()
model.fit(train_ds)
model.summary()
model.evaluate(test_ds)
model.save("project/model")

After fit, model.summary() prints the model structure, and model.evaluate(test_ds) returns the metrics for the test split. The final line writes a TensorFlow SavedModel to project/model, which the README notes is compatible with Yggdrasil Decision Forests. For a gradient boosted trees variant, the same file structure applies with a different model class from the tfdf.keras namespace; the README's usage example uses RandomForestModel, so treat that as the verified starting point.

The README Tells New Users to Leave

The most important line in the repository is the note at the top: the maintainers recommend migrating to Yggdrasil Decision Forests, stating that YDF trains the same models as TF-DF but is faster and has more functionality. A project whose own front page points elsewhere is not a neutral signal, and anyone evaluating TF-DF for new work should read that note before anything else.

This does not make TF-DF broken. It makes it a compatibility and integration layer for people already inside TensorFlow. If your reason for choosing it is "I want decision forests in Python", the recommendation points you at the sibling library instead. If your reason is "I have a Keras pipeline and need a forest in it", the layer is exactly what you want and the migration note is a forward-looking statement rather than a defect report.

A second limitation is platform. Linux and macOS only, with Windows routed through WSL plus Linux. Teams on native Windows builds have no supported path in the README.

A third is version coupling. The pip package wraps TensorFlow and Keras, so the Keras 3 changes visible in the release history can affect code that subclasses or extends the model classes. The README does not document rollback or downgrade procedures; documentation/known_issues.md and CHANGELOG.md are the places the repository keeps that kind of information.

TF-DF Versus Plain scikit-learn Forests

The honest comparison is scikit-learn's RandomForestClassifier and GradientBoostingClassifier, which also train forests in Python and have no TensorFlow dependency at all. The difference is not the algorithm family; it is the artifact and the runtime. A scikit-learn forest pickles, and serving it usually means a Python process or a conversion step. A TF-DF forest is a SavedModel, which means it enters the same deployment path as any other TensorFlow model, including TensorFlow Serving style workflows and graph-level tooling.

The cost side is real. Installing TF-DF pulls in TensorFlow, which is a much larger dependency than scikit-learn. If your only goal is a tabular baseline in a notebook, scikit-learn gets you there with less to install and no platform restriction. TF-DF earns its weight when the forest has to live next to neural network models in one pipeline, or when you want the model format to be convertible to YDF's C++, JavaScript, CLI and Go runtimes later. Choose on the deployment target, not on the training API, because both libraries train a forest with a few lines of Python.

Maintenance, Licence and Upgrade Costs

The repository is not archived, and the last push was on 2026-05-19. The release history shows v1.12.0 on 2025-03-13, v1.10.1 on 2025-03-28 and v1.11.0 on 2024-10-28, so the version numbers in that list are not in strict chronological order, which is worth knowing before you assume v1.12.0 supersedes v1.10.1 in every respect. Check CHANGELOG.md for what each release actually changed.

Licensing is Apache-2.0, the same permissive licence used across much of the TensorFlow ecosystem. That generally means you can use, modify and redistribute the code with the usual obligations around notices and patent terms, but the LICENSE file is the authority and this is not legal advice.

The upgrade cost is dominated by TensorFlow and Keras version alignment rather than by TF-DF itself. Because the library is a Keras layer over a C++ core, a Keras major version change can reach your model code, and the README does not describe a downgrade procedure. Budget time for pinning TensorFlow alongside TF-DF in your environment file, and read documentation/known_issues.md before upgrading a production pipeline. The migration path to YDF is documented separately, and since the model formats are compatible, an exported SavedModel is not stranded if you move.

Editorial conclusion

Adopt TF-DF if you already have a Keras or TensorFlow pipeline, need a forest model that can be saved as a SavedModel alongside neural networks, or want a tree-based baseline without leaving Python. Do not adopt it for a new project if you have no TensorFlow dependency: the README itself recommends migrating to Yggdrasil Decision Forests, which trains the same models and is described as faster with more functionality. Before committing, verify three things: that your platform is Linux or macOS (Windows requires WSL plus Linux), that the pip package resolves against your installed TensorFlow version, and whether the Keras 3 API changes in v1.12.0 affect the model classes you plan to subclass. The migration guide at ydf.readthedocs.io is the single most useful page to read before you write training code, because it tells you what the successor library does differently while the model format stays compatible.

Frequently asked questions

What are the cons of decision trees?

The README does not discuss the statistical weaknesses of decision trees; it describes TF-DF as a library to train, run and interpret decision forest models and points readers to the Yggdrasil Decision Forests introduction page for background on decision forests.

Is XGBoost a decision tree?

The README does not mention XGBoost. It does state that TF-DF supports gradient boosted trees alongside random forests, which is the tree ensemble family it trains.

What do decision trees mean?

The README does not define decision trees itself; it links to the Yggdrasil Decision Forests introduction page for an explanation of decision forests, and describes TF-DF as a library for training, running and interpreting them in TensorFlow.

Can you provide an example of a decision tree?

The README's usage example is a complete end-to-end run rather than a single tree diagram: it loads CSV files with Pandas, converts them with tfdf.keras.pd_dataframe_to_tf_dataset, trains a tfdf.keras.RandomForestModel, evaluates it and saves it. The repository also includes examples/minimal.py and other files under examples/.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. README
  4. Releases
  5. tensorflow/decision-forests on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/tensorflow-decision-forests.svg)](https://hysenlabs.com/projects/tensorflow-decision-forests)