GRF: Generalized Random Forests for Causal Inference in R
Generalized Random Forests
At a glance
- What is it?
- grf-labs/grf is a GPL-3.0 R package with a C++ core that implements non-parametric causal inference and treatment effect estimation. It provides causal forests, quantile forests, survival forests, and least-squares regression forests, all with honest estimation and formal confidence intervals.
- Who is it for?
- GRF is the right tool for statisticians and empirical researchers who need to estimate heterogeneous treatment effects with confidence intervals, handle instrumental variables, or run non-parametric survival analysis in R. Teams doing standard supervised learning for prediction should use ranger or xgboost instead, since GRF's splitting criterion is not optimized for prediction accuracy.
- Can I use it commercially?
- Yes, with conditions. GPL-3.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
- Is it still maintained?
- Yes. The repository last received commits 154 days ago.
- What is it written in?
- Mainly C++, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What GRF Solves and Who Uses It
Standard random forests are prediction tools: they estimate a conditional mean or class probability, but they do not answer causal questions or produce confidence intervals with formal coverage guarantees. GRF is designed for researchers who need to estimate how a treatment effect varies across individuals or subgroups, not just predict an outcome.
The README describes GRF as providing non-parametric methods for heterogeneous treatment effects estimation with support for right-censored outcomes, multiple treatment arms or outcomes, and instrumental variables. On top of that it provides least-squares regression, quantile regression, and survival regression, all with support for missing covariates.
The primary users are researchers in economics, public health, and social sciences who run randomized experiments or observational studies and need to characterize treatment effect heterogeneity. The package is also used in policy evaluation, where the question is not just whether a treatment works on average but whether it works better for some subgroups than others.
GRF is an R package. It has a C++ core for performance and uses the ranger fast random forest library as its foundation. The README acknowledges the ranger authors as contributors: the GRF repository started as a fork of ranger.
Honest Estimation and Confidence Intervals
The central methodological contribution that distinguishes GRF from standard random forests is the combination of honest estimation and a splitting criterion designed for causal inference.
Honest estimation means that the data used to choose splits in a tree is kept separate from the data used to populate the leaves. The README defines this as using one subset of the data for choosing splits and another for populating the leaves of the tree. This separation is what allows GRF to produce confidence intervals with valid coverage properties: without honesty, using the same observations to both build and estimate the model creates a form of in-sample bias.
The splitting criterion in GRF is not minimizing prediction error. The algorithm, described in detail in the GRF reference at grf-labs.github.io/grf/REFERENCE.html, splits on the variable and threshold that best separates local treatment effect estimates. This means the tree structure is directly optimized for heterogeneity rather than for out-of-bag prediction accuracy.
Confidence intervals are available for least-squares regression forests and for treatment effect estimation. The theoretical guarantees come from the paper published in the Annals of Statistics (2019) by Athey, Tibshirani, and Wager. The GRF reference documents where these guarantees apply and where they do not.
Available Forest Types and Their Statistical Use Cases
The README describes the following forest types available in GRF:
- Causal forests for heterogeneous treatment effects under unconfoundedness - Instrumental variable forests for when treatment assignment is endogenous - Survival forests with support for right-censored outcomes, described in a 2023 JRSS-B paper by Cui, Kosorok, Sverdrup, Wager, and Zhu - Quantile regression forests for estimating conditional quantiles - Least-squares regression forests with confidence intervals - Local linear forests, described in a paper by Friedberg, Tibshirani, Athey, and Wager in the Journal of Computational and Graphical Statistics (2020)
The causal forest is the most commonly cited forest type. It takes a covariate matrix X, an outcome vector Y, and a treatment indicator W. The assumption is that treatment assignment W is conditionally independent of potential outcomes given X, which is the standard unconfounded observational study assumption or, automatically satisfied in a randomized experiment.
The instrumental variable forest is for settings where unconfoundedness does not hold: a binary instrument Z is provided alongside X, Y, and W, and the forest estimates the local average treatment effect for compliers.
All forest types handle missing covariates natively, which the README lists as a supported feature.
Installing and Running a Causal Forest
GRF is available on CRAN:
install.packages("grf")Conda users can install from conda-forge:
conda install -c conda-forge r-grfThe development version can be installed from source with devtools:
devtools::install_github("grf-labs/grf", subdir = "r-package/grf")Installing from source requires a C++17 compiler. On Windows, the RTools toolchain is also required.
The README provides a complete working example for causal forest estimation:
library(grf)
n <- 2000
p <- 10
X <- matrix(rnorm(n * p), n, p)
X.test <- matrix(0, 101, p)
X.test[, 1] <- seq(-2, 2, length.out = 101)
W <- rbinom(n, 1, 0.4 + 0.2 * (X[, 1] > 0))
Y <- pmax(X[, 1], 0) * W + X[, 2] + pmin(X[, 3], 0) + rnorm(n)
tau.forest <- causal_forest(X, Y, W)
tau.hat.oob <- predict(tau.forest)
hist(tau.hat.oob$predictions)This example generates a dataset with heterogeneous treatment effects driven by the first covariate, trains a causal forest, and plots the distribution of out-of-bag treatment effect estimates. Out-of-bag prediction means each observation's estimate uses only trees that were not trained on that observation, providing a valid held-out estimate without a separate test set.
Limitations and Wrong-Fit Cases
GRF requires that the researcher specify the correct causal structure before choosing a forest type. Passing an observational study to causal_forest when unconfoundedness does not hold gives numerically valid output but causally invalid estimates. The README does not perform this check automatically; the choice of forest type encodes the researcher's assumptions.
The confidence intervals require the honest estimation setting and a reasonably large sample size. The GRF reference documents the theoretical conditions, but the practical implication is that confidence intervals on small datasets should be interpreted cautiously. The README points to the algorithm reference for troubleshooting, not to a list of minimum sample sizes.
GRF is an R package. It has no Python interface in this repository. Python users who need causal forest methods must look elsewhere. The README mentions the conda installation of the R package but not a Python wrapper.
The package is not designed for prediction accuracy competitions. If the goal is predictive performance rather than causal inference or heterogeneity estimation, standard random forests (ranger, randomForest) or gradient-boosted trees will typically outperform GRF because their splitting criterion is optimized for prediction, not for separating local treatment effect estimates.
Comparison with ranger, Maintenance, and License
ranger is the fast C++ random forest implementation that GRF forked from. ranger supports regression, classification, and survival random forests with speed as its primary goal. It does not implement honest estimation, causal splitting criteria, or treatment effect confidence intervals. Researchers who need only standard prediction tasks and want the fastest possible forest implementation should use ranger. Researchers who need causal inference, formal confidence intervals, or heterogeneous treatment effect estimates should use GRF.
The README acknowledges ranger directly and credits its authors for providing a useful foundation. GRF adds the honest splitting mechanism, the local parameter estimation objective, and the supporting statistical theory on top of ranger's C++ infrastructure.
The development of GRF is supported by the National Institutes of Health, the National Science Foundation, the Sloan Foundation, the Office of Naval Research, and Schmidt Futures, according to the README. The references section lists publications through 2025, and the last push to the master branch was on 2026-04-30. The repository is not archived.
GRF is released under the GPL-3.0 license. The GPL requires that any derivative work that is distributed publicly also be released under GPL-3.0 or a compatible license. For academic research and internal use, this poses no restriction. For incorporating GRF into a commercial product that is distributed to others, the GPL terms require distributing the source of the combined work.
Editorial conclusion
GRF is the right tool for statisticians and empirical researchers who need to estimate heterogeneous treatment effects with confidence intervals, handle instrumental variables, or run non-parametric survival analysis in R. Teams doing standard supervised learning for prediction should use ranger or xgboost instead, since GRF's splitting criterion is not optimized for prediction accuracy. Before applying causal_forest to an observational dataset, verify that the unconfoundedness assumption holds for your study design, since GRF does not test this assumption automatically. The GPL-3.0 license is permissive for academic use but constrains redistribution in commercial products.
Frequently asked questions
What is a generalized random forest?
A generalized random forest is a non-parametric statistical method that uses a forest of trees to estimate a local parameter rather than a prediction. GRF builds trees by splitting on variables that best separate local parameter estimates, and uses honest estimation to produce valid confidence intervals.
What are GRF's main functions?
GRF provides causal forests for heterogeneous treatment effects, instrumental variable forests, survival forests for right-censored outcomes, quantile regression forests, least-squares regression forests with confidence intervals, and local linear forests. All support missing covariates.
What are the benefits of using GRF?
GRF provides honest estimation that separates the data used for splitting from the data used for leaf estimation, which enables formal confidence intervals with valid coverage. It supports causal inference tasks that standard random forests do not address, including treatment effect heterogeneity and instrumental variable estimation.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/grf-labs-grf)