Library / SDK
dmlc/xgboost avatar
dmlc/xgboost

XGBoost: Distributed Gradient Boosting Across Python, R, and JVM

Scalable, Portable and Distributed Gradient Boosting (GBDT, GBRT or GBM) Library, for Python, R, Java, Scala, C++ and more. Runs on single machine, Hadoop, Spark, Dask, Flink and DataFlow

28,796 stars8,912 forksC++Apache-2.0

At a glance

What is it?
XGBoost implements gradient-boosted decision trees in a C++ core with language wrappers for Python, R, Java, and Scala. The same code runs on a single machine or scales to distributed clusters including Spark, Dask, and Kubernetes.
Who is it for?
Engineers working on structured tabular data in Python, R, or JVM environments who need a distributed-capable gradient boosting implementation will find the Apache-2.0 license and multi-language package ecosystem directly useful. Teams working on unstructured input types such as raw images or audio sequences, or needing deep learning architectures, should look elsewhere.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 4 days ago.
What is it written in?
Mainly C++, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 26, 2026, and from our analysis. They are not legal advice.

Editorial analysis

A Gradient Boosting Library Built Around a C++ Core

XGBoost, short for eXtreme Gradient Boosting, is a machine learning library designed for engineers who need a gradient-boosted decision tree implementation that works across multiple languages without sacrificing performance. The C++ core implements the gradient boosting framework and exposes it through wrappers for Python, R, Java, and Scala. The Python package is the most common entry point, but the same algorithms are available to teams working in JVM-based languages through the jvm-packages/ directory in the repository.

The library originated from a research project at the University of Washington and was introduced in the paper "XGBoost: A Scalable Tree Boosting System" by Tianqi Chen and Carlos Guestrin, presented at the SIGKDD 2016 Conference on Knowledge Discovery and Data Mining. That paper is the canonical reference for the algorithm's design. The repository also includes a CITATION file pointing to that work. The homepage is at xgboost.readthedocs.io, where the full documentation lives.

Parallel Tree Boosting: How the Algorithm Works

Gradient boosting builds an ensemble by adding trees one at a time, each correcting residual errors from the previous ensemble. XGBoost calls its implementation parallel tree boosting. The parallelism applies to how individual trees are constructed: the algorithm can use multiple cores when searching for the best split points within a single tree, not only when building trees sequentially.

The README describes the library as implementing machine learning algorithms under the Gradient Boosting framework and providing what it calls GBDT (Gradient Boosted Decision Trees), GBRT (Gradient Boosted Regression Trees), and GBM (Gradient Boosting Machine) variants. All three names refer to the same underlying technique. The README states the implementation can solve problems beyond billions of examples, meaning the design avoids memory bottlenecks when the training set is very large. The repository description characterizes the library as scalable, portable, and distributed.

Installing XGBoost for Python, R, and JVM

The Python package is available on PyPI at pypi.python.org/pypi/xgboost and on conda-forge as py-xgboost at anaconda.org/conda-forge/py-xgboost. The R package is listed on CRAN at cran.r-project.org/web/packages/xgboost. The README links to all three registries directly. The full installation guide is at xgboost.readthedocs.io, which covers environment-specific steps for each language.

For JVM users, the repository contains a jvm-packages/ directory that provides the Java and Scala interface to the same C++ core. The repository structure separates language wrappers clearly: R-package/ holds the R code, python-package/ holds the Python interface, and jvm-packages/ covers the JVM side. The demo/ directory organizes worked examples by task, including demo/guide-python/ for Python walkthroughs, demo/dask/ for distributed execution examples, and demo/c-api/ for users working directly with the C interface. The documentation at xgboost.readthedocs.io is the authoritative starting point for setup details beyond finding the right package registry.

Scaling to Spark, Dask, Kubernetes, and Other Clusters

The README states that the same code running on a single machine also runs on Kubernetes, Hadoop, SGE, Dask, Spark, PySpark, and Google Cloud DataFlow without modification. This cross-environment portability is a central design goal. The stated capability to handle problems beyond billions of examples requires moving past single-machine memory constraints, and the distributed execution support is the mechanism for that.

The demo/dask/ subdirectory in the repository contains examples specific to Dask-based execution. The repository description also explicitly names Flink as a supported environment alongside the others. This matters for teams whose training data lives in a managed cluster environment: the gradient boosting code can run in a local Python environment during development, then shift to a cluster for production training without rewriting core algorithm calls. The distributed support also means that the same model built during development can be used for inference across different deployment targets.

Where XGBoost Is the Wrong Tool

XGBoost is designed around tree-based gradient boosting. The README does not document support for architectures outside this framework. Teams working on raw image data, audio sequences, or other inputs that benefit from convolutional or recurrent architectures will need a separate toolset. The library requires training data to be expressed as feature vectors, which is inherent to tree-based methods.

The README also does not document rollback procedures for model versioning, an integrated feature store, or deployment packaging beyond what the language ecosystems provide. Teams needing those capabilities will need additional infrastructure. The documentation at xgboost.readthedocs.io likely covers the supported use cases in more detail, but the README itself draws no explicit boundaries on input types. Users who discover they need a neural network architecture after starting with XGBoost will face a migration rather than a configuration change.

XGBoost vs LightGBM: Level-Wise vs Leaf-Wise Tree Growth

LightGBM is Microsoft's gradient boosting framework and the most direct alternative to XGBoost. The RELATED SEARCHES data shows people frequently compare the two. The key architectural difference is in how each library builds individual trees. LightGBM uses a leaf-wise growth strategy, expanding the leaf with the largest loss reduction at each step. XGBoost uses level-wise growth, expanding all leaves at a given depth before moving deeper. These are stable, fundamental design choices that produce different behavior on various dataset shapes.

Both libraries provide Python packages and can run in distributed environments. XGBoost's origin in a SIGKDD 2016 paper gives it a documented algorithmic foundation that has been cited and examined in the research literature. LightGBM was released by Microsoft later and also has documented research backing. The choice between them often comes down to existing tooling, language requirements beyond Python, and the specific data shapes in a project. XGBoost provides JVM wrappers that LightGBM does not offer at the same level of official support.

Releases, License, and What to Verify Before Adopting

XGBoost is released under the Apache-2.0 license, which permits commercial use and modification without requiring source disclosure. The most recent release is v3.4.2, published on 2026-09-15. Release v3.3.1 was also published on the same date, suggesting parallel maintenance of multiple lines. The repository received its most recent push on 2026-09-26 and is not archived. Release notes are at xgboost.readthedocs.io/en/latest/changes/index.html. The project has a SECURITY.md for reporting vulnerabilities.

The CITATION file in the repository provides the BibTeX entries for the SIGKDD 2016 paper, which is useful for teams using XGBoost in academic or audited settings where model provenance matters. For teams adopting it in production, the key items to verify are the Python version constraints in the current PyPI package, any breaking changes in the v3.x series documented in the release notes, and whether the distributed environment they plan to use (Spark, Dask, or Kubernetes) matches the configuration examples in the demo/ directory.

Editorial conclusion

Engineers working on structured tabular data in Python, R, or JVM environments who need a distributed-capable gradient boosting implementation will find the Apache-2.0 license and multi-language package ecosystem directly useful. Teams working on unstructured input types such as raw images or audio sequences, or needing deep learning architectures, should look elsewhere. Before adopting, verify that the Python version requirements in the PyPI package match your environment and review the v3.x release notes at xgboost.readthedocs.io/en/latest/changes/index.html for breaking changes since your target version.

Frequently asked questions

What is XGBoost used for?

XGBoost is used for supervised machine learning on structured and tabular data using gradient-boosted decision tree ensembles. The README describes it as solving many data science problems in a fast and accurate way, and it was presented at the SIGKDD 2016 data mining conference.

Is XGBoost AI or ML?

XGBoost is a machine learning library. The README describes it as implementing machine learning algorithms under the Gradient Boosting framework, distinct from rule-based systems but also distinct from deep neural network approaches.

How do you use XGBoost in Python?

The Python package is available on PyPI at pypi.python.org/pypi/xgboost and on conda-forge as py-xgboost. The documentation at xgboost.readthedocs.io covers Python-specific usage in detail, and the demo/guide-python/ directory in the repository contains worked examples.

How do you install XGBoost?

Python users can install XGBoost from PyPI or from conda-forge as py-xgboost. R users install from CRAN. For Java and Scala, the repository's jvm-packages/ directory provides the interface. The full installation guide is at xgboost.readthedocs.io.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
  4. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/dmlc-xgboost.svg)](https://hysenlabs.com/projects/dmlc-xgboost)