CatBoost: gradient boosting that handles categorical features without manual encoding
A fast, scalable, high performance Gradient Boosting on Decision Trees library, used for ranking, classification, regression and other machine learning tasks for Python, R, Java, C++. Supports computation on CPU and GPU.
At a glance
- What is it?
- CatBoost is a gradient boosting library from Yandex with native support for categorical features and CPU/GPU training. Here is what the repository and documentation actually tell you, and where the trade-offs sit.
- Who is it for?
- Adopt CatBoost if your tabular data carries high-cardinality categorical columns and you would otherwise spend days on target encoding or one-hot expansion; the library handles them in the boosting loop itself. Do not adopt it if you need a small dependency footprint, since the Python package wraps a compiled core, or if your team has already standardised on another GBDT library and your categorical handling is solved.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly C++, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What CatBoost solves that plain one-hot encoding does not
Gradient boosting over decision trees is a well-trodden method. The friction shows up on tabular data with categorical columns: a user ID, a product SKU, a city name. One-hot encoding explodes the feature space, and target encoding leaks the label unless you are careful with folds. CatBoost's stated advantage is support for both numerical and categorical features, handled inside the algorithm rather than in a preprocessing step. The README frames the main advantages as superior quality compared with other GBDT libraries on many datasets, best-in-class prediction speed, and fast GPU and multi-GPU support. The audience is the practitioner with a messy table: ranking, classification and regression tasks where some columns are strings. The repository ships bindings for Python, R, Java and C++, so the same trained model can be scored from a service written in a different language than the one used for training.
How the categorical handling and GPU path fit together
CatBoost is a gradient boosting method over decision trees, and the documentation describes the algorithm's main stages separately from the prediction stages. The categorical support is the part that distinguishes it from a generic GBDT: rather than requiring the caller to encode strings first, the library accepts categorical columns and the algorithm deals with them during training. The project's own reference paper is titled "CatBoost: gradient boosting with categorical features support", which is a fair summary of the design intent. On the compute side, training runs on CPU by default and GPU support is described as out-of-the-box, with multi-GPU variants. The build layout reflects that: the top-level CMake files include separate configurations such as CMakeLists.linux-x86_64-cuda.txt and CMakeLists.windows-x86_64-cuda.txt alongside the non-CUDA ones, so a GPU build is a distinct compilation target rather than a runtime toggle in the C++ source. For distributed work the README points at Apache Spark and a CLI mode for distributed learning. Prediction is also exposed through a standalone model API documented in catboost/CatboostModelAPI.md, which matters if you want to evaluate a trained model without the training stack.
Installing CatBoost in Python and training a first classifier
The README directs readers to the installation guide for the Python package, the R package, the command line, and the Apache Spark package. It also carries PyPI and conda-forge badges, so both a pip route and a conda route exist. The README does not reproduce the install commands inline, so the exact line to run is on the linked Python installation page rather than in the repository's front page. What the README does show is the shape of the API: the Python package is imported as catboost, and the classifier and regressor are the two entry points most tutorials start from. The library's distinguishing argument is that you pass categorical columns to the model instead of encoding them yourself, which the documentation covers under the algorithm's main stages. Once a model is trained, the README points to the model API documentation in catboost/CatboostModelAPI.md for evaluating it inside another application, and to a separate tutorials repository at github.com/catboost/tutorials for worked examples covering training modes, cross-validation, parameter tuning and feature importance. If you are working in a notebook, the install step is whatever the linked Python installation page specifies; the README does not describe a separate notebook-specific installer.
Where CatBoost is the wrong tool
The library is a compiled C++ core with language bindings, so the Python package is not a small pure-Python dependency. If your deployment target is a constrained environment where you count megabytes, that matters. The GPU path is also not free: the repository keeps separate CUDA CMake configurations, and the documentation describes GPU training as a distinct mode rather than something that silently accelerates every workload. On small datasets the overhead of moving data to the device can outweigh the benefit, and the README makes no claim about a threshold. There is a documentation note worth flagging: the README tells readers that if they cannot open the documentation in a browser they should allow yastatic.net and yastat.net domains, which implies the docs site loads assets from those hosts. That is an unusual dependency for a documentation page and can bite in restricted corporate networks. Finally, CatBoost is a supervised tabular method. It is not a replacement for a neural approach on text or images, and the README does not present it as one.
CatBoost against XGBoost and LightGBM
The related searches around this project are dominated by comparisons, so it is worth being precise about the difference in approach rather than declaring a winner. XGBoost and LightGBM are also gradient boosting over decision trees. The distinction CatBoost states for itself is categorical feature support as a first-class part of the algorithm, plus GPU and multi-GPU training described as out-of-the-box. XGBoost historically expects you to encode categorical variables before training, and LightGBM has its own categorical handling with a different implementation. The README links a benchmarks repository at github.com/catboost/benchmarks for quality comparisons against other GBDT libraries, and that is the honest place to look rather than a summary sentence. If your data is all numeric and your pipeline is already tuned, switching libraries is mostly a rewrite of the training script with uncertain payoff. If your data is mostly categorical strings, the difference in preprocessing effort is the real argument.
Maintenance, releases and licence
The repository is not archived, and the last push was on 2026-09-10. Recent releases listed are node-package-v1.27.0 on 2026-02-21, v1.2.10 on 2026-02-19 and v1.2.9 on 2026-02-18. Note the two version lines: a Node package versioned 1.27.0 and a core version at 1.2.x. That is a real upgrade consideration, because the Node binding and the core library move on separate schedules, and pinning one does not pin the other. For Python and R users the version to track is the 1.2.x line. The project is licensed under Apache License, Version 2.0, with copyright attributed to YANDEX LLC, 2017-2026, and the README points to the LICENSE file for details. Apache-2.0 is a permissive licence that includes an explicit patent grant, which is generally friendlier for commercial embedding than a copyleft licence, but the exact obligations depend on how you redistribute and this is not legal advice. The repository carries a SECURITY.md and a CONTRIBUTING.md, and the README lists a Telegram group, GitHub Discussions and a Stack Overflow tag as support channels.
Editorial conclusion
Adopt CatBoost if your tabular data carries high-cardinality categorical columns and you would otherwise spend days on target encoding or one-hot expansion; the library handles them in the boosting loop itself. Do not adopt it if you need a small dependency footprint, since the Python package wraps a compiled core, or if your team has already standardised on another GBDT library and your categorical handling is solved. Before committing, verify the package installs cleanly on your Python version, check that the GPU build configuration matches your CUDA setup, and confirm the model export path covers the prediction runtime you actually deploy to.
Frequently asked questions
Is CatBoost better than XGBoost?
The README states that CatBoost has superior quality compared with other GBDT libraries on many datasets and links a separate benchmarks repository for the comparisons. The honest answer is that it depends on the dataset, and the project's own benchmark repository is the place to check rather than a summary claim.
What is CatBoost used for?
The README describes it as a gradient boosting method over decision trees used for ranking, classification, regression and other machine learning tasks. The repository provides bindings for Python, R, Java and C++, plus Apache Spark and a command line interface for distributed training.
How can I use CatBoost in Python?
Install the catboost package following the Python installation page linked from the README, then import CatBoostClassifier or CatBoostRegressor and call fit. Categorical columns are passed through the cat_features argument instead of being encoded beforehand, which is the library's stated advantage over generic GBDT tooling.
Who invented CatBoost?
The licence attributes copyright to YANDEX LLC, 2017-2026, and the README lists the reference paper authors as Anna Veronika Dorogush, Andrey Gulin, Gleb Gusev, Nikita Kazeev, Liudmila Ostroumova Prokhorenkova and Aleksandr Vorobev.
How do I use CatBoost on a GPU?
The README describes fast GPU and multi-GPU support for out-of-the-box training, and the repository keeps separate CUDA build configurations such as CMakeLists.linux-x86_64-cuda.txt. The README does not state a dataset size threshold at which GPU training becomes worthwhile.
How do I install CatBoost?
The README points to separate installation guides for the Python package, the R package, the command line and the Apache Spark package, and carries PyPI and conda-forge badges. The exact commands are on those linked installation pages rather than in the README itself.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/catboost-catboost)