Library / SDK
serengil/chefboost avatar
serengil/chefboost

ChefBoost: decision trees with categorical features in a few lines of Python

A Lightweight Decision Tree Framework supporting regular algorithms: ID3, C4.5, CART, CHAID and Regression Trees; some advanced techniques: Gradient Boosting, Random Forest and Adaboost w/categorical features support for Python

488 stars101 forksPythonMIT

At a glance

What is it?
ChefBoost wraps ID3, C4.5, CART, CHAID, regression trees, gradient boosting, random forest and adaboost behind one fit call, and it emits the trained tree as a Python if-statement file. The appeal is the categorical handling; the cost is that the output is code, not a standard model object.
Who is it for?
ChefBoost fits small tabular datasets with mixed numeric and nominal columns, teaching, and any workflow where you want to read the learned rules or ship them as plain Python. It is the wrong tool for large data, sparse or high-dimensional inputs, and anything needing GPU training or a serialization format other tools can load.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 140 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem ChefBoost targets: nominal columns that other libraries make you encode

Most Python tree libraries expect a numeric matrix. If a column holds 'Sunny', 'Rain' and 'Overcast', you one-hot encode it or map it to integers, and the integer mapping invents an order that does not exist. ChefBoost takes the opposite position. The README states that it handles both numeric and nominal features and target values, so no pre-processing step is required before building a tree. That single design choice explains most of the rest of the library: the split search has to work on category sets, the generated model has to compare strings, and the saved artifact is a Python file rather than a numeric tree structure.

The intended user is someone with a small tabular dataset and a categorical column they do not want to encode. The bundled example is the classic golf dataset loaded from dataset/golf.txt with a Decision target column. That is a teaching-sized problem, and the library is honest about that scale. It is not positioned as a replacement for a production gradient boosting stack.

How ChefBoost picks splits: five metrics behind one config key

The mechanism is a recursive split search. According to the README, regular decision tree algorithms find the best feature and the best split point maximizing the information gain, then build trees recursively in child nodes. Which metric counts as "best" depends on the algorithm you select. ID3 uses entropy and information gain, C4.5 uses entropy and gain ratio, CART uses GINI, CHAID uses chi square, and the regression tree uses standard deviation. All five are reachable through the same algorithm key in the config dictionary, so switching from ID3 to CHAID is a one-word change rather than a different API.

The ensemble modes reuse that same base learner. Gradient boosting is described as building a tree and then building another based on the previous one's error, with predictions being the sum of each tree's result. Random forest splits the dataset into several sub datasets, builds a tree per subset, and averages the predictions. Adaboost is listed alongside them. The configuration surface for these is small: enableGBM with epochs, learning_rate and max_depth, or enableRandomForest with num_of_trees.

One consequence of the recursive information-gain search is worth naming. Greedy top-down splitting is exactly what random forest inherits, so the usual caveats about axis-aligned splits and unstable trees apply here as they do anywhere else. ChefBoost does not add a pruning stage that the README describes, and it does not document early stopping for the boosting path beyond the epochs setting.

Installing ChefBoost and building a first tree from the golf dataset

The README gives PyPI as the installation route. The package declares python_requires >=3.6 in setup.py and pulls pandas, numpy, tqdm and psutil from requirements.txt.

bash
pip install chefboost

Import the module under its aliased name, which is how every example in the README refers to it.

python
from chefboost import Chefboost as chef

The minimal training call takes a pandas DataFrame, a config dictionary and the name of the target column. The README's example reads dataset/golf.txt and selects C4.5.

python
import pandas as pd

df = pd.read_csv("dataset/golf.txt")
config = {'algorithm': 'C4.5'}
model = chef.fit(df, config = config, target_label = 'Decision')

After fit returns, the trained tree is written as Python if statements under outputs/rules. The README shows the shape of that file: a findDecision function whose parameters are the feature names and whose branches compare string values such as Outlook == 'Rain'. Prediction on a new row goes through chef.predict with the values in the same order as the function signature.

python
prediction = chef.predict(model, param = ['Sunny', 'Hot', 'High', 'Weak'])

To skip training later, save the model and reload it. The README notes that restoration requires the .py and .pkl files to sit under outputs/rules.

python
chef.save_model(model, "model.pkl")
model = chef.load_model("model.pkl")
prediction = chef.predict(model, ['Sunny',85,85,'Weak'])

The generated rules file is the real interface, and that cuts both ways

ChefBoost does not hand back an opaque tree object you inspect through a library API. It writes executable Python. The README shows restoreTree loading outputs/rules/rules.py and returning an object with a findDecision method, so you can call the tree without the ChefBoost training path at all. The same property is what makes transfer learning possible in the README's framing: you restore a built tree and continue from it.

Read that as a deployment decision, not just a convenience. A rules.py file can be reviewed in a pull request, diffed between versions, and imported by a service that has no ChefBoost dependency. It also means every retrain produces a new source file, and any hand edit to that file is lost on the next fit. The README does not document a merge strategy or a stable ordering for the generated branches, so treating outputs/rules as generated code you never touch is the safer reading. The predict path also expects positional values, as the README's ['Sunny',85,85,'Weak'] example shows, which makes column order part of your model's contract rather than something the artifact records for you.

Where ChefBoost is the wrong choice

The README describes ChefBoost as lightweight and demonstrates it on a 14-row golf dataset. Nothing in the repository suggests it is built for wide data. The split search runs over category sets and string comparisons, and the output is a nested if statement per path, so a deep tree on a high-cardinality column produces a large Python file rather than a compact numeric array. If your features are thousands of sparse indicator columns, one-hot encoding plus a library that indexes columns by integer is the more natural fit, and ChefBoost's categorical advantage buys you nothing.

Scale is the second boundary. There is no documented GPU path, no distributed training, and no mention of out-of-core data. The requirements list is four small packages. A dataset that fits in memory and a model you want to read is the target; a dataset that does not fit is not.

The third boundary is interoperability. Because the artifact is a Python module plus a pickle, a service written in Go, Java or C++ cannot consume it without reimplementing the rules or running a Python sidecar. Model registries and serving frameworks that expect a standardized serialization format will need a conversion step the README does not describe.

ChefBoost against scikit-learn's tree module

The obvious comparison is scikit-learn, which also ships decision trees, random forests and gradient boosting in Python. The difference is in the input contract. scikit-learn's tree estimators require numeric input, so categorical columns go through OneHotEncoder or OrdinalEncoder first, and the resulting split conditions are thresholds on the encoded columns. ChefBoost takes the DataFrame as-is and produces equality tests on the original category strings, which is why the README can claim no pre-processing is needed.

That difference propagates. A scikit-learn tree exposes arrays and a predict method that any Python process can load; ChefBoost exposes a generated findDecision function. scikit-learn's ecosystem covers pipelines, cross-validation helpers and model selection utilities that ChefBoost does not attempt. ChefBoost covers CHAID, which scikit-learn does not provide, and it makes the learned rules directly readable as source. If your categorical columns are the hard part of the problem and the dataset is small, ChefBoost removes a preprocessing stage. If your problem is numeric, large, or needs the surrounding tooling, the encoding step is cheaper than the constraints.

Licence, maintenance and what upgrading costs

ChefBoost is MIT licensed, as stated in the README badge and in the classifiers block of setup.py. MIT is permissive: you can use, modify and redistribute it, including in closed-source products, provided the copyright notice and licence text are retained. That is a description of the licence terms as the repository states them, not legal advice; if the licence matters to your organization, read the LICENSE file and get your own counsel.

The repository is not archived, and the last push was on 2026-05-13. That is roughly four months before the date of this article, so the project has seen recent activity, though the material retrieved lists no releases, and setup.py still carries version 0.0.19. Do not read a version number that low as a maturity signal in either direction; read it as a project that has not cut a 1.0. Because it is distributed through PyPI, upgrading means pinning a version and reinstalling, and the main upgrade risk is the generated rules format. If a new version changes how outputs/rules files are written, any saved .pkl paired with an older rules.py may not restore cleanly. The README does not document a compatibility guarantee across versions, so pinning chefboost in your requirements alongside the model artifacts is the practical safeguard.

Editorial conclusion

ChefBoost fits small tabular datasets with mixed numeric and nominal columns, teaching, and any workflow where you want to read the learned rules or ship them as plain Python. It is the wrong tool for large data, sparse or high-dimensional inputs, and anything needing GPU training or a serialization format other tools can load. Before adopting it, check that the algorithm you need is actually listed in the README table, run the golf dataset through your own config to see the generated outputs/rules/rules.py, and confirm that your prediction path can carry the same column order you trained on.

Frequently asked questions

What is ChefBoost in Python?

It is a lightweight decision tree framework that supports ID3, C4.5, CART, CHAID and regression trees, plus gradient boosting, random forest and adaboost. Its distinguishing feature is that it accepts numeric and nominal features and target values without pre-processing.

How do I install ChefBoost?

The README gives pip install chefboost as the installation route, which pulls in the library and its prerequisites. You then import it with from chefboost import Chefboost as chef.

Does ChefBoost need one-hot encoding for categorical columns?

No. The README states that ChefBoost handles both numeric and nominal features and target values, so no pre-processing is required before building trees. The generated rules compare the original category values directly, for example Outlook == 'Rain'.

How do I reuse a ChefBoost model without retraining?

Use chef.save_model to write a .pkl file, then chef.load_model to restore it and chef.predict for new instances. The README notes that restoration requires the .py and .pkl files to be stored under outputs/rules.

Official sources

  1. Issues
  2. License: MIT
  3. Project website
  4. README
  5. serengil/chefboost on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/serengil-chefboost.svg)](https://hysenlabs.com/projects/serengil-chefboost)