TabICL does its learning in predict, and its older checkpoints used four times the default ensemble
TabICLv2: An open tabular foundation model
At a glance
- What is it?
- An open tabular foundation model whose classifier and regressor take the same training data as any scikit-learn estimator and then do the learning at prediction time, in one forward pass through a pre-trained transformer. Three model generations ship, only the newest one handles regression, and the default ensemble size was cut from thirty-two to eight after the fact, which quietly breaks comparability with older published numbers.
- Who is it for?
- TabICL is worth benchmarking on your own tables if you currently spend time tuning gradient boosting, because the claim is specific: no hyperparameter tuning, and beating heavily tuned boosting on roughly four datasets in five on the benchmark it names. Three things to check before you rely on the headline numbers.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 9, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Fit is cheap, predict is where the learning happens
The API is scikit-learn shaped, and the comments in the basic example explain the whole design in two lines.
from tabicl import TabICLClassifier, TabICLRegressor
clf = TabICLClassifier()
clf.fit(X_train, y_train) # downloads checkpoint on first use, otherwise cheap
clf.predict(X_test) # in-context learning happens hereSo fit does not train anything. It fetches the checkpoint on first use and otherwise does almost no work. The learning occurs during predict, when the training data is supplied as context to the pre-trained transformer. Both fit and predict run in a single forward pass through that transformer, which is why there is no training loop and no hyperparameter search: the training data is the prompt.
That is also why the checkpoint is downloaded rather than trained. And it is why the same estimator class can serve both classification and regression, with separate checkpoints for each.
The accuracy claim is stated against a named benchmark and a named comparison set: the model is competitive on two tabular benchmarks, needs no hyperparameter tuning, and beats heavily tuned gradient boosting on about eighty percent of datasets in one of them. Eighty percent is not everywhere, and the honest reading is that you should check the fifth.
Three generations ship, and only the newest one does regression
The model table lists three checkpoints, and the shape of it tells you which one to use.
The current version is the default for both tasks, with separate classifier and regressor checkpoints dated to the same day. It is described as strongly improved over the first version through better synthetic pre-training data, architectural improvements, and better pre-training itself, at comparable runtime.
The two older versions have no regression checkpoint at all. The first version is the original model from the earlier conference paper, and the intermediate one is a post-training of that version on an early form of the current version's prior. Both are classification only, and the intermediate one is marked as having no associated paper.
So regression is a capability of the current generation, not of the library. If you have an older notebook pinning a version by name, you will get a classifier with no regressor available, and nothing in the error would explain why.
There is one more line in that section worth reading twice. The two older versions originally used thirty-two ensemble members, and the default was later reduced to eight. Any published number produced with the older versions therefore used four times the current default, which is a difference large enough to change both runtime and accuracy.
Seven parameters exist, and the default is not tuned by you
There is a full parameter example in the readme showing every option with its default, and reading it tells you what the model is actually doing.
The ensemble size defaults to eight members, with the comment that more is better but slower. That is the one knob a practitioner reaches for first, and it is also the one the project just reduced.
Two of the parameters are about permutation rather than tuning. One chooses a feature permutation strategy with a default named for a sampling method, and another chooses a class permutation strategy with a different default. These are the ensemble's source of diversity: rather than averaging over random seeds, the model averages over different orderings of the columns and of the classes.
Two are about data hygiene. A normalisation method defaults to trying none, which implies the alternatives are opt in, and an outlier threshold defaults to a z-score of four, described as the threshold for detection and clipping. So extreme values are clipped rather than dropped.
Two are about how the answer is produced. A temperature parameter controls prediction confidence, and a flag chooses between averaging logits and averaging probabilities, which is a real difference when one class dominates.
The regressor takes the same parameters minus the four classification-specific ones, which is a clean indication of which of these are about labels rather than about data.
The speed number is one GPU and one dataset shape
Performance claims are worth quoting precisely, because this one is anchored very tightly.
The stated figure: on one specific GPU, fit and predict together on a dataset of fifty thousand samples and one hundred features completes in under ten seconds, which the readme says is ten times faster than a competing in-context tabular model. A GPU is recommended for larger datasets.
That is one hardware SKU, one dataset shape, and one metric. It is a good signal that the approach is fast, and it is not a claim about your data. The relevant variable is almost certainly width and row count rather than anything semantic.
The scalability claim is given as a range rather than a point: strong performance on benchmarks running from three hundred to a hundred thousand training samples and up to two thousand features. Beyond that, larger datasets are supported through CPU and disk offloading, and the readme attaches a caveat in the same sentence: accuracy may degrade at some point. That is the most honest clause in the section, because it admits the scaling mechanism trades quality for size at an unspecified threshold.
The speedup for repeated inference comes from caching. Because the training data is context rather than weights, the key and value projections of the training set can be cached during fit and reused across predict calls, so later calls only process the new data. The cost is additional memory, and the readme says so rather than leaving you to discover it.
Three save switches, and one of them is about privacy
Serialisation is a short API with three boolean switches, and the combinations are meaningful.
A fitted estimator saves to a single file with three flags. One controls whether the model weights go in, and the alternative is to reload from the checkpoint on load, which keeps the file small. One controls whether the training data goes in. One controls whether the key value cache goes in, and only applies if a cache exists.
The privacy case is the third combination. When a cache exists and is saved, you can exclude the cached training data from the file, and the readme says this may be useful for data privacy.
That is worth pausing on. The cached projections are derived from your training set, and a pickled model that contains them carries information about your rows. Setting one flag to false produces an artefact that looks like a fitted model and contains no data. If you are shipping a model to someone, or storing one in a shared location, that distinction is the whole ballgame.
The cost is that a saved model without weights is not self-contained. It reloads from the checkpoint, which means the checkpoint is still a dependency at load time rather than at fit time.
Fine-tuning is a separate class, and its output loads into the zero-shot estimator
Zero-shot use is the default, and the readme is clear about when to depart from it: when one downstream dataset matters enough to spend a few minutes on.
The fine-tuning path is a full PyTorch training loop rather than a light adapter. It specifies the optimiser as AdamW with a warmup schedule, gradient clipping, early stopping against a held-out split, and multi-GPU runs. Three ensemble sizes are configurable and they are not the same number: members per training meta-batch, ensemble size for end of epoch validation, and the ensemble size of the estimator you actually call at prediction time, which defaults back to eight.
That three-way split is the detail that matters. Training, validation and inference ensembles are different sizes by design, so the number you tune for training quality is not the number that governs your serving cost.
The regressor takes the same parameters with one addition, a choice of evaluation metric between three options. Both classes point at their own docstring for the full surface rather than the readme repeating it.
Then the useful part: the checkpoint written during fine-tuning follows the pretraining checkpoint schema, so it loads straight back into the ordinary zero-shot estimator by path. A fine-tuned model is therefore interchangeable with a downloaded one at inference time, which means you can fine-tune offline and serve with the same class your pipeline already uses.
The dependency comments say out loud what each package is for
The project file carries comments next to its dependencies, and they are unusually candid about what is load bearing.
The deep learning framework has a floor version whose stated reason is a specific attention kernel rather than any API the code calls. That is a heavier coupling than a version number suggests: you are pinned to a minimum because of one accelerated kernel, not because the library needs a new release.
A tensor operation library is required, and the comment says it is used only for rotary position embeddings and could be replaced. A process utility library is required, and the comment says it is used for memory calculation at inference, which is how the CPU and disk offloading path knows what it has to fit. The rest are the ordinary scientific stack plus a progress bar and the model hub client.
The documentation extra is the interesting one. It pulls two different explainability libraries rather than one, a numeric compiler added specifically, with a comment that it is there to make the installer's life easier, and a gallery extension that executes the example notebooks as part of the documentation build.
There are five optional extras in total, covering forecasting, explainability, fine-tuning, pre-training and everything at once. And there is a platform note worth heeding: installing the deep learning framework through pip can fail on Intel Macs, in which case the readme tells you to install it through conda first. Since the release dropped support for that combination elsewhere in the ecosystem, this is the project being helpful about a gap it does not control.
Editorial conclusion
TabICL is worth benchmarking on your own tables if you currently spend time tuning gradient boosting, because the claim is specific: no hyperparameter tuning, and beating heavily tuned boosting on roughly four datasets in five on the benchmark it names. Three things to check before you rely on the headline numbers. The speed claim is anchored to one GPU and one dataset shape, so measure on your own. The accuracy figures outside that benchmark are not published in the readme, so treat them as unverified. And if you compare against an older version, check the ensemble default, because it was reduced after those versions shipped and the results are not comparable across the change.
Frequently asked questions
What is TabICL?
An open tabular foundation model for classification and regression, published as a pip-installable package that follows the scikit-learn estimator interface. It does not train during fit; instead the training data is supplied as context during predict, and both steps run in a single forward pass through a pre-trained transformer, so there is no hyperparameter tuning. The current version is the default for both tasks.
How does TabICL compare with TabPFN?
The readme gives one direct comparison: on a specific GPU, fitting and predicting a dataset of fifty thousand samples and a hundred features takes under ten seconds, which it states is ten times faster than a competing in-context tabular model. It also claims no hyperparameter tuning is needed, and that it beats heavily tuned gradient boosting on roughly eighty percent of datasets in one named benchmark.
Does TabICL support regression as well as classification?
The current version does, with a default regressor checkpoint alongside the classifier one. The two older versions listed in the readme are classification only and have no regression checkpoint at all, so requesting a regressor while pinned to an older version would fail for a reason the error would not explain.
Do I need a GPU to run TabICL?
Not to use it, but the readme recommends a GPU for larger datasets and anchors its speed claim to one specific GPU. Datasets much larger than the benchmark range can be handled through CPU and disk offloading, with the caveat stated in the same sentence: accuracy may degrade at some point. The readme also warns that installing the deep learning framework through pip can fail on Intel Macs and gives a conda alternative.
What is the key-value cache in TabICL for?
Faster repeated inference on the same training data. Because the training set is context rather than weights, its projections can be cached during fit and reused across predict calls so later calls only process new data. The cost is additional memory, which the readme names rather than hides, and because the cached projections derive from your rows you can save a model without them, which the readme notes may be useful for data privacy.
How do I fine-tune TabICL on my own dataset?
Install the fine-tuning extra, then use the dedicated fine-tuned estimator classes rather than the default ones. They run a full training loop with a warmed-up optimiser, gradient clipping, early stopping against a held-out split, and multi-GPU support, with separate ensemble sizes for training, validation and inference. The checkpoint they write follows the pretraining schema, so it loads straight back into the ordinary zero-shot estimator.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/soda-inria-tabicl)