LimiX: a structured-data foundation model you run from Python, not a training project
LimiX: Unleashing Structured-Data Modeling Capability for Generalist Intelligence https://arxiv.org/abs/2509.03505
At a glance
- What is it?
- LimiX is an Apache-2.0 tabular foundation model from limix-ldm-ai that handles classification, regression and missing-value imputation with one checkpoint. The interesting part is the inference-only workflow, and the parts the README leaves thin.
- Who is it for?
- Adopt LimiX if you have an NVIDIA GPU, a fixed CUDA and Python 3.12 environment, and a stream of tabular problems where retraining per dataset is the bottleneck. Skip it if you need CPU inference, if missing value imputation matters and you were planning to use LimiX-2M, or if you need feature-level attribution from the model itself.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 14 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What LimiX is trying to replace in a tabular pipeline
Most tabular work still starts with a per-dataset pipeline. You pick a model family, tune it, handle missing values separately, and repeat the whole exercise for the next table. LimiX is an attempt to collapse that into one checkpoint. The README describes it as "the first structured-data foundation model built for general intelligence", and the stated goal is joint distribution modeling of variables and missing values so that a single model covers classification, regression, missing value imputation and tabular generation without task-specific network design.
The intended user is not someone training a model from scratch. The repository ships inference_classifier.py, inference_regression.py, an inference/ directory, an examples/ directory with three demo scripts, and a model/ directory. There is no training script listed among the top-level entries. That layout tells you what the project expects: you download a checkpoint and call it. If your job is to fine-tune a bespoke architecture on proprietary data, this is not the entry point.
The claim that LimiX "surpasses XGBoost, classic tabular deep learning models and existing tabular foundation models on benchmarks of 10 mainstream structured datasets" comes from the project's own README and the technical report. Treat it as the authors' evaluation, not an independent one. The benchmark images in doc/ cover BCCO, TabArena and TabZilla splits for classification and regression, plus a missing-value imputation figure.
How a LimiX inference call actually moves data
The mechanism described in the README is a transformer with dual attention. Features X and targets Y are embedded from what the README calls a "prior knowledge base" into token representations, then attention is applied across both the sample dimension and the feature dimension. That second axis is the distinguishing part: most tabular deep learning treats columns as fixed positions, while LimiX attends over them to find salient features. The resulting representations go to regression and classification heads.
The practical consequence is that task type is selected by which predictor class you instantiate, not by rebuilding the network. The LimiXPredictor constructor in the README exposes a single interface with device, model_path, mix_precision, inference_config, categorical_features_indices, outlier_remove_std, softmax_temperature, mask_prediction, inference_with_DDP and seed. Several of those are opinionated defaults rather than neutral knobs. outlier_remove_std defaults to 12, softmax_temperature to 0.9, and mix_precision to True. The README documents the parameters in a table but does not explain how outlier_remove_std interacts with genuinely heavy-tailed targets, which is the kind of gap that shows up only when you run it on your own data.
inference_config accepts either a list or a string, and the config/ directory holds the configuration files. The README does not spell out the schema of those files, so reading config/ is the only way to know what a valid value looks like.
Installing LimiX and running a first classification job
The README recommends the Dockerfile over a manual build. The build command takes the base image as a build argument, which is how you pin the CUDA version:
docker build --network=host -t limix/infe:v1 --build-arg FROM_IMAGES=nvidia/cuda:12.2.0-base-ubuntu22.04 -f Dockerfile .The Dockerfile itself creates a conda environment named LimiX with python=3.12.7, installs torch==2.7.1, torchvision==0.22.1 and torchaudio==2.7.1, then installs a prebuilt flash_attn wheel and scikit-learn. Note the comment at the top of the Dockerfile: the flash_attn wheel must be downloaded first, from the Dao-AILab release URL, and placed next to the Dockerfile because the COPY instruction expects it there. The build will fail without that file.
The manual path is documented as Option 2 and is more fragile. You fetch the same wheel, then install pinned versions:
wget -O flash_attn-2.8.0.post2+cu12torch2.7cxx11abiTRUE-cp312-cp312-linux_x86_64.whl https://github.com/Dao-AILab/flash-attention/releases/download/v2.8.0.post2/flash_attn-2.8.0.post2+cu12torch2.7cxx11abiTRUE-cp312-cp312-linux_x86_64.whl
pip install python==3.12.7 torch==2.7.1 torchvision==0.22.1 torchaudio==2.7.1
pip install flash_attn-2.8.0.post2+cu12torch2.7cxx11abiTRUE-cp312-cp312-linux_x86_64.whlThen clone the source and move into it:
git clone https://github.com/limix-ldm/LimiX.git
cd LimiXThe wheel filename encodes cu12, torch2.7 and cp312. If your Python or CUDA differs, that exact file will not install, and the README gives no alternative build.
For the first real run, the model download table lists two checkpoints. LimiX-16M supports classification, regression and missing value imputation. LimiX-2M supports classification and regression only. Both are hosted on Hugging Face under stableai-org. The repository ships examples/demo_classification.py, examples/demo_regression.py and examples/demo_missing_value_imputation.py; the README does not print their contents, so open the file to see which inference_config and categorical_features_indices values it passes.
Where LimiX is the wrong tool
The dependency chain is the first hard limit. A pinned flash_attn wheel built for cu12, torch 2.7 and cp312, plus mix_precision defaulting to True, means the documented path assumes an NVIDIA GPU with a matching CUDA runtime. The README does not document a CPU-only inference path. If your deployment target is a CPU-only container or an environment where you cannot control the CUDA version, the recommended install route does not apply to you.
The second limit is task coverage versus model size. Missing value imputation is listed for LimiX-16M and not for LimiX-2M. The release notes for LimiX-2M, published on 2025-11-10, say the smaller variant offers "significantly lower GPU memory usage and faster inference speed" and that the retrieval mechanism was enhanced, but the model table is explicit that imputation is out of scope for it. Choosing the small model to save memory costs you a task.
The third is scope of the claim. The README's performance statement is about "10 mainstream structured datasets" and the figures in doc/ are benchmark splits. That is not the same as your table. Nothing in the README describes a calibration or validation step before you trust the outputs, and outlier_remove_std=12 will silently clip a heavy-tailed target. Verify on a held-out slice of your own data before wiring LimiX into anything that acts on its predictions.
LimiX against XGBoost and the tabular foundation model family
The obvious comparison is XGBoost, and the difference is not accuracy on a leaderboard. XGBoost trains a new model per dataset. You supply the features, you tune the hyperparameters, and the artifact is specific to that table. LimiX inverts this: the checkpoint is fixed and your table is the input at inference time. That is why there is no training script in the repository layout. The trade is control for convenience. With gradient boosting you can inspect split gains and feature importances directly; with LimiX you get predictions from a transformer whose internal attribution the README does not describe.
The second comparison is to other tabular foundation models, which the README groups together and claims to surpass on its benchmarks. The architectural difference it highlights is the attention over both sample and feature dimensions, plus the retrieval mechanism that LimiX-2M's release notes say was enhanced. The retrieval_extension/ directory is where that lives. If you are evaluating LimiX against another in-context tabular model, the meaningful question is not which wins on TabArena but whether the retrieval path can use your own reference set, and the README does not document that workflow.
Maintenance, licence and what an upgrade costs
The last push to the default branch was on 2026-06-16, and the repository is not archived. The release history is short and specific: V1.0.1 on 2025-09-19, V1.1.0 on 2025-11-10, with LimiX-2M announced on the same day as V1.1.0. The news section also records that LimiX-2M was accepted to ICML 2026, with a separate arXiv entry. That is a research-release cadence rather than a library cadence, and you should plan upgrades around checkpoints rather than around a semantic versioning guarantee.
Upgrade cost concentrates in two places. The dependency pins are exact: python=3.12.7, torch==2.7.1, and a flash_attn wheel whose filename encodes the CUDA and Python ABI. Moving any one of those means finding or building a matching wheel, and the README offers no fallback. The second is the inference_config format. It accepts a list or a string and lives in config/, but the README does not version or document the schema, so a config that works with one checkpoint may not be valid for the next.
The licence is Apache-2.0. The README states that "all model resources are open-sourced under the Apache 2.0 License", and LICENSE.txt sits at the repository root. Apache-2.0 includes an explicit patent grant and requires that you retain notices, which matters if you redistribute the model inside a product. This is a description of the licence text, not legal advice for your situation.
Editorial conclusion
Adopt LimiX if you have an NVIDIA GPU, a fixed CUDA and Python 3.12 environment, and a stream of tabular problems where retraining per dataset is the bottleneck. Skip it if you need CPU inference, if missing value imputation matters and you were planning to use LimiX-2M, or if you need feature-level attribution from the model itself. Before wiring it into a pipeline, open config/ to learn the inference_config schema the README leaves undocumented, and run examples/demo_classification.py against a held-out slice of your own table to see how outlier_remove_std=12 treats your target distribution.
Frequently asked questions
Is Windows or Linux better for running LimiX?
The README documents only Linux paths. The recommended install builds a Docker image from an nvidia/cuda:12.2.0-base-ubuntu22.04 base, and the manual path uses wget and a CUDA-specific flash_attn wheel filename. No Windows instructions appear in the README.
Who mainly uses LimiX?
The repository is aimed at engineers and researchers who want to run tabular classification, regression or missing value imputation from a pretrained checkpoint instead of training a model per dataset. There is no training script among the top-level entries, only inference scripts and example demos.
Why would you use LimiX instead of training your own model?
The README's argument is that one model covers classification, regression, missing value imputation and tabular generation without task-specific network design, so you skip per-dataset architecture work. The cost is that the checkpoint is fixed and the README does not describe feature-level attribution.
What is the downside of using LimiX?
The documented install assumes an NVIDIA GPU with a matching CUDA runtime and a pinned flash_attn wheel, and no CPU-only path is described. LimiX-2M does not support missing value imputation, and the README does not document the inference_config schema.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/limix-ldm-ai-limix)