plexe: building tabular ML models from a natural language intent
✨ Build a machine learning model from a prompt
At a glance
- What is it?
- Plexe is a Python framework from plexe-ai that turns a dataset plus a plain-language intent into a packaged model. It is aimed at tabular work, it runs on LLM agents, and the repository is still labelled alpha.
- Who is it for?
- Adopt plexe if your data is tabular, your intent fits in one sentence, and you want a packaged artifact rather than a notebook. Do not adopt it if you need a documented rollback path, if you cannot send data or schema details to an LLM provider, or if your task is not tabular classification or regression.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Activity is slowing. The repository last received commits 6 months ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem plexe targets: tabular models without the boilerplate
Most tabular machine learning work is not the modelling. It is picking a metric, writing the feature pipeline, trying a handful of gradient boosting configurations, and then rebuilding all of it as a deployable artifact. Plexe's README frames the project as a way to skip that: you supply a dataset in Parquet, CSV, ORC or Avro plus an intent such as "predict whether a passenger was transported", and the system runs an agentic loop that returns a best solution, metrics and a report.
The audience is narrow but real. It is engineers who already know what a train/test split is and do not want a no-code tool, but who also do not want to hand-write the same XGBoost pipeline for the fourth time. The README's own examples are churn, house prices and passenger transport, all tabular prediction problems. Nothing in the README suggests plexe is intended for sequence models, recommendation systems or unstructured text.
The project is Apache-2.0 licensed and published on PyPI. Its pyproject.toml still carries the classifier "Development Status :: 3 - Alpha", which is worth reading literally: the interface described in the README is the interface of a project that has not declared itself stable.
Fourteen agents, six phases, and a model package at the end
The README describes a multi-agent architecture: 14 specialised AI agents across a 6-phase workflow. The phases it names are data analysis and task identification, metric selection, hypothesis-driven model search, performance and robustness evaluation, and packaging for deployment. The agents are built on smolagents, and LLM calls go through LiteLLM, which is why the config file can route different agents to different providers.
The iteration loop is the core of it. You pass max_iterations, and each iteration is a search step: the system proposes a model configuration, trains it, evaluates it, and uses the result to inform the next hypothesis. allowed_model_types constrains the search space to something like ["xgboost"] if you already know what you want. enable_final_evaluation runs the chosen model against a held-out test set.
The output is the part that matters for deployment. Plexe writes a self-contained package to work_dir/model/, also archived as model.tar.gz. The README states the package has no dependency on plexe. Its layout is artifacts/ for the trained model and feature pipeline pickles, src/ for the inference predictor and pipeline code, schemas/ for input and output JSON schemas, config/ for hyperparameters, evaluation/ for metrics and analysis reports, plus model.yaml and a README with example code. That separation is the design decision that makes the tool usable in production: you build with plexe and ship without it.
Installing plexe and running a first model
The core install pulls XGBoost, Keras and scikit-learn. CatBoost, LightGBM, PyTorch, PySpark and AWS support are extras, grouped either by framework or by task, so a tabular run with local Spark looks like this:
pip install "plexe[tabular,pyspark]"Python must be >= 3.10 and < 3.13. Plexe reaches LLM providers through LiteLLM, and the README says the team actively tests only openai/* and anthropic/* models, so export the keys for those two before anything else:
export OPENAI_API_KEY=<your-key>
export ANTHROPIC_API_KEY=<your-key>The CLI entry point takes a dataset URI and an intent. Running the README's example should produce a search over the configured number of iterations and leave a package in the working directory:
python -m plexe.main \
--train-dataset-uri data.parquet \
--intent "predict whether a passenger was transported" \
--max-iterations 5The same run from Python returns three objects. The first is the best solution, whose performance attribute is a float:
from plexe.main import main
from pathlib import Path
best_solution, metrics, report = main(
intent="predict whether a passenger was transported",
data_refs=["train.parquet"],
max_iterations=5,
work_dir=Path("./workdir"),
)
print(f"Performance: {best_solution.performance:.4f}")After a run, the Streamlit dashboard reads the working directory and shows the search tree and evaluation reports:
python -m plexe.viz --work-dir ./workdirFor a containerised path, the repository ships a Makefile with make build, make test-quick (described as roughly one iteration) and make run-titanic, and a multi-stage Dockerfile with pyspark as the default target and databricks as an alternative target built with --target databricks.
Where plexe gets awkward: cost, providers and the missing rollback story
Every iteration is a set of LLM calls. The hypothesiser, feature processor and model definer are each routed to a model, and the README's own config example points them at gpt-5-mini and claude-sonnet-4-5. Raising max_iterations therefore raises both compute and token spend, and the README does not give any guidance on what a reasonable ceiling is. The 5 and 10 values in the examples are the only numbers offered.
Provider support is broader in theory than in practice. The README is explicit: plexe should work with most LiteLLM providers, but only openai/* and anthropic/* are actively tested, and issues with other providers are something the team asks users to report. An ollama/llama3 entry appears in the config example, so local models are contemplated, but the same note applies to them.
The packaging story has a gap. The README documents what goes into work_dir/model/ and states the package does not depend on plexe, but it does not document rollback, versioning of generated packages, or how to replace a deployed model with a newly generated one. There is also no rollback or promotion mechanism described for the search itself: if iteration 7 is worse than iteration 3, the README does not explain how you would recover iteration 3's artifact.
Finally, the framework is not a substitute for domain judgement. It selects an evaluation metric on your behalf. If your problem has an asymmetric cost structure, a rare positive class, or a leakage risk in the feature columns, nothing in the described workflow guarantees the chosen metric reflects that.
How plexe differs from AutoML libraries and hand-written pipelines
The closest comparison is a classical AutoML library such as auto-sklearn or FLAML. Those search a predefined space of preprocessing steps and estimators with Bayesian optimisation or bandit methods. The search is deterministic given a seed, the objective is a scalar the caller supplies, and the whole run is reproducible without network access. Plexe replaces the fixed search space with an LLM proposing hypotheses, which means the search path depends on the model provider, the prompt, and the data summary the agents see. That buys flexibility in feature engineering and metric choice. It costs reproducibility and it costs money per iteration.
The other comparison is the pipeline you would write yourself with scikit-learn and XGBoost. That gives you full control and no external API calls, but you own the feature encoding, the metric selection and the packaging. Plexe's contribution is that packaging step: a directory with artifacts/, src/, schemas/, config/, evaluation/, model.yaml and a README, plus a tarball. If you already have a packaging convention, the value proposition shrinks considerably.
For teams already on a platform, the WorkflowIntegration interface is the escape hatch. The README shows passing integration=MyCustomIntegration() into main and points at plexe/integrations/base.py for the full interface, which is how you would wire plexe into existing storage, tracking or deployment infrastructure instead of adopting its output layout.
Maintenance, upgrades and what the licence permits
The repository is not archived. The most recent push recorded is 2026-03-06, which is more than six months before today, so this is not a project to describe as actively developed. The release history shows v1.4.2, v1.4.3 and v1.4.4 all landing within a few days in early March 2026, which reads as a burst of patch releases rather than a steady cadence. Treat the version you pin as the version you will be running for a while.
Upgrade cost is shaped by the dependency list in pyproject.toml. Plexe pins tensorflow >=2.20.0,<2.21.0, keras >=3.12.0,<3.13.0, xgboost >=3.1.1,<3.2.0, pandas >=2.3.3,<2.4.0 and scikit-learn >=1.7.2,<1.8.0, and it splits numpy by Python version because numpy 2.x requires 3.11 or newer. Those upper bounds protect you from breakage but they also mean plexe will hold back a shared environment, particularly if another service in the same environment needs a newer pandas or scikit-learn.
On licensing: the project is Apache-2.0, which permits commercial use and modification and includes a patent grant. That covers the plexe code. It does not cover the LLM providers you route to, and it does not cover the datasets you feed in. The README does not address data residency or what is transmitted to providers, so if your columns are regulated, that is a question for your own review rather than something the documentation answers. None of this is legal advice.
Editorial conclusion
Adopt plexe if your data is tabular, your intent fits in one sentence, and you want a packaged artifact rather than a notebook. Do not adopt it if you need a documented rollback path, if you cannot send data or schema details to an LLM provider, or if your task is not tabular classification or regression. Verify first that your Python version is in the >=3.10,<3.13 range, that your provider keys work through LiteLLM, and that the generated package under work_dir/model/ loads without plexe installed.
Frequently asked questions
Which Python versions does plexe support?
The installation section states that plexe requires Python >= 3.10 and < 3.13. The pyproject.toml dependency block matches that range, and it splits numpy versions because numpy 2.x needs Python 3.11 or newer.
Which dataset formats can I pass to plexe?
The README says you provide a tabular dataset in Parquet, CSV, ORC or Avro, alongside a natural language intent. The CLI takes it through --train-dataset-uri and the Python API through data_refs.
Does the model plexe produces require plexe at inference time?
No. The README states the output package at work_dir/model/ has no dependency on plexe, so you build the model with plexe and deploy it anywhere. The package includes artifacts, source code, schemas, config, evaluation reports, model.yaml and a README.
Which LLM providers does plexe work with?
Plexe calls providers through LiteLLM, so any supported provider can in principle be configured. The README notes that only openai/* and anthropic/* models are actively tested and asks users to report issues with other providers.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/plexe-ai-plexe)