predikit: wrapping scikit-learn and XGBoost models as LLM-callable tools
The missing bridge between your ML models and your AI agents. MCP tools use the same Pydantic input validation and model execution as direct invoke() calls.
At a glance
- What is it?
- predikit generates OpenAI and LangChain tool schemas from a Pydantic input model and routes prediction calls through the same validation path. It is a thin adapter for Python teams already serving sklearn or XGBoost estimators to an agent.
- Who is it for?
- Adopt predikit if you already have fitted sklearn or XGBoost estimators and want them exposed to an agent without hand-writing JSON Schema, and verify first that your Pydantic field names match the training column names exactly, because the library maps features by name and raises ValueError on any mismatch.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 12 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The glue code predikit removes between .predict() and a tool call
A trained estimator answers one question: given an array of numbers, what is the label or value. An agent needs a different contract. It needs a name, a description, a JSON Schema describing each argument, and a callable that rejects malformed input before the model sees a stack trace. Most teams write that layer by hand, once per model, and it drifts.
predikit targets that gap. The README states the project wraps any trained scikit-learn or XGBoost model as an LLM-callable tool with auto-generated JSON schemas and typed I/O. The intended reader is a Python developer who already has a fitted estimator in memory or in a registry and wants it reachable from OpenAI function calling, LangChain, or any tool-calling API that accepts JSON Schema. The pyproject classifiers list Development Status 4 - Beta, so treat the API as settled enough to use but not frozen.
The library does not train, tune, or serve models over HTTP. It is an adapter between a Pydantic v2 input model and an estimator's predict call, plus a serializer that turns that adapter into whatever schema the calling framework expects.
How ModelTool turns a Pydantic class into a validated prediction call
The core object is ModelTool. Its constructor takes the fitted estimator, a name and description the LLM sees, a Pydantic BaseModel describing inputs, and an output_name with output_description for the prediction key. Construction builds the schema; the call path is separate.
The README describes .invoke(input_dict) as validate, predict, return {output_name: value}. That ordering is the whole design. Pydantic v2 validates and coerces the dictionary the agent produced, the validated fields are mapped onto the feature vector, the estimator predicts, and the result comes back under the key you named. .ainvoke() is the async version of the same path, which matters when the calling agent is itself async and you do not want a blocking predict inside the event loop.
Export is one call per framework. .to_openai() returns a dict in OpenAI function-calling shape. .to_langchain() returns a StructuredTool. .to_callable() returns a plain Python function for anything else. ToolRegistry groups several ModelTools and returns a list from .to_openai() or .to_langchain(), with .get("name") to retrieve one. ModelEnsemble runs several tools and reconciles outputs through one of five strategies: "collect" merges outputs into one dict and allows different output_name values, "mean" averages numeric outputs, "vote" takes a majority class vote, and "weighted_mean" and "weighted_vote" do the same with a weights list. The mean and vote strategies require all tools to share output_name, which is the constraint that decides whether an ensemble is a drop-in or a rewrite.
Installing predikit and wrapping an Iris classifier
The README gives pip as the install path, with optional extras for each integration. Base install pulls pydantic>=2.0, scikit-learn>=1.2 and numpy>=1.24 according to pyproject; xgboost, langchain, mlflow, snowflake and cli are extras.
pip install predikit
# Optional extras
pip install predikit[xgboost] # XGBoost support
pip install predikit[langchain] # LangChain StructuredTool export
pip install predikit[mlflow] # MLflow Model Registry loader
pip install predikit[snowflake] # Snowflake Model Registry loaderThe quick start trains a logistic regression on the Iris dataset, declares a Pydantic model with one field per feature, and wraps the estimator. The README's own snippet is truncated mid-construction, but the shape is clear from the Core API section: the tool takes model, name, description, input_schema, output_name and output_description.
from pydantic import BaseModel, Field
from sklearn.datasets import load_iris
from sklearn.linear_model import LogisticRegression
from predikit import ModelTool
X, y = load_iris(return_X_y=True)
clf = LogisticRegression(max_iter=200).fit(X, y)
class IrisInput(BaseModel):
sepal_length: float = Field(description="Sepal length in cm")
sepal_width: float = Field(description="Sepal width in cm")
petal_length: float = Field(description="Petal length in cm")
petal_width: float = Field(description="Petal width in cm")A first real use is two lines once the tool exists. Pass the tool's OpenAI schema to the API, then call the tool locally with the arguments the model returned.
tool = ModelTool(model=clf, name="classify_iris", ...)
tool.to_openai() # OpenAI function schema, ready to pass to the API
tool.invoke({"sqft": 2200}) # -> {"price_usd": 370730}The second line in that README example is a regression tool, not the Iris one, so expect .invoke() to return a dict keyed by your output_name with the estimator's prediction as the value. If the field names do not match the training columns, the call raises ValueError before the estimator runs. The repository also ships runnable examples under examples/, including 01_basic_sklearn.py, 04_confidence_routing.py and 05_multi_model_ensemble.py, which is where to look for the confidence and ensemble paths the README only summarizes.
The field naming rule is the sharpest edge in the library
predikit maps inputs to features by name, not position. The README states this in bold terms: schema field names must exactly match the column names the model was trained on. If you trained on a DataFrame with columns sqft and bedrooms, the Pydantic fields must be sqft and bedrooms, not square_footage and beds.
The failure is loud, which is the good part. A mismatch raises ValueError naming the missing features and listing both what the schema has and what the model expects. You find out at construction or first call rather than from a silently wrong prediction. The bad part is that the constraint is not negotiable in code. There is no alias map, no rename hook documented in the README. If your training pipeline produced cryptic column names, that ugliness is now part of the schema the LLM reads.
The README's tip says that if you trained with a numpy array there are no feature names to check against. That is the second edge. Estimators fitted on raw arrays bypass the name check entirely, which means the validation guarantee that makes the library worth using is weakest exactly where the input is least self-describing. For a numpy-trained model you are back to trusting positional order, and nothing in the README suggests predikit can recover names that were never stored.
Confidence thresholds and registry loaders, and where the documentation runs thin
The README's comparison table lists confidence handling as a feature: confidence_threshold plus on_low_confidence, so a low-confidence prediction can be routed somewhere other than the agent. The examples directory has 04_confidence_routing.py, which is the concrete reference. The README does not spell out the accepted values of on_low_confidence or what happens to the returned dict when a prediction falls below the threshold, so read the example before designing around it.
Registry loaders are the other convenience. from_mlflow() and from_snowflake() are listed as loaders behind the mlflow and snowflake extras, with examples 06_mlflow_loader.py and 07_snowflake_loader.py. The README does not document what these loaders return, whether they construct a ModelTool directly or a bare estimator, or how they handle registry stages and versions. If your deployment story depends on pulling a specific model version from MLflow, that is the first thing to verify in the example rather than assume from the table.
The same thinness applies to the roadmap and CHANGELOG, which are present in the repository but not reproduced in the README. Version numbers are the reliable signal: pyproject lists 0.6.3 while the most recent tagged release is v0.6.2, so the working tree is one patch ahead of the last release.
What predikit does not do, and the honest alternative
predikit is not a model server. There is no HTTP endpoint, no container, no batching layer, no request queue. The tool runs in the same Python process as the agent. That is fine for a single agent with a modest model, and wrong for a shared prediction service that several agents or languages need to reach, because every caller must import the library and hold the estimator in memory.
It is also not a training or feature-engineering framework. The name-matching rule means feature construction stays your problem. If your production features are computed by a separate pipeline, predikit validates the values it receives but cannot tell you whether those values were computed the way the training data was.
The direct alternative is writing the adapter yourself. A hand-written tool is a Pydantic model plus a function that calls model.predict and a dict you pass to the OpenAI API. That is perhaps thirty lines per model, and it removes the dependency entirely. The difference in approach is maintenance surface: hand-written adapters must be updated for each framework's schema changes, while predikit centralizes that translation in .to_openai() and .to_langchain(). If you have one model and no plans for ensembles or registries, the hand-written version is competitive. The gap widens with the number of models, because ToolRegistry and the five ensemble strategies are the parts that are genuinely tedious to reimplement.
For teams that need a network boundary rather than an in-process adapter, a model serving framework is the right tool and predikit is not, regardless of how well the schema generation works.
Licence, maintenance and what an upgrade actually costs
The project is MIT licensed, stated in both the README and pyproject. MIT permits commercial use, modification and redistribution with the licence text retained. That is permissive and carries no copyleft obligation on your own code. This is not legal advice; if you redistribute predikit inside a product, have your own counsel confirm the notice requirements.
The repository is not archived, and the last push was on 2026-08-22. Releases v0.6.0 and v0.6.1 landed on 2026-08-15 and v0.6.2 on 2026-08-22, so the recent cadence is fast. The classifier still says Beta, and the version is 0.x, which under common convention means the maintainer reserves the right to break the API between minor versions. Budget for reading the CHANGELOG before each minor upgrade rather than pinning and forgetting.
The upgrade cost is concentrated in two places. First, the constructor signature of ModelTool and ModelEnsemble, since those are what your code instantiates. Second, the shape of .to_openai() output, since that is what you hand to the LLM provider and what the provider's own API changes will force. The dependency floor is low: Python 3.10 or newer, Pydantic 2, scikit-learn 1.2, numpy 1.24. Optional extras add their own version floors, and the mcp extra is pinned to mcp>=1.28,<2, which is the one dependency with an upper bound and therefore the one most likely to block an upgrade.
Editorial conclusion
Adopt predikit if you already have fitted sklearn or XGBoost estimators and want them exposed to an agent without hand-writing JSON Schema, and verify first that your Pydantic field names match the training column names exactly, because the library maps features by name and raises ValueError on any mismatch. Skip it if your models are PyTorch or TensorFlow graphs, if you need a network endpoint rather than an in-process tool, or if you cannot rename columns, since the numpy-array path leaves the library with no feature names to check against.
Frequently asked questions
Does predikit work with models other than scikit-learn and XGBoost?
The README and pyproject list scikit-learn and XGBoost as the supported models, with XGBoost behind the xgboost extra. No other estimator families are documented.
Why does predikit raise a ValueError about missing model features?
predikit maps inputs to features by name, not position, so the Pydantic schema field names must exactly match the column names the model was trained on. When they do not match, the error names the missing features and lists both what the schema has and what the model expects.
Can predikit serve a model over HTTP?
No. The tool runs in the same Python process as the caller and exposes .invoke(), .ainvoke(), .to_openai(), .to_langchain() and .to_callable(). No network endpoint is documented.
What ensemble strategies does predikit support?
ModelEnsemble supports five: "collect", "mean", "vote", "weighted_mean" and "weighted_vote". The mean and vote strategies require all tools to share the same output_name, while "collect" merges outputs and allows different ones.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/tejas-ta-predikit)