Open-source project
mindsdb/lightwood avatar
mindsdb/lightwood

Lightwood: JSON-AI Pipelines for Tabular and Time-Series Problems

Lightwood is Legos for Machine Learning.

511 stars101 forksPythonGPL-3.0

At a glance

What is it?
Lightwood is an AutoML framework from MindsDB that turns a pandas DataFrame plus a target column into a generated Python training pipeline. It suits engineers who want to override individual pipeline steps rather than accept a black box, and it ships under GPL-3.0 with a Python 3.10 to 3.13 dependency range.
Who is it for?
Adopt Lightwood if you have a pandas DataFrame, a single column you want to predict, and a reason to intervene in the pipeline rather than accept whatever a black-box AutoML service returns. Do not adopt it if you need a stable API surface across upgrades, if you cannot take a GPL-3.0 dependency, or if you are not prepared to read generated Python.
Can I use it commercially?
Yes, with conditions. GPL-3.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
Is it still maintained?
Yes. The repository last received commits 33 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The Problem Lightwood Targets: Repetitive Pipeline Code, Not Model Choice

Most AutoML tools answer the question "which model wins on this table." Lightwood answers a different one: "how do I stop rewriting the same preprocessing, encoding and training scaffolding for every new dataset." The README states the goal directly, that users should focus on what they want to do with their data "without needing to write repetitive boilerplate code around machine learning and data preparation."

That framing matters for who should care. If you are a data scientist who already has a preferred gradient boosting setup and just wants hyperparameter search, Lightwood is more machinery than you need. If you are an application engineer who keeps being handed CSVs and asked for a prediction endpoint, the abstraction is aimed at you. The framework handles numbers, dates, categories, tags, text, arrays and multimedia formats, and it has a time-series mode for problems with between-row dependencies. Those are the shapes of data that usually force hand-written preprocessing, which is exactly where the boilerplate complaint comes from.

How JSON-AI Works: Three Steps, Three Overridable Objects

Lightwood splits the pipeline into preprocessing and cleaning, feature engineering, and model building and training. Each step has a named object you can replace.

Preprocessing runs per column. Lightwood performs what the README calls "a brief statistical analysis" to guess the data type, writes that guess into a JSON-AI syntax object, then cleans each column according to the identified type and splits the data into train, dev and test. The cleaner and splitter objects own those two jobs.

Feature engineering converts cleaned columns into numerical representations through encoders. The distinction the README draws is between rule-based encoders, which transform data by fixed instructions such as normalizing a numeric column, and learned encoders, which produce a representation only after training, the way a language model produces a [CLS] token. Encoders are assigned per column based on data type, and you can override that assignment either at the column level or at the data-type level. All of them inherit from BaseEncoder.

The final stage is the mixer, defined in the README as a predictive model that takes encoded feature data and outputs a prediction for the target. Mixers inherit from BaseMixer, and the framework is predominantly PyTorch based while stating it can support other models. The important architectural consequence is that the JSON-AI object is not just configuration. It is an intermediate representation that Lightwood compiles into Python code, which means the pipeline you inspect is the pipeline that runs.

Installing Lightwood and Training a First Predictor

Install from PyPI. The README notes that depending on your environment you may need pip instead of pip3, and it recommends a virtual environment.

bash
pip3 install lightwood

The dependency declarations in pyproject.toml pin the supported interpreter range to Python 3.10 through 3.13. The README's development section is stricter, telling contributors to use a version in the range >=3.8, < 3.11, which is narrower than what the packaging metadata allows. If you are installing the released package rather than cloning the repository, trust the metadata; if you are running the test suite locally, the README's own instructions are the ones to follow.

The core workflow is four calls. You load a DataFrame, declare which column you want to predict, generate the JSON-AI description, compile it to Python, and train. The README gives this example:

python
import pandas as pd
from lightwood.api.high_level import (
    ProblemDefinition,
    json_ai_from_problem,
    code_from_json_ai,
    predictor_from_code,
)

if __name__ == '__main__':
    df = pd.read_csv("https://raw.githubusercontent.com/mindsdb/benchmarks/main/benchmarks/datasets/hdi/data.csv")
    pdef = ProblemDefinition.from_dict({"target": "Development Index"})
    json_ai = json_ai_from_problem(df, problem_definition=pdef)
    code = code_from_json_ai(json_ai)
    predictor = predictor_from_code(code)
    predictor.learn(df)

After learn completes, the README splits the preprocessed data and prints predictions:

python
test_df = predictor.split(predictor.preprocess(df))["test"]
preds = predictor.predict(test_df).iloc[:10]
print(preds)

You should see a ten-row table. The README also shows two optional print statements, json_ai.to_json() and print(code), and those are worth running the first time you use the framework. The generated code is where you find out which encoders were chosen for each column and which mixer was selected, and it is the artifact you edit when defaults are wrong.

Where Lightwood Gets in the Way

The generated-code model is the framework's main selling point and its main hazard. Once you call code_from_json_ai, you hold a Python file that describes your pipeline. Nothing in the README describes a round trip from edited code back into JSON-AI, so if you hand-modify the generated pipeline, regenerating it from the problem definition will discard your changes. The documentation does not describe a merge or diff mechanism for that case.

The dependency list is another real constraint. The pins include torch at ~2.8.0, pandas at exactly 2.2.3, scikit-learn at exactly 1.5.2, xgboost capped at <=1.8.0, and dill at exactly 0.3.6. Installing Lightwood into an environment that already has a different scikit-learn or pandas version means one of the two has to move. This is not unusual for an ML framework, but it is a hard boundary worth checking before you commit.

The time-series mode is where the optional extras matter. pyproject.toml declares pystan, prophet and neuralforecast under an extra_ts extra, librosa under audio, torchvision and pillow under image, and lightgbm under extra. A plain pip3 install lightwood does not pull those in, so a time-series problem that needs one of those backends requires installing the corresponding extra. The README's installation section does not walk through the extras, which leaves you reading pyproject.toml to find out what is optional.

Finally, Lightwood assumes tabular input in a pandas DataFrame. The README describes multimedia formats as supported data types, but the entry point is still a DataFrame, and the time-series support is scoped to problems with between-row dependencies. If your data lives in a graph, a streaming topic, or a feature store, this is the wrong layer.

Lightwood Versus a Plain scikit-learn or PyTorch Pipeline

The honest comparison is not against another AutoML product but against the code you would write yourself. A hand-built scikit-learn pipeline gives you a Pipeline object, a ColumnTransformer, and a grid search. You write more lines, but every line is yours, there is no generated intermediate, and your dependency versions are whatever you choose.

Lightwood's difference is the JSON-AI layer. You get a declarative description of the whole pipeline that can be inspected, printed and compiled, plus the ability to swap a single encoder or mixer without rewriting the rest. The README's BYOM section states that user architectures are supported "so long as you follow the abstractions provided within each step," and the tutorials cover custom cleaner, custom splitter, custom explainer and custom mixer. That is a real extension surface, not a configuration flag.

The cost of that surface is the abstraction itself. To replace a mixer you inherit from BaseMixer and conform to its interface; to replace an encoder you inherit from BaseEncoder. Mixing Lightwood with a hand-rolled PyTorch training loop means understanding both the framework's contracts and your own. For a one-off model, that overhead does not pay for itself. For a team producing pipelines repeatedly from similar data, the encoder assignment rules and the code generation are the parts that save time.

Licence, Maintenance Signals and Upgrade Cost

Lightwood is licensed GPL-3.0-only, declared both in pyproject.toml and in the repository's LICENSE file. This is a copyleft licence, which is a different proposition from the permissive licences most Python ML tooling uses. The framework is also a dependency of MindsDB, a commercial product, which is worth knowing when you evaluate the licence in your own context. What that means for your distribution obligations depends on how you link and ship the code, and that is a question for your legal team, not for this article.

The repository is not archived, and the last push was on 2026-08-28. The most recent tagged release listed is v25.12.1.0 from 2025-12-02, preceded by v25.9.1.0 and v25.7.5.1. The version string in pyproject.toml matches v25.12.1.0, so the packaging metadata tracks the release tags rather than the branch head.

Upgrade cost is dominated by the pinned dependency set. Because pandas, scikit-learn and dill are pinned to exact versions, a major-version bump in any of them requires a coordinated Lightwood release. If your application also depends on those libraries directly, plan for the upgrade to move several packages at once. The README's development instructions also ask contributors to run `python -m unittest discover tests` from the cloned directory, which is the check the project itself uses to validate a working tree.

Editorial conclusion

Adopt Lightwood if you have a pandas DataFrame, a single column you want to predict, and a reason to intervene in the pipeline rather than accept whatever a black-box AutoML service returns. Do not adopt it if you need a stable API surface across upgrades, if you cannot take a GPL-3.0 dependency, or if you are not prepared to read generated Python. Verify first that your Python version falls inside the declared range, that your target column survives the automatic type inference the way you expect, and that the generated code is something your team is willing to own, because after code_from_json_ai it is your file, not the framework's.

Frequently asked questions

What is Lightwood?

It is an AutoML framework from MindsDB that generates and customizes machine learning pipelines through a declarative syntax called JSON-AI. It works with pandas DataFrames and supports numbers, dates, categories, tags, text, arrays and multimedia formats, plus a time-series mode.

How do I install Lightwood?

The README gives `pip3 install lightwood`, noting that you may need pip instead of pip3 depending on your environment, and recommends a virtual environment. The packaging metadata declares support for Python 3.10 through 3.13.

Do I need to specify more than the target column?

No. The README states that the only thing a user needs to specify in the ProblemDefinition dictionary is the name of the column to predict, via the target key. Lightwood infers the data types of the remaining columns through a statistical analysis.

Can I replace the model Lightwood picks?

Yes. Mixers inherit from BaseMixer and encoders inherit from BaseEncoder, and the README says user architectures are supported as long as they follow the abstractions provided within each step. The tutorials cover custom cleaner, custom splitter, custom explainer and custom mixer.

What licence does Lightwood use?

GPL-3.0-only, declared in pyproject.toml and in the repository's LICENSE file. That is a copyleft licence, so the obligations differ from permissive Python ML libraries.

Official sources

  1. Issues
  2. License: GPL-3.0
  3. mindsdb/lightwood on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/mindsdb-lightwood.svg)](https://hysenlabs.com/projects/mindsdb-lightwood)