Library / SDK
pyg-team/pytorch-frame avatar
pyg-team/pytorch-frame

PyTorch Frame: A Modular Deep Learning Library for Heterogeneous Tabular Data

Tabular Deep Learning Library for PyTorch

799 stars73 forksPythonMIT

At a glance

What is it?
PyTorch Frame separates tabular modeling into Materialization, FeatureEncoder, TableConv and Decoder, so numerical, categorical, text and image columns can share one model. It suits PyTorch users who need column-type flexibility, not teams looking for a drop-in GBDT replacement.
Who is it for?
Adopt PyTorch Frame if your tables mix column types that GBDTs handle badly, such as text embeddings or image embeddings, and you already train in PyTorch. Do not adopt it if your data is plain numerical and categorical and a tuned XGBoost or CatBoost already wins in the project's own benchmark, or if you need a stable API across versions: the project labels itself Development Status 4 - Beta.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 22 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The column-type problem PyTorch Frame was built for

Tree-based models such as GBDT handle numerical and categorical columns well and remain the default choice for tabular prediction. The README states the limitation directly: those models have integration difficulties with downstream models and struggle with complex column types such as texts, sequences and embeddings. A table with a product description, a timestamp and a thumbnail image does not fit a gradient-boosted tree without manual feature engineering for each type.

PyTorch Frame targets that gap. It is a deep learning extension for PyTorch for heterogeneous tabular data, and the README lists the column types it supports: numerical, categorical, multicategorical, text_embedded, text_tokenized, timestamp, image_embedded and embedding. The intended audience is PyTorch users who want to train a neural network over a pandas DataFrame without writing a separate encoder for every column, and researchers who want to swap model components rather than rewrite a training loop.

The second stated goal is integration with large language models. The README points to examples/llm_embedding.py and examples/transformers_text.py, which encode text through OpenAI, Cohere, Voyage AI or Hugging Face and then train those embeddings alongside other column types. That is a different use case from tabular classification on clean numeric features, and it is the one where the library's design pays off.

Materialization, FeatureEncoder, TableConv, Decoder

The architecture is a four-stage pipeline, and the README names each stage. Materialization converts a raw pandas DataFrame into a TensorFrame, the internal representation used for training. FeatureEncoder then maps the TensorFrame to hidden column embeddings with shape [batch_size, num_cols, channels]. TableConv models column-wise interactions over those hidden embeddings, and Decoder pools the result into an embedding or prediction per row.

The important design decision is that TableConv keeps the [batch_size, num_cols, channels] shape. Column identity is preserved through the stack rather than flattened into a single feature vector at the start. That is what allows attention-style interaction between columns, and it is also why the library can accept a text column and a numerical column in the same tensor: each has its own encoder, and the interaction happens after encoding.

The modular split is not just organizational. The README says the setup lets users experiment with different architectures, and the repository backs that up. Under examples/ there are separate scripts for ExcelFormer, TabTransformer, TabNet, Trompt, Trompt with multi-GPU, Revisiting Deep Learning Models, TabPFN classification and a tuned GBDT baseline. Each one is an assembly of the same three pieces. If you want to change how columns interact, you replace TableConv and leave the encoder and decoder alone.

The cost of this design is indirection. A simple MLP on a numeric table becomes an encoder, a convolution block and a decoder, which is more code than a torch.nn.Sequential. That overhead only makes sense when the column types or the interaction modeling justify it.

Installing PyTorch Frame and running the tutorial example

The package is published on PyPI as pytorch-frame, and pyproject.toml declares requires-python >= 3.10 with classifiers listed through Python 3.14. The README's installation section points to the documentation site for the current instructions, so check there for the exact command matching your PyTorch build. The distribution name is the one declared in pyproject.toml under [project]:

toml
name="pytorch-frame"

The base dependencies are numpy, pandas, torch, tqdm, pyarrow and Pillow. If you want the benchmark and baseline tooling, the full extra pulls in scikit-learn, xgboost, optuna, optuna-integration, catboost, lightgbm, datasets and torchmetrics. Those extras are declared in pyproject.toml as follows, and the version constraints on xgboost are worth reading before you install:

toml
[project.optional-dependencies]
full=[
    "scikit-learn",
    "xgboost>=1.7.0, <2.0.0",
    "optuna>=3.0.0",
    "optuna-integration",
    "mpmath==1.3.0",
    "catboost",
    "lightgbm",
    "datasets",
    "torchmetrics",
]

If your environment already has XGBoost 2.x installed, the full extra will conflict with it. Install the base package and add XGBoost yourself if you only need the deep models.

For a first real run, the repository ships examples/tutorial.py. The Quick Tour in the README describes what a model looks like: an encoder that maps a TensorFrame to [batch_size, num_cols, channels], a stack of convolutions that preserves that shape, and a decoder that pools to [batch_size, out_channels]. The tutorial example follows that pattern end to end on a real dataset, so it is the fastest way to see the data flow before wiring in your own DataFrame. Run it from a checkout of the repository so the example can import torch_frame from the local tree, or install the package first and run the file directly.

The first thing to verify on your own data is that every column maps to a supported stype. The documentation has a dedicated page on handling heterogeneous stypes, and a column that fits none of the listed types will need to be converted or dropped before Materialization.

Where PyTorch Frame is the wrong tool

The project's own benchmark directory compares deep tabular models against GBDTs, and the README links to it without claiming a clean win. That is the honest framing: for many tabular problems, a tuned gradient-boosted tree is competitive or better, and the library ships a tuned GBDT baseline in examples/tuned_gbdt.py precisely so you can check that on your data. If your columns are all numerical and low-cardinality categorical, the deep models here are unlikely to justify the training cost.

The API stability question is separate. pyproject.toml classifies the project as Development Status 4 - Beta. The release history shows 0.2.4 in January 2025, 0.2.5 in February 2025 adding Python 3.13 and PyTorch 2.6 support, and 0.3.0 in November 2025 described as broader compatibility and usability enhancements. Pre-1.0 releases with compatibility-focused changelogs mean you should pin a version and read the changelog before upgrading, particularly if you subclass FeatureEncoder or TableConv.

There is also a data-size boundary. Column-wise interaction over [batch_size, num_cols, channels] scales with the number of columns, so a table with thousands of sparse columns is a poor fit compared with a model that selects or hashes features first. The README does not document a column-count limit, so treat this as a design consequence of the TableConv stage rather than a stated constraint, and measure it on your own table.

How it differs from PyTorch Tabular

PyTorch Tabular is the other well-known PyTorch library in this space, and the two make different bets. PyTorch Tabular is oriented around a configuration-driven workflow: you describe the model and training setup, and the library assembles a standard supervised pipeline around a chosen architecture.

PyTorch Frame instead exposes the model internals as composable classes. The README describes the modular design of FeatureEncoder, TableConv and Decoder, and the examples show each published model rebuilt from those parts. The practical difference is where you spend your effort. With a configuration-driven library, adapting to an unusual column type means working within its data-handling abstractions. With PyTorch Frame, you write an encoder for that column and pass it in, which is more work but leaves the rest of the pipeline untouched.

The second difference is scope. PyTorch Frame explicitly supports text_embedded, text_tokenized, image_embedded and embedding column types, and the README documents integration with OpenAI, Cohere, Voyage AI and Hugging Face for producing text embeddings. It also integrates with PyG for learning over relational databases, with RelBench named as the example. If your table is a flat numeric matrix, that extra surface area is unused weight.

Licence, releases and the cost of staying current

PyTorch Frame is MIT licensed, with the licence file declared in pyproject.toml as license = "MIT" and license-files = ["LICENSE"]. MIT is permissive: you can use, modify and redistribute the code, including in commercial products, provided the copyright notice and licence text are retained. That is a description of the licence terms, not legal advice; if your organisation has specific obligations around attribution or bundled dependencies, have counsel review the LICENSE file and the transitive dependency licences.

The upgrade cost is the part to plan for. The base dependency set is small and stable: numpy, pandas, torch, tqdm, pyarrow and Pillow. The full extra is where version pressure lives, because it pins xgboost>=1.7.0, <2.0.0 and mpmath==1.3.0. That exact mpmath pin is the kind of constraint that collides with other scientific packages in a shared environment. If you only need the deep models, install the base package and keep the baseline tooling in a separate environment.

Maintenance looks current: the last push to the default branch was on 2026-09-07, and the repository is not archived. The release cadence is irregular rather than frequent, with three releases listed between January 2025 and November 2025. For a pre-1.0 research-oriented library, that is a reasonable signal, but it does mean you should not expect a long-term support branch. Pin the version in your lockfile and test the upgrade against your own model code before moving.

Editorial conclusion

Adopt PyTorch Frame if your tables mix column types that GBDTs handle badly, such as text embeddings or image embeddings, and you already train in PyTorch. Do not adopt it if your data is plain numerical and categorical and a tuned XGBoost or CatBoost already wins in the project's own benchmark, or if you need a stable API across versions: the project labels itself Development Status 4 - Beta. Before committing, run the examples/tutorial.py script against your own DataFrame, check that every column maps to one of the supported stypes, and confirm the pinned xgboost>=1.7.0, <2.0.0 constraint in the full extra does not conflict with the rest of your environment.

Frequently asked questions

Is PyTorch still relevant in 2026?

The README does not address PyTorch's overall standing. It does state that PyTorch Frame is a deep learning extension for PyTorch and builds directly upon it, so the library's usefulness depends on the PyTorch runtime you already have.

Is TensorFlow losing to PyTorch?

The README makes no comparison between PyTorch and TensorFlow. It only describes PyTorch Frame as an extension for PyTorch, with dependencies on torch and integration with other PyTorch libraries such as PyG.

Is ChatGPT built on PyTorch?

The README does not discuss ChatGPT or which framework it uses. Its only mention of large language models is that PyTorch Frame can encode text through OpenAI, Cohere, Voyage AI or Hugging Face embeddings and train those alongside other column types.

Does Nvidia use PyTorch?

The README does not mention Nvidia or any hardware vendor. It documents the Python and PyTorch versions supported, including 0.2.5 adding Python 3.13 and PyTorch 2.6 support.

Official sources

  1. License: MIT
  2. Project website
  3. pyg-team/pytorch-frame on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/pyg-team-pytorch-frame.svg)](https://hysenlabs.com/projects/pyg-team-pytorch-frame)