# TransmogrifAI: AutoML for Structured Data on Apache Spark

> Salesforce's Scala library wraps Spark ML in a strongly typed feature DSL and automates feature engineering, validation and model selection. It targets structured tabular data and teams already on Spark 2.4 and Scala 2.11.

**salesforce/TransmogrifAI** — TransmogrifAI (pronounced trăns-mŏgˈrə-fī) is an AutoML library for building modular, reusable, strongly typed machine learning workflows on Apache Spark with minimal hand-tuning

- Repository: https://github.com/salesforce/TransmogrifAI
- Website: https://transmogrif.ai
- Stars: 2,276 · Forks: 399
- Language: Scala
- License: BSD-3-Clause
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/salesforce-transmogrifai

## The problem TransmogrifAI targets: tabular ML glue code

Most tabular machine learning work is not modeling. It is writing feature extraction code, deciding which columns to keep, and comparing a handful of classifiers by hand. TransmogrifAI exists to compress that work into a declarative workflow. The README states the library was developed to accelerate machine learning developer productivity through automation, with an API that enforces compile-time type-safety, modularity and reuse. That sentence is the whole design brief.

The audience is specific. The README lists three cases: building production machine learning applications in hours rather than months, building models without a machine learning PhD, and building modular, reusable, strongly typed workflows. A data engineer who already knows Spark and Scala fits all three. A data scientist working in Python notebooks does not, and the library makes no attempt to serve them.

The claim in the README that automation achieves accuracies close to hand-tuned models with almost 100x reduction in time is the project's own marketing line, not an independently verified benchmark. Treat it as a design goal. The mechanism behind it, automated feature engineering followed by automated model selection, is visible in the Titanic example and is worth reading on its own terms.

## How the workflow works: features, transmogrify, sanityCheck, model selection

The README's Titanic example is the clearest description of the data flow. Data is read into a Spark DataFrame, then FeatureBuilder.fromDataFrame extracts a response feature and a set of predictor features. The response is typed as RealNN, which tells the library the target is numeric and not missing. Predictors arrive as a collection of typed features rather than raw columns.

From there the pipeline runs in four stages. predictors.transmogrify() performs automated feature engineering, turning raw typed features into a feature vector. survived.sanityCheck(featureVector, removeBadFeatures = true) runs automated feature validation and selection, and the flag tells it to drop features that fail. BinaryClassificationModelSelector().setInput(survived, checkedFeatures).getOutput() performs automated model selection. Finally an OpWorkflow is constructed with the input dataset and the result features, and train() fits the model.

The type system is the part that distinguishes this from writing Spark ML directly. Features carry their semantic type, so a feature declared as numeric cannot be silently passed where a text feature is expected. The README frames this as compile-time type-safety, which means many wiring mistakes surface at build time instead of at runtime on a cluster.

The model summary output shows what the selector actually did on that dataset: it evaluated Logistic Regression and Random Forest models with 3 folds and the AuPR metric, considered 3 Logistic Regression configurations and 16 Random Forest configurations, and selected a Random Forest classifier. The summary also prints hold-out and training metrics and top insights computed by correlation. That reporting is generated by the workflow, not assembled by the user.

## Installing TransmogrifAI and running a first workflow

The README does not give an install command. It points at the Quick Start and Documentation section and the project homepage at transmogrif.ai, and the Maven Central badge identifies the artifact as com.salesforce.transmogrifai:transmogrifai-core_2.11. The version shown in the badge is 0.7.0. If you use Gradle or Maven, that coordinate is what you add to your build; the repository ships a build.gradle and a pom.xml, so both build paths exist in the tree.

The badges also pin the runtime: Spark 2.4 and Scala 2.11. Those are not suggestions. A project on Spark 3.x or Scala 2.12 will need its own build from source, and the README does not document that path.

The Titanic example is the first real use. It needs a SparkSession and the library imports, then reads a CSV into a case class.

```scala
import com.salesforce.op._
import com.salesforce.op.readers._
import com.salesforce.op.features._
import com.salesforce.op.features.types._
import com.salesforce.op.stages.impl.classification._
import org.apache.spark.SparkConf
import org.apache.spark.sql.SparkSession

implicit val spark = SparkSession.builder.config(new SparkConf()).getOrCreate()
import spark.implicits._
```

With the session in place, the README reads the data through DataReaders.Simple.csvCase[Passenger], then extracts features and builds the pipeline in a few lines.

```scala
val passengersData = DataReaders.Simple.csvCase[Passenger](path = pathToData).readDataset().toDF()

val (survived, predictors) = FeatureBuilder.fromDataFrame[RealNN](passengersData, response = "survived")

val featureVector = predictors.transmogrify()
val checkedFeatures = survived.sanityCheck(featureVector, removeBadFeatures = true)
val pred = BinaryClassificationModelSelector().setInput(survived, checkedFeatures).getOutput()

val model = new OpWorkflow().setInputDataset(passengersData).setResultFeatures(pred).train()
println("Model summary:\n" + model.summaryPretty())
```

What the reader should see is a model summary. On the Titanic data the README shows the selector evaluating Logistic Regression and Random Forest across 3 folds, then printing the chosen Random Forest parameters, hold-out and training metrics including Precision, Recall, F1, AuROC and AuPR, and a short list of top insights ranked by correlation. If your run prints that block, the workflow completed.

## Where TransmogrifAI is the wrong tool

The library is built for structured data on Spark. The README's own framing is structured data, and every example works from a DataFrame of typed columns. There is no documented path for image, audio or free-text modeling. If your problem is a sequence model or a transformer fine-tune, this is not the layer you want.

The version constraints are the sharper limitation. Spark 2.4 and Scala 2.11 are what the badges state. Spark 2.4 reached the end of its line well before the most recent release listed here, and Scala 2.11 is older still. A team on a current Spark distribution will spend its first week on build compatibility rather than on modeling, and the README does not describe an upgrade path.

The release history reinforces the point. The newest release listed is 0.7.0 from 2020-06-11, with 0.6.1 from 2019-09-12 and 0.6.0 from 2019-07-12 before it. The repository's last push was on 2026-06-02, so there is activity in the tree, but the published artifact line has not moved in years. Anyone expecting new estimators or new Spark support from a dependency bump should check the changelog rather than assume.

Finally, the automation is opinionated. The selector chooses among the model families it knows, and the Titanic output shows Logistic Regression and Random Forest. If your problem needs a specific estimator, a custom loss, or a feature transformation the DSL does not expose, you will fight the abstraction. The README does not document an escape hatch for arbitrary Spark ML stages inside an OpWorkflow.

## TransmogrifAI compared with plain Spark ML pipelines

The honest alternative is Spark MLlib, which TransmogrifAI sits on top of. With Spark ML you assemble a Pipeline from StringIndexer, OneHotEncoder, VectorAssembler and an Estimator, then run a CrossValidator over a ParamGridBuilder grid. It is more code, and every column type decision is yours to make and to keep consistent between training and serving.

The difference in approach is where the automation lives. Spark ML automates execution: it distributes the fit and the grid search. TransmogrifAI automates the decisions: which features to build, which to drop, and which model family and hyperparameters to try. The transmogrify and sanityCheck stages have no direct Spark ML equivalent. They are the reason the Titanic example is roughly ten lines instead of a hundred.

The cost is control and coupling. A Spark ML pipeline is portable across Spark versions with modest edits, and the stages are documented by Apache. A TransmogrifAI workflow is portable only within the versions the library supports, and the behavior of transmogrify is a library implementation detail rather than something you configure stage by stage. If you need to explain exactly which transformations ran, Spark ML gives you a graph you wrote yourself. TransmogrifAI gives you a summary and a set of insights.

For a team already fluent in Spark ML and happy with its pipelines, switching buys less than the README implies. For a team repeatedly writing the same feature engineering boilerplate across projects, the typed feature DSL is the part worth evaluating.

## Maintenance, upgrades and the BSD-3-Clause licence

The repository is not archived, and the last push was on 2026-06-02. That is recent enough that the tree is being touched, but it says nothing about the release cadence. The Maven Central badge points at 0.7.0, and the release list confirms 0.7.0 as the newest entry, dated 2020-06-11. Plan your upgrade budget around the artifact line, not around commit activity.

Upgrade cost has two axes. The first is the TransmogrifAI version itself; with only three releases listed, the surface is small and the changelog in the repository is the place to check what moved. The second is the Spark and Scala pair. Because the library compiles against Spark 2.4 and Scala 2.11, moving your cluster forward means either staying behind or building the library against a newer pair yourself. The README does not document a supported path for the latter.

The licence is BSD-3-Clause, stated in the README badge and present as a LICENSE file at the repository root. That is a permissive licence, which in practice means you can use the library in commercial products provided you keep the copyright notice and the disclaimer, and you do not use the project's name to endorse your own work. This is a description of the licence text, not legal advice; read the LICENSE file and talk to your own counsel about your distribution model.

One operational detail worth noting for regulated environments: the repository includes a SECURITY.md, which is where the project states how vulnerabilities should be reported. If your process requires a disclosure channel, that file is the entry point.

## Conclusion

Adopt TransmogrifAI if you already run Spark 2.4 with Scala 2.11 and want a typed, reusable workflow for tabular prediction rather than a notebook of ad hoc feature code. Do not adopt it if you need text, image or deep learning pipelines, or if you cannot pin your Spark and Scala versions to what the README states. Before committing, verify that the published transmogrifai-core_2.11 artifact resolves in your build, that your data fits the csvCase reader or another reader in the readers module, and that the last push date of 2026-06-02 is consistent with the release you plan to deploy, since the newest release listed is 0.7.0 from 2020-06-11.

## FAQ

### What is TransmogrifAI and who is it for?

It is an AutoML library written in Scala that runs on Apache Spark, aimed at structured data and at developers who want to build modular, strongly typed machine learning workflows with less hand-tuning. The README lists production applications, users without a machine learning PhD, and reusable workflows as the target cases.

### How do I install TransmogrifAI?

The README does not give install steps; it points to the Quick Start and Documentation section and the project homepage. The Maven Central badge identifies the artifact as com.salesforce.transmogrifai:transmogrifai-core_2.11 at version 0.7.0, and the repository contains both a build.gradle and a pom.xml.

### Which Spark and Scala versions does TransmogrifAI require?

The README badges state Spark 2.4 and Scala 2.11. Those are the versions the published artifact targets, and the README does not document a supported path for building against newer versions.

## Sources

- [License: BSD-3-Clause](https://github.com/salesforce/TransmogrifAI/blob/master/LICENSE)
- [Project website](https://transmogrif.ai)
- [README](https://github.com/salesforce/TransmogrifAI/blob/master/README.md)
- [Releases](https://github.com/salesforce/TransmogrifAI/releases)
- [salesforce/TransmogrifAI on GitHub](https://github.com/salesforce/TransmogrifAI)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/salesforce-transmogrifai
