# Pentaho Data Integration (Kettle): building PDI from source

> Pentaho Data Integration, also called Kettle, is a Java ETL suite whose Spoon desktop client and Carte server ship from a Maven build. This article covers what the repository contains, how to build and package it, and where the project stops being the right tool.

**pentaho/pentaho-kettle** — Pentaho Data Integration ( ETL ) a.k.a Kettle

- Repository: https://github.com/pentaho/pentaho-kettle
- Stars: 8,401 · Forks: 3,579
- Language: Java
- License: NOASSERTION
- Published: 2026-09-22 · Updated: 2026-09-22 · Language: en
- Canonical page: https://hysenlabs.com/projects/pentaho-pentaho-kettle

## What Pentaho Data Integration is, and who the repository is for

Pentaho Data Integration is an extract, transform, load suite written in Java and distributed under the name Kettle. The repository is not a single application. It is a Maven multi-module build whose modules map onto distinct runtime pieces: core holds the core implementation, engine holds the PDI engine, ui holds the user interface, dbdialog holds the database dialog, engine-ext holds engine extensions, plugins holds the core plugins, and assemblies produces the distribution archive. A separate integration module holds cross-module integration tests.

The audience is therefore narrower than the product name suggests. Someone who wants to click through a graphical ETL tool and load CSV files into a warehouse is not the reader of this repository. The reader is a developer or platform engineer who needs to build the distribution, patch a plugin, or embed the engine in another Java application. The README describes build prerequisites and test conventions, not end-user workflows, and it points people who want help to the Hitachi Vantara community forum rather than to documentation inside the repository.

That split matters when you evaluate the project. The community edition and the commercial product share a name and a codebase history, but this repository documents only the source build. Anything about support tiers, licensing terms for the commercial edition, or prebuilt installers is outside what the README covers.

## The module layout and how a build turns into a runnable distribution

The data flow of the build is straightforward once you read the module list. Source in core, engine, ui and the plugin modules is compiled and tested by Maven. The assemblies module then aggregates the compiled artifacts into distribution archives. The README states that packaged results land in the target/ sub-folders of assemblies/*, and gives a concrete example: a distribution of the Desktop Client (CE) can be found in assemblies/client/target/pdi-ce-*-SNAPSHOT.zip.

That single path tells you a lot. The desktop client is the artifact most people mean when they say Spoon, and it is produced by the client assembly rather than by the root build directly. The engine and the graphical client are separate concerns that happen to be packaged together for the desktop case. A Carte server deployment, by contrast, would come from a different assembly, and the repository root carries CarteAPIDocumentation.md and Carte-jmeter.jmx as evidence that the server side has its own API surface and its own load-testing setup.

Two build flags appear in the README and are worth knowing before you run anything. Passing -Drelease triggers obfuscation or uglification as needed, which means a release build is not byte-identical to an ordinary one. Passing -Dmaven.test.skip=true skips tests, and the README explicitly discourages it. There is also a checkstyle check available through mvn checkstyle:check, which the contributing section treats as part of the normal pull request path.

## Building Pentaho Data Integration with Maven

The prerequisites are stated plainly: Maven version 3 or later, Java JDK 11, and a specific settings.xml placed in your <user-home>/.m2 directory. That settings file is hosted in the pentaho/maven-parent-poms repository, which means the build depends on artifact resolution configuration that lives outside this repository. If your organisation mirrors Maven Central or blocks external repositories, that file is the first thing to reconcile.

The full build is a single command:

```bash
mvn clean install
```

This compiles every module and runs the tests. The README notes that you can add -Drelease to trigger obfuscation and/or uglification, and that -Dmaven.test.skip=true skips tests, with the caveat that you should not do so. Expect the first run to be slow, because it resolves the full dependency tree before compiling anything.

To produce the distributable archives rather than just install artifacts into the local repository, use the package goal:

```bash
mvn clean package
```

The README states the packaged results appear in the target/ sub-folders of assemblies/*. For the desktop client community edition, the expected output is assemblies/client/target/pdi-ce-*-SNAPSHOT.zip. Unzip that archive and you have the client distribution. If the file is not there, the assembly step failed or was skipped, and the console output is where you find out why.

For a quicker iteration loop while developing, you can run the unit tests alone:

```bash
mvn test
```

And to remote debug a single Java unit test, the README gives this sequence, with 5005 as the default port:

```bash
cd core
mvn test -Dtest=<<YourTest>> -Dmaven.surefire.debug
```

The integration tests are separate and are gated behind a flag:

```bash
mvn verify -DrunITs
```

A single integration test can be selected with -Dit.test=<<YourIT>>, and adding -Dmaven.failsafe.debug attaches the debugger on the same default port 5005.

## Writing tests against the PDI environment

The contributing section describes two JUnit ClassRules that exist specifically because PDI tests mutate global state. RestorePDIEnvironment is intended for core tests and RestorePDIEngineEnvironment for engine tests, and both live under src/test/java in their respective modules. The README shows the intended usage pattern:

```java
public class MyTest {
  @ClassRule public static RestorePDIEnvironment env = new RestorePDIEnvironment();
  #setUp()...
  @Test public void testSomething() { 
    assertTrue( myMethod() ); 
  }
}
```

The snippet as printed in the README is not compilable Java; #setUp() is shorthand for the setup method, not valid syntax. Treat it as an illustration of where the ClassRule goes rather than as a file to copy. The real signal is that the project expects tests to restore environment state after themselves, which is a common pattern in codebases where a static engine configuration is shared across test classes. If you contribute a plugin or an engine change, following that convention is what keeps the suite from leaking state between tests.

## Where Pentaho Data Integration is the wrong choice

The most concrete limitation is distribution. The README documents how to build the project and how to package it, but it does not document a download location for a prebuilt binary, and the repository homepage field is empty. The only release listed is 5.2.0.2-C-185-R, a customer patch from 2015. Anyone expecting to pull a current installer from this repository will not find one described here.

Second, the build chain is opinionated. Java JDK 11 and Maven 3 or later are hard prerequisites, and a settings.xml from a different repository must be installed before the build resolves dependencies. Teams standardised on Gradle, or running an older JDK, are outside the supported path. The -Drelease flag also means the artifact you build locally may differ from a release artifact in ways that affect debugging, because obfuscation is applied.

Third, this is not a lightweight library. The module list includes a full user interface, a database dialog layer, and a plugin system. If your task is to move a few thousand rows between two systems on a schedule, the assembly, build and test overhead is disproportionate. A script and a scheduler will do the same job with less to maintain.

Finally, the repository is not the place to look for operational documentation. The README directs questions to the Hitachi Vantara community forum, and the only operational artifact visible at the root is CarteAPIDocumentation.md plus a JMeter test plan. Deployment topology, clustering behaviour and upgrade procedures are not covered in the README.

## How this differs from a hosted or commercial ETL platform

The natural comparison is a commercial data integration platform, and the difference is not primarily feature count. It is where the logic lives and who operates it. In Pentaho Data Integration, transformations and jobs are artifacts produced by the desktop client and executed by the engine, and the repository gives you the source for both. You can read the engine code, patch a plugin, and run the whole thing on your own hardware. The Carte API documentation at the repository root suggests the server exposes an API for remote execution, which is the hook a self-hosted deployment would use.

A hosted platform inverts that. You get a managed control plane, a browser interface and a support contract, and you give up the ability to inspect or modify the execution engine. For a team without Java engineers, that trade is usually correct. For a team that needs to embed transformation logic inside an existing Java service, or that must run ETL inside a network where no external control plane is permitted, the source build is the only option that fits.

The honest framing is that Pentaho Data Integration is not competing on ease of first use. It is competing on the ability to own the stack. That is a real advantage for some organisations and a pure cost for others.

## Maintenance, licensing and what to check before adopting

The repository is not archived, and the last push was on 2026-09-21, so the source tree is receiving commits. That is a statement about commit activity, not about release cadence. The only release listed is the 2015 customer patch 5.2.0.2-C-185-R, and the README describes snapshot artifacts (pdi-ce-*-SNAPSHOT.zip), which means the documented build output is a snapshot rather than a versioned release. If your process requires pinned, versioned artifacts, that gap is the first thing to resolve with the maintainers.

The licence field is NOASSERTION, and the repository root contains a LICENSE.TXT. NOASSERTION means the automated licence detection could not classify the file, not that the project is unlicensed. Read LICENSE.TXT directly, and if you intend to redistribute the built artifacts or embed the engine in a product, have someone qualified review it. Nothing in the README states the terms under which the community edition may be redistributed.

Upgrade cost is dominated by the build chain rather than by API churn, because the README pins Java JDK 11 and Maven 3. Moving to a newer JDK means verifying that the whole module set, including the plugin modules and the assembly steps, still builds. The checkstyle check (mvn checkstyle:check) and the two environment ClassRules are the guardrails the project itself relies on, so running them is the cheapest way to find out whether a change breaks the build conventions.

## Conclusion

Adopt Pentaho Data Integration if you need a Java-based ETL suite whose transformation logic lives in files you can diff and version, and you are prepared to build it yourself, since the README documents no prebuilt binary. Do not adopt it if you want a hosted service or a vendor support contract, because the repository covers the source build only. Before committing, verify that a distribution ZIP appears under assemblies/client/target/ after mvn clean package, and confirm that your team can maintain a Java 11 plus Maven 3 build chain.

## FAQ

### What is Pentaho Kettle?

Pentaho Data Integration, also known as Kettle, is a Java ETL suite. The repository is a Maven multi-module build containing a core implementation, a PDI engine, a user interface, a database dialog, engine extensions, core plugins and an assemblies module that produces the distribution archive.

### Is Pentaho Kettle open source?

The source code is published in a public repository under the pentaho organisation and is not archived. The licence field is reported as NOASSERTION and a LICENSE.TXT sits at the repository root, so the exact terms need to be read from that file rather than assumed.

### Is Pentaho Kettle free?

The repository covers the source and its build instructions, and does not state pricing for any edition. It also does not describe a prebuilt binary download, so the question of cost is not answerable from what the README documents.

### What are the alternatives to Pentaho Kettle?

The README does not name or compare any competing ETL tool. It does describe the design trade-off: this project gives you the source for the engine and the client so you can self-host and modify them, in contrast to a managed platform where you cannot inspect the execution engine.

## Sources

- [Issues](https://github.com/pentaho/pentaho-kettle/issues)
- [pentaho/pentaho-kettle on GitHub](https://github.com/pentaho/pentaho-kettle)
- [README](https://github.com/pentaho/pentaho-kettle/blob/master/README.md)
- [Releases](https://github.com/pentaho/pentaho-kettle/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/pentaho-pentaho-kettle
