# tech.ml.dataset: Functional Tabular Data for Clojure on the JVM

> tech.ml.dataset (TMD) is a Clojure library for tabular data processing that brings columnar storage and immutable datasets to the JVM. It targets data engineers who want the ergonomics of Pandas or R's data.table without leaving the Clojure ecosystem.

**techascent/tech.ml.dataset** — A Clojure high performance data processing system

- Repository: https://github.com/techascent/tech.ml.dataset
- Stars: 763 · Forks: 34
- Language: Clojure
- License: EPL-1.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/techascent-tech-ml-dataset

## Tabular Data on the JVM Without Python's State Problems

tech.ml.dataset (TMD) targets Clojure engineers who process structured, tabular data and want to stay on the JVM without switching to Python or R. The core design premise is that datasets are immutable values, not mutable objects. In Python, a Pandas DataFrame is modified in place by default. In TMD, every transformation returns a new dataset, leaving the original untouched.

That functional approach matters for teams building data pipelines where correctness is hard to verify. State bugs in Pandas pipelines are a common source of incorrect results, particularly in notebook environments where cells are executed out of order. TMD sidesteps that class of problem by design, and it composes naturally with Clojure's own sequence and transducer machinery. The README describes this as making datasets easier to reason about, and it shows in code that mixes dataset transforms with standard Clojure operations.

The library is aimed at pragmatic, data-intensive work on the JVM. The README describes it as similar to Python's Pandas or R's data.table, but built from the ground up around Clojure idioms rather than ported or wrapped from another language.

## Columnar Storage, Primitive Arrays, and String Tables

The README highlights several mechanisms that reduce memory footprint compared to row-oriented storage. Columnar layout stores all values of a single column together in memory, which lets numeric operations skip over irrelevant fields and keeps CPU caches warm. Primitive arrays avoid the overhead of boxed Java objects: a column of longs in TMD is a raw long array, not an array of java.lang.Long instances.

Packed datetime types encode timestamps as integers rather than object references, which can reduce per-value memory significantly for date-heavy tables. The README links to an external benchmark showing the memory impact of this on real data, though it does not state exact compression ratios for arbitrary inputs.

TMD also uses string tables, which replace repeated string values with integer indices. A column that holds country codes or category labels stores each unique string once and references it by index everywhere else. This can substantially reduce memory for columns with low cardinality.

The underlying numeric work is delegated to tech.v3.datatype, the dtype-next project from the same author. That library provides the array and primitive-operation infrastructure. TMD sits on top of it and handles dataset-level concerns: joining, grouping, sorting, and loading from files.

## Installing tech.ml.dataset and Verifying the Setup

The library is distributed through Clojars, the Clojure community artifact host. The README directs you to the Clojars page (clojars.org/techascent/tech.ml.dataset) for exact dependency coordinates to paste into deps.edn, project.clj, or whichever build tool you use. The repository itself contains a deps.edn file, indicating the project uses the Clojure CLI toolchain.

After adding the dependency, the README provides a verification step. Require the main namespace and call ->>dataset on any Clojure sequence of maps:

```clojure
user> (require 'tech.v3.dataset)
nil
user> (->> (System/getProperties)
           (map (fn [[k v]] {:k k :v (apply str (take 40 (str v)))}))
           (tech.v3.dataset/->>dataset {:dataset-name "My Truncated System Properties"}))
```

The ->>dataset function converts a sequence of maps into a typed, column-oriented dataset. The output is a tabular view of the JVM's system properties, with column names derived from the map keys and types inferred from the values. Seeing that table print confirms the library loaded correctly.

The documentation site at techascent.github.io/tech.ml.dataset covers a Getting Started guide, a long-form Walkthrough with real data examples, and a Quick Reference for frequently used functions. A Java API with Javadoc is also provided for teams that need to call TMD from Java code without writing Clojure; a sample program lives at java_test/java/jtest/TMDDemo.java in the repository.

## Where tech.ml.dataset Falls Short

The README does not document a streaming or out-of-core processing API. TMD is designed for in-memory work. Very large tables will consume JVM heap proportional to their size even with the columnar storage savings. Teams dealing with datasets that exceed available memory have no documented path within TMD itself.

The library has no GitHub releases. Version history lives in CHANGELOG.md and the Clojars artifact list. There are no formal release tags with associated notes, which can complicate reproducible builds if Clojars metadata changes or becomes unavailable.

TMD is a Clojure library first. If your team works primarily in Python, Java, or Scala, the idiomatic data-processing libraries for those languages are more appropriate. The Java API exists, and the repository documents it with Javadoc, but the primary interface is Clojure. A Java engineer who wants to call TMD from pure Java will work against the grain of the library's design.

## tech.ml.dataset Compared to Python's Pandas

Python's Pandas is the most widely used tabular data library in data science. Both libraries load tabular files, group and aggregate rows, and join tables on keys. The differences are structural.

Pandas runs in a Python process and integrates naturally with NumPy, scikit-learn, and Jupyter. TMD runs on the JVM and integrates naturally with Clojure libraries, Java code, and the JVM ecosystem. The choice between them is primarily a question of which runtime your team already operates in.

Pandas DataFrames are mutable by default. A function that receives a DataFrame can modify it and the caller may not notice. TMD datasets are immutable; a function that transforms a dataset returns a new one. That difference has practical consequences for pipeline correctness and testing.

The README mentions independent benchmarks that compare TMD's speed, without describing those benchmarks in detail or stating specific margins. For teams within the JVM ecosystem, the README also notes that tablecloth provides a higher-level API built on top of TMD with additional features, and that tech.ml provides machine learning pathways for regression and classification. There is also tmducken, a binding to a high-performance in-process SQL database that works alongside TMD.

## Maintenance Calendar and EPL-1.0 Licence

The last push to the repository was on 2026-09-04. The project is not archived. The repository includes a CONTRIBUTORS.md file, a CHANGELOG.md tracking changes, and active community channels: a Zulip stream at clojurians.zulipchat.com in the tech.ml.dataset.dev section, and a Slack channel in the Clojurians workspace at the data science channel.

The library is distributed under the Eclipse Public License 1.0 (EPL-1.0). EPL-1.0 is a weak copyleft licence. If you modify the library's source code and distribute the modified version, EPL requires you to publish those modifications under the same licence. Using the library as a dependency in an application without modifying it does not trigger that obligation. Teams with strict legal review requirements should verify EPL-1.0 compatibility with their own licence policies before adopting.

## Conclusion

Clojure data engineers building JVM data pipelines get real value from tech.ml.dataset's immutable dataset model and columnar memory layout. The library has no out-of-core processing API; if working datasets exceed JVM heap, there is no documented workaround within TMD itself. Python teams have no reason to switch runtimes. The Getting Started guide at techascent.github.io/tech.ml.dataset/000-getting-started.html walks through the first dataset operations.

## FAQ

### Does tech.ml.dataset support calling from Java as well as Clojure?

The README describes a Java API with Javadoc documentation and a sample program at java_test/java/jtest/TMDDemo.java in the repository. The primary interface is Clojure, but a Java API exists for teams that need to call TMD from Java code.

### What provides the numeric operations underneath tech.ml.dataset?

The README identifies tech.v3.datatype, also known as the dtype-next project, as the underlying numeric subsystem that tech.ml.dataset builds on. It provides the array and primitive-operation infrastructure.

### Is there a higher-level API for tech.ml.dataset?

The README describes tablecloth from the scicloj project as an alternative API with additional features built on top of tech.ml.dataset. It is listed as the recommended starting point for users who want more abstractions.

## Sources

- [Issues](https://github.com/techascent/tech.ml.dataset/issues)
- [License: EPL-1.0](https://github.com/techascent/tech.ml.dataset/blob/master/LICENSE)
- [README](https://github.com/techascent/tech.ml.dataset/blob/master/README.md)
- [techascent/tech.ml.dataset on GitHub](https://github.com/techascent/tech.ml.dataset)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/techascent-tech-ml-dataset
