tech.ml.dataset: Functional Tabular Data for Clojure on the JVM
A Clojure high performance data processing system
At a glance
- What is it?
- tech.ml.dataset (TMD) is a Clojure library for tabular data processing that brings columnar storage and immutable datasets to the JVM. It targets data engineers who want the ergonomics of Pandas or R's data.table without leaving the Clojure ecosystem.
- Who is it for?
- Clojure data engineers building JVM data pipelines get real value from tech.ml.dataset's immutable dataset model and columnar memory layout. The library has no out-of-core processing API; if working datasets exceed JVM heap, there is no documented workaround within TMD itself.
- Can I use it commercially?
- Yes, with conditions. EPL-1.0 is a weak copyleft licence: you can use it inside commercial and closed-source software, but if you distribute changes to its own files, you must publish those changes under the same licence.
- Is it still maintained?
- Yes. The repository last received commits 27 days ago.
- What is it written in?
- Mainly Clojure, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 26, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Tabular Data on the JVM Without Python's State Problems
tech.ml.dataset (TMD) targets Clojure engineers who process structured, tabular data and want to stay on the JVM without switching to Python or R. The core design premise is that datasets are immutable values, not mutable objects. In Python, a Pandas DataFrame is modified in place by default. In TMD, every transformation returns a new dataset, leaving the original untouched.
That functional approach matters for teams building data pipelines where correctness is hard to verify. State bugs in Pandas pipelines are a common source of incorrect results, particularly in notebook environments where cells are executed out of order. TMD sidesteps that class of problem by design, and it composes naturally with Clojure's own sequence and transducer machinery. The README describes this as making datasets easier to reason about, and it shows in code that mixes dataset transforms with standard Clojure operations.
The library is aimed at pragmatic, data-intensive work on the JVM. The README describes it as similar to Python's Pandas or R's data.table, but built from the ground up around Clojure idioms rather than ported or wrapped from another language.
Columnar Storage, Primitive Arrays, and String Tables
The README highlights several mechanisms that reduce memory footprint compared to row-oriented storage. Columnar layout stores all values of a single column together in memory, which lets numeric operations skip over irrelevant fields and keeps CPU caches warm. Primitive arrays avoid the overhead of boxed Java objects: a column of longs in TMD is a raw long array, not an array of java.lang.Long instances.
Packed datetime types encode timestamps as integers rather than object references, which can reduce per-value memory significantly for date-heavy tables. The README links to an external benchmark showing the memory impact of this on real data, though it does not state exact compression ratios for arbitrary inputs.
TMD also uses string tables, which replace repeated string values with integer indices. A column that holds country codes or category labels stores each unique string once and references it by index everywhere else. This can substantially reduce memory for columns with low cardinality.
The underlying numeric work is delegated to tech.v3.datatype, the dtype-next project from the same author. That library provides the array and primitive-operation infrastructure. TMD sits on top of it and handles dataset-level concerns: joining, grouping, sorting, and loading from files.
Installing tech.ml.dataset and Verifying the Setup
The library is distributed through Clojars, the Clojure community artifact host. The README directs you to the Clojars page (clojars.org/techascent/tech.ml.dataset) for exact dependency coordinates to paste into deps.edn, project.clj, or whichever build tool you use. The repository itself contains a deps.edn file, indicating the project uses the Clojure CLI toolchain.
After adding the dependency, the README provides a verification step. Require the main namespace and call ->>dataset on any Clojure sequence of maps:
user> (require 'tech.v3.dataset)
nil
user> (->> (System/getProperties)
(map (fn [[k v]] {:k k :v (apply str (take 40 (str v)))}))
(tech.v3.dataset/->>dataset {:dataset-name "My Truncated System Properties"}))The ->>dataset function converts a sequence of maps into a typed, column-oriented dataset. The output is a tabular view of the JVM's system properties, with column names derived from the map keys and types inferred from the values. Seeing that table print confirms the library loaded correctly.
The documentation site at techascent.github.io/tech.ml.dataset covers a Getting Started guide, a long-form Walkthrough with real data examples, and a Quick Reference for frequently used functions. A Java API with Javadoc is also provided for teams that need to call TMD from Java code without writing Clojure; a sample program lives at java_test/java/jtest/TMDDemo.java in the repository.
Where tech.ml.dataset Falls Short
The README does not document a streaming or out-of-core processing API. TMD is designed for in-memory work. Very large tables will consume JVM heap proportional to their size even with the columnar storage savings. Teams dealing with datasets that exceed available memory have no documented path within TMD itself.
The library has no GitHub releases. Version history lives in CHANGELOG.md and the Clojars artifact list. There are no formal release tags with associated notes, which can complicate reproducible builds if Clojars metadata changes or becomes unavailable.
TMD is a Clojure library first. If your team works primarily in Python, Java, or Scala, the idiomatic data-processing libraries for those languages are more appropriate. The Java API exists, and the repository documents it with Javadoc, but the primary interface is Clojure. A Java engineer who wants to call TMD from pure Java will work against the grain of the library's design.
tech.ml.dataset Compared to Python's Pandas
Python's Pandas is the most widely used tabular data library in data science. Both libraries load tabular files, group and aggregate rows, and join tables on keys. The differences are structural.
Pandas runs in a Python process and integrates naturally with NumPy, scikit-learn, and Jupyter. TMD runs on the JVM and integrates naturally with Clojure libraries, Java code, and the JVM ecosystem. The choice between them is primarily a question of which runtime your team already operates in.
Pandas DataFrames are mutable by default. A function that receives a DataFrame can modify it and the caller may not notice. TMD datasets are immutable; a function that transforms a dataset returns a new one. That difference has practical consequences for pipeline correctness and testing.
The README mentions independent benchmarks that compare TMD's speed, without describing those benchmarks in detail or stating specific margins. For teams within the JVM ecosystem, the README also notes that tablecloth provides a higher-level API built on top of TMD with additional features, and that tech.ml provides machine learning pathways for regression and classification. There is also tmducken, a binding to a high-performance in-process SQL database that works alongside TMD.
Maintenance Calendar and EPL-1.0 Licence
The last push to the repository was on 2026-09-04. The project is not archived. The repository includes a CONTRIBUTORS.md file, a CHANGELOG.md tracking changes, and active community channels: a Zulip stream at clojurians.zulipchat.com in the tech.ml.dataset.dev section, and a Slack channel in the Clojurians workspace at the data science channel.
The library is distributed under the Eclipse Public License 1.0 (EPL-1.0). EPL-1.0 is a weak copyleft licence. If you modify the library's source code and distribute the modified version, EPL requires you to publish those modifications under the same licence. Using the library as a dependency in an application without modifying it does not trigger that obligation. Teams with strict legal review requirements should verify EPL-1.0 compatibility with their own licence policies before adopting.
Editorial conclusion
Clojure data engineers building JVM data pipelines get real value from tech.ml.dataset's immutable dataset model and columnar memory layout. The library has no out-of-core processing API; if working datasets exceed JVM heap, there is no documented workaround within TMD itself. Python teams have no reason to switch runtimes. The Getting Started guide at techascent.github.io/tech.ml.dataset/000-getting-started.html walks through the first dataset operations.
Frequently asked questions
Does tech.ml.dataset support calling from Java as well as Clojure?
The README describes a Java API with Javadoc documentation and a sample program at java_test/java/jtest/TMDDemo.java in the repository. The primary interface is Clojure, but a Java API exists for teams that need to call TMD from Java code.
What provides the numeric operations underneath tech.ml.dataset?
The README identifies tech.v3.datatype, also known as the dtype-next project, as the underlying numeric subsystem that tech.ml.dataset builds on. It provides the array and primitive-operation infrastructure.
Is there a higher-level API for tech.ml.dataset?
The README describes tablecloth from the scicloj project as an alternative API with additional features built on top of tech.ml.dataset. It is listed as the recommended starting point for users who want more abstractions.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/techascent-tech-ml-dataset)