Tablesaw: a Java dataframe and visualization library for JVM data work
Java dataframe and visualization library
At a glance
- What is it?
- Tablesaw is an Apache-2.0 Java library for loading, cleaning, transforming and charting tabular data without leaving the JVM. It fits teams already writing Java who want dataframe-style operations and Plotly-backed charts, and it is a poor fit for anyone expecting pandas-scale ecosystem coverage.
- Who is it for?
- Adopt Tablesaw if your data pipeline already lives in Java and you want filtering, grouping and descriptive statistics without shelling out to Python. Skip it if your work depends on a wide ecosystem of statistical or machine learning packages, or if you need a stable 1.0 API surface.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 53 days ago.
- What is it written in?
- Mainly Java, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The gap Tablesaw fills for JVM developers
Java has never had a comfortable story for exploratory data work. You can read a CSV with a parser, hold rows in a List<Map<String, Object>>, and write loops to compute a mean, but the code grows quickly and the types get loose. Tablesaw's pitch is that a Java developer can load a file, filter it, group it and summarize it with a chain of method calls, and then hand the same table to a plotting wrapper. The README frames the audience plainly: "If you work with data in Java, it may save you time and effort."
The library also positions itself as a preparation step rather than a modelling tool. The README lists Smile, Tribuo, H20.ai and DL4J as libraries you can feed from a Tablesaw table. That framing is honest about scope. Tablesaw does the tabular plumbing; the modelling stays in whatever library you already chose. For a team whose services are written in Java and whose analysts keep asking for a small amount of statistics, that division of labour is reasonable. For a team doing heavy statistical modelling, it is a detour.
How Tablesaw represents a table and moves data through it
The core abstraction is a Table made of typed columns. Rather than one generic column type, Tablesaw carries column classes for the kinds of data it expects: numbers, strings, dates, booleans and so on. That matters because the descriptive statistics the README lists (mean, min, max, median, sum, product, standard deviation, variance, percentiles, geometric mean, skewness, kurtosis) only make sense on numeric columns, and the typed design lets the API expose them where they apply instead of failing at runtime on a string column.
The data flow is conventional. Readers pull from a source, whether that is a CSV file, a TSV file, a JSON document, an Excel workbook, an HTML table, a fixed width text file or an RDBMS, and the README notes those sources can be local or remote over http or S3. From there the operations are the ones you would expect from a dataframe: append or join tables, add and remove columns or rows, sort, group, filter, edit, transpose, run map and reduce operations, and handle missing values. Writers push the result back out to CSV, JSON, HTML or fixed width files.
The visualization layer is a wrapper around the Plotly JavaScript library, not a native Java renderer. That is a design decision with visible consequences: chart output depends on a JavaScript runtime being present in whatever environment renders it, which is why the README steers readers toward Jupyter notebooks through BeakerX or IJava rather than promising a Swing window. The repository layout reflects the same modularity, with separate top-level directories for core, excel, html, json, jsplot, arrow, beakerx and saw.
Installing tablesaw-core and running a first filter
Tablesaw is distributed as Maven artifacts, so the install step is a dependency entry. The README gives the coordinates and tells you to look up the version number in the release notes rather than hardcoding one. The most recent release listed is v0.44.4 from 2026-01-09.
<dependency>
<groupId>tech.tablesaw</groupId>
<artifactId>tablesaw-core</artifactId>
<version>VERSION_NUMBER_GOES_HERE</version>
</dependency>The README also lists optional modules you add only when you need them: tablesaw-beakerx for BeakerX, tablesaw-excel for Excel workbooks, tablesaw-html for HTML, tablesaw-json for JSON and tablesaw-jsplot for charts. Keeping them separate means a project that only reads CSV does not pull in the Excel or plotting dependencies.
For a first real use, the documented starting point is the getting started page at jtablesaw.github.io/tablesaw/gettingstarted, followed by the user guide. The README does not inline a complete load-and-query example, so the honest advice is to open the user guide rather than guess at API names. What the README does confirm is the shape of the work: read a file into a Table, then apply the transformation operations (sort, group, filter, edit) and the statistics operations to that table. If you are experimenting rather than building, the README recommends Jupyter through BeakerX or IJava, and it points to Gary Sharpe's tutorials for CSV and JSON importing as worked examples.
Where Tablesaw stops being the right tool
The clearest limitation is ecosystem depth. Tablesaw gives you descriptive statistics and data preparation. It does not give you regression, clustering or model training; the README explicitly routes that work to Smile, Tribuo, H20.ai or DL4J. If your task is statistical modelling rather than tabular wrangling, you are adding a dependency to reach a library you still have to add separately.
Parquet support is a second gap, and an instructive one. It lives in tablesaw-parquet, a project the README marks as "outside of this organization", hosted at github.com/tlabs-data. That means the format many data teams actually store their tables in is maintained on a different release cadence by different people. If Parquet is central to your pipeline, you are depending on two projects with two maintenance schedules, and the README's own phrasing makes that boundary explicit rather than hiding it.
The visualization wrapper is a third constraint. Because charts go through Plotly's JavaScript library, the natural home for them is a notebook, and the README's recommendations all point that way: BeakerX, IJava, Google Colab. A backend service that needs to emit a chart image is working against the grain of how the library is documented. And the version numbering is worth noting on its own: v0.44.4 signals pre-1.0, which in practice means the API can still move between releases.
Tablesaw against pandas and the JVM alternatives
The obvious comparison is pandas. Both give you a typed tabular container with filtering, grouping, joining and descriptive statistics, and both wrap a JavaScript charting library (Plotly in Tablesaw's case, matplotlib by default in pandas). The difference is not the operations, it is the surrounding world. pandas sits inside an ecosystem where scikit-learn, statsmodels and a large body of published notebooks assume you are already holding a dataframe. Tablesaw sits inside Java, where the equivalent modelling libraries exist but expect you to convert or adapt.
So the choice is usually made for you by where the rest of the code lives. If your ingestion service, your scheduled jobs and your deployment pipeline are Java, moving data out to Python for a group-by is friction, and Tablesaw removes it. If your team already writes Python, Tablesaw offers nothing pandas does not.
Within the JVM there is also the option of writing the aggregation by hand over a stream, or of using a general-purpose collections library. That works for one-off sums and counts. It stops working when you need joins, missing-value handling, grouping across multiple keys and a chart at the end, which is precisely the set of operations the README enumerates. Tablesaw's value is that those operations are named, tested and documented in one place rather than reinvented per project.
Release cadence, licence and what an upgrade costs
The repository is not archived, and the last push was on 2026-08-08, roughly five weeks before this writing. The three most recent releases, v0.44.2, v0.44.3 and v0.44.4, all landed on 2026-01-09, which suggests a batch of patch releases rather than a steady stream. The README directs readers to the release notes for version numbers, which is the right place to check before pinning a dependency.
Upgrade cost is mostly a function of the 0.x version line. Because the project has not reached 1.0, there is no compatibility promise to lean on, and a minor version bump can carry API changes. The modular split helps here: if you use only tablesaw-core, an upgrade touches one artifact, and the optional modules (excel, html, json, jsplot, beakerx) can move independently. Teams that depend on Parquet carry the extra cost of tracking a separate repository.
On licensing, Tablesaw is Apache-2.0, and the repository carries a LICENSE.txt at the top level. Apache-2.0 is a permissive licence that permits commercial use and modification, and it includes an explicit patent grant. It is not a copyleft licence, so it does not oblige you to publish your own source. That is the general shape of the licence; whether it satisfies your organisation's policy is a question for your legal team, not something to settle from a README badge.
Editorial conclusion
Adopt Tablesaw if your data pipeline already lives in Java and you want filtering, grouping and descriptive statistics without shelling out to Python. Skip it if your work depends on a wide ecosystem of statistical or machine learning packages, or if you need a stable 1.0 API surface. Before committing, verify the current release version in the release notes, confirm which optional modules (tablesaw-json, tablesaw-excel, tablesaw-jsplot) your use case needs, and check whether the Parquet path through the external tablesaw-parquet project is maintained to your satisfaction.
Frequently asked questions
What is Tablesaw used for?
Tablesaw is a Java dataframe and visualization library. The README describes it as supporting loading, cleaning, transforming, filtering and summarizing data, plus descriptive statistics, and as a way to prepare data for machine learning libraries such as Smile, Tribuo, H20.ai and DL4J.
How do I use Tablesaw in a Java project?
Add the tablesaw-core artifact as a Maven dependency, taking the version number from the release notes, then add optional modules such as tablesaw-json, tablesaw-excel or tablesaw-jsplot only if you need those formats or charts. The README points to the getting started page and the user guide for the API itself.
How do I install Tablesaw?
There is no installer; Tablesaw is a set of Maven artifacts. The README gives the tech.tablesaw groupId and tablesaw-core artifactId, and lists the supporting artifacts tablesaw-beakerx, tablesaw-excel, tablesaw-html, tablesaw-json and tablesaw-jsplot.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/jtablesaw-tablesaw)