# data.table: R's data.frame replacement with a bracket that does SQL

> data.table gives R a high-performance version of data.frame with indexed joins, update by reference and a delimited file reader called fread. It is for R users whose data has outgrown base R, and the bracket syntax is the price of admission.

**Rdatatable/data.table** — R's data.table package extends data.frame:

- Repository: https://github.com/Rdatatable/data.table
- Website: http://r-datatable.com
- Stars: 3,921 · Forks: 1,054
- Language: R
- License: MPL-2.0
- Published: 2026-09-23 · Updated: 2026-09-23 · Language: en
- Canonical page: https://hysenlabs.com/projects/rdatatable-data-table

## What data.table does that base data.frame does not

Base R's data.frame copies on almost every modification. data.table is a reimplementation that adds an index, updates columns by reference so no copy is made, and internally parallelizes many common operations across CPU threads, according to the README. Those three properties are the reason it exists. The README describes it as "a high-performance version of base R's data.frame with syntax and feature enhancements for ease of use, convenience and programming speed."

The audience is R users working on one machine with data that no longer fits comfortably in base R's model: aggregations over hundreds of millions of rows, joins between tables where a merge() call would materialize intermediate copies, or repeated column assignment inside a loop. The README cites benchmarks on up to two billion rows and mentions 100GB in RAM, which sets the scale the project aims at. If your data is small, the syntax cost buys you little.

The project states it has no dependencies other than base R itself, and that its R dependency is kept "as old as possible for as long as possible", currently R 3.5.0 from 2018. For production environments where you cannot freely upgrade R, that constraint is the feature.

## DT[i, j, by]: the bracket as a query language

The core mechanism is a single bracket method. The README frames it as a mapping onto SQL: FROM[WHERE, SELECT, GROUP BY] corresponds to DT[i, j, by]. The i argument filters rows, j is any R expression (not a fixed list of columns), and by groups the result. Because j accepts arbitrary R expressions from any package, you are not limited to a backend's function set, and columns of type list are supported.

This differs from base R in two visible ways. Columns do not need the DT$ prefix, and the by argument is built into the bracket rather than requiring a separate call. The README's own example uses the iris dataset:

```r
library(data.table)
DT = as.data.table(iris)

# FROM[WHERE, SELECT, GROUP BY]
# DT  [i,     j,      by]

DT[Petal.Width > 1.0, mean(Petal.Length), by = Species]
#      Species       V1
#1: versicolor 4.362791
#2:  virginica 5.552000
```

The join machinery is the part that has no base R equivalent. The README lists ordered joins (rolling forwards, backwards, nearest, and limited staleness), overlapping range joins described as similar to IRanges::findOverlaps, non-equi joins using >, >=, <, <=, aggregate on join via by=.EACHI, and update on join. In base R, a rolling join or a range join is manual work with findInterval or a loop. Here it is an argument.

Reshaping is covered by dcast (pivot, wider, spread) and melt (unpivot, longer, gather). File input and output are covered by fread and fwrite, which the README calls fast and feature rich for delimited files.

## Installing data.table and running a first grouped aggregation

The README gives the CRAN install first. The development version has two paths: update_dev_pkg() installs it only if a newer build is available, and the explicit repos argument forces the install from the project's GitLab-hosted repository.

```r
install.packages("data.table")

# latest development version (only if newer available)
data.table::update_dev_pkg()

# latest development version (force install)
install.packages("data.table", repos="https://rdatatable.gitlab.io/data.table")
```

After either path, library(data.table) should load without warnings, and the README points to example(data.table) as a source of runnable examples. The vignette "Introduction to data.table" is linked from the Getting started section, and the wiki page Getting started is the other entry point named there.

For a first real use, the pattern in the README is a grouped aggregation over a filtered table. Convert a data.frame with as.data.table, then pass a logical condition as i, an aggregate expression as j, and a column name as by. The output is a data.table with the grouping column plus a default column name for the aggregate, as the README's iris output shows with V1. If you need to name it, the j expression is where the name goes, not a separate step.

If you build from source instead of CRAN, the repository ships a Makefile with the sequence the maintainers use: clean, build, install, test, check. The build target runs R CMD build with --no-build-vignettes, and the test target runs test.data.table() in a fresh R session.

```bash
make build
make install
make test
```

The Makefile hardcodes the tarball name data.table_1.18.99.tar.gz, so a version bump in DESCRIPTION means editing the Makefile targets that reference it.

## Where the by-reference model and the syntax bite

Update by reference is the sharpest edge in the package. The README lists "fast add/update/delete columns by reference by group using no copies at all" as a feature. The benefit is memory and speed. The cost is aliasing: because no copy is made, two names pointing at the same data.table see the same modification. That is the intended semantics, and it surprises people who expect R's usual copy-on-modify behavior. It is also why the README calls out "careful API lifecycle management" as a project value: changes to reference semantics have to be staged rather than dropped into a release.

The syntax itself is a real adoption barrier. DT[i, j, by] is compact and reads well once internalized, but j is evaluated in the frame of the table, so a variable named the same as a column resolves to the column. Debugging a bracket expression is harder than debugging a chain of function calls because the error points into the bracket rather than at a named intermediate. Teams that standardize on dplyr's verb-per-line style will find the density of a single bracket harder to review.

The wrong-tool case is a database. data.table is an in-memory, single-process library. The README's scale claims are about RAM, explicitly: 100GB in RAM. If your data exceeds memory, or if you need transactions, concurrent writers, or a query planner that spills to disk, data.table is not the layer for that. It also does not give you a server: there is no daemon, no connection protocol, no user permissions model. The README positions the package against base R and against database backends in one specific respect, that any R function from any package can be used in queries rather than a backend's subset, which is a statement about expressiveness, not about data size beyond memory.

## data.table against dplyr and base R

The closest comparison is dplyr. Both operate on in-memory R data, and both are widely used. The difference in approach is where the work happens. dplyr composes named verbs (filter, mutate, summarise, group_by) that pipe into each other, and it translates to SQL when the input is a database table, which data.table does not do. data.table puts the same operations into one bracket with positional arguments and keeps everything in R's own evaluation, with no translation layer. If you need one code path that runs against both a local frame and a remote database, dplyr's backend abstraction is the point of the design; data.table has no equivalent and does not claim one.

The other alternative is base R itself. data.frame plus merge plus aggregate plus reshape is what data.table replaces. Base R is always available and needs no install, and for small tables it is fine. The README's argument is that data.table keeps the data.frame interface while changing the performance characteristics: subsetting works the way you expect, but without the DT$ prefix and with by built in. That continuity is deliberate, and it is why as.data.table exists as a conversion rather than a separate class hierarchy.

One more distinction worth naming: data.table ships fread and fwrite, so it covers ingestion and export as well as transformation. If you are assembling a pipeline from separate packages for reading, transforming and writing, data.table is a single dependency that covers all three. The README's no-dependencies claim matters here, because it means adding data.table does not pull a tree of other packages into your DESCRIPTION.

## Maintenance, licence and the cost of staying current

The repository is not archived, and the last push was on 2026-09-18. The NEWS file at the top level is the changelog the README tells users to read, and there are two older files, NEWS.0.md and NEWS.1.md, alongside it, which suggests the project keeps historical release notes rather than discarding them.

The licence is MPL-2.0, the Mozilla Public License 2.0. MPL-2.0 is a file-level copyleft licence: modifications to files covered by the licence must be made available under the same licence, while larger works that combine the covered files with other code can be distributed under other terms. That is the general shape of the licence, not legal advice; if you are vendoring or modifying data.table source, read the LICENSE file in the repository and get your own counsel.

Upgrade cost is unusually low for a package this size, and that is a design choice rather than an accident. The README states the R dependency is kept "as old as possible for as long as possible", currently R 3.5.0 (2018), and that the project continuously tests against that version. It also states there are no dependencies other than base R. Together those two facts mean an upgrade cannot break you by dragging in a newer R or a transitive package. The remaining risk is behavioural: the README lists "careful API lifecycle management" as a project value, which implies deprecations are staged, but the README does not document a specific deprecation policy or a rollback procedure. Check the NEWS file for the version you are moving to before you upgrade a production script.

The governance structure is documented in GOVERNANCE.md, and the README notes the project uses a custom governance agreement and is fiscally sponsored by NumFOCUS. There is a CODEOWNERS file, so review responsibility is assigned by path. For a team deciding whether to depend on the project long term, those files are the ones to read, not the badge list.

## Conclusion

Adopt data.table if you work in R and your data has outgrown base data.frame or dplyr on a single machine: the by-reference column updates, ordered and non-equi joins, and fread cover ground that base R does not. Do not adopt it if your team cannot absorb the DT[i, j, by] syntax or if you need a database backend with transactions; data.table is in-memory and single-process. Before committing, check the DESCRIPTION file for the minimum R version (the README states R 3.5.0) and run test.data.table() after installing, since the bundled test suite is the only correctness check you can run locally.

## FAQ

### How do I install data.table in R?

Run install.packages("data.table") for the CRAN release. For the development version, the README gives data.table::update_dev_pkg(), which installs only if a newer build is available, or install.packages("data.table", repos="https://rdatatable.gitlab.io/data.table") to force the install.

### How do I install the data.table package in R from the development repository?

The README gives two options: data.table::update_dev_pkg() installs the latest development build only if a newer one is available, and install.packages("data.table", repos="https://rdatatable.gitlab.io/data.table") forces the install from the project's GitLab-hosted repository.

### How do I use data.table in R?

Use the bracket method DT[i, j, by], which the README maps onto SQL as FROM[WHERE, SELECT, GROUP BY]. The i argument filters rows, j is any R expression, and by groups the result. The README's example is DT[Petal.Width > 1.0, mean(Petal.Length), by = Species].

### What is data.table in R?

It is a package that provides a high-performance version of base R's data.frame with syntax and feature enhancements, according to the README. It adds fread and fwrite for delimited files, ordered and non-equi joins, and column updates by reference, with no dependencies other than base R.

### How do I create a data.table in R?

The README's usage example converts an existing data.frame with as.data.table(iris). Once converted, the bracket method applies to it the same way it applies to any data.table, and the README's Getting started section points to example(data.table) for more.

## Sources

- [Issues](https://github.com/Rdatatable/data.table/issues)
- [License: MPL-2.0](https://github.com/Rdatatable/data.table/blob/master/LICENSE)
- [Project website](http://r-datatable.com)
- [Rdatatable/data.table on GitHub](https://github.com/Rdatatable/data.table)
- [README](https://github.com/Rdatatable/data.table/blob/master/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/rdatatable-data-table
