# dplyr: the grammar of data manipulation in R, and where it stops being enough

> dplyr gives R a small set of verbs for filtering, selecting, mutating, arranging and summarising data frames, and the same verbs also drive Arrow, SQL, data.table, DuckDB and Spark backends. Here is how it installs, how the verbs compose, and which workloads belong somewhere else.

**tidyverse/dplyr** — dplyr: A grammar of data manipulation

- Repository: https://github.com/tidyverse/dplyr
- Website: https://dplyr.tidyverse.org/
- Stars: 5,072 · Forks: 2,110
- Language: R
- License: NOASSERTION
- Published: 2026-09-23 · Updated: 2026-09-23 · Language: en
- Canonical page: https://hysenlabs.com/projects/tidyverse-dplyr

## The problem dplyr solves is verb sprawl, not speed

Base R can filter, sort and aggregate, but each task has its own calling convention. Subsetting uses bracket indexing with a comma. Sorting uses order() inside brackets. Aggregation uses aggregate() or tapply(), each with a different argument order. The cost is not that any one call is hard, it is that the reader has to re-derive the intent every time.

dplyr replaces that with five verbs the README lists explicitly: mutate() adds variables that are functions of existing ones, select() picks variables by name, filter() picks cases by value, summarise() reduces many values to one summary, and arrange() changes row order. group_by() turns any of them into a per-group operation. That is the whole grammar, and it is small on purpose.

The audience is anyone doing exploratory or pipeline work on tabular data in R: analysts who move between datasets weekly, and package authors who want a stable interface over a storage engine. The README points new users at the data transformation chapter of R for Data Science rather than at the reference documentation, which tells you the intended entry point is a book, not a function index.

## How the verbs compose, and what group_by actually changes

A dplyr pipeline is a sequence of verbs applied left to right, each returning a data frame or tibble that the next verb consumes. The README's own examples show the shape. filter(species == "Droid") on the starwars dataset returns a tibble with the matching rows and all fourteen columns. select(name, ends_with("color")) returns four columns. mutate(name, bmi = mass / ((height / 100)^2)) then select(name:mass, bmi) creates a derived column and narrows the output in the same pipeline.

group_by() is the piece that changes semantics rather than shape. After group_by(species), summarise() produces one row per species instead of one row overall, and n() counts rows within each group. The README example chains summarise() straight into filter(n > 1, mass > 50), which is worth pausing on: the filter runs against the summarised output, not the original rows, so it removes whole species rather than individual characters. Reading a pipeline in the wrong order is the most common way to get a plausible but wrong answer.

The backends are the second half of the architecture. dplyr itself operates on data frames and tibbles. For anything else, the README names five translation layers: arrow for larger-than-memory data including remote cloud storage via the Apache Arrow C++ engine Acero, dbplyr for relational databases which translates dplyr code to SQL, dtplyr for large in-memory data which translates to data.table code, duckplyr for large in-memory data which translates to DuckDB queries with an automatic R fallback when translation is not possible, and sparklyr for data in Apache Spark. Your verbs stay the same; the execution engine does not.

## Installing dplyr and running a first grouped summary

dplyr is on CRAN. The README gives two install paths: the whole tidyverse, or dplyr alone. Install the single package if you do not want the rest of the tidyverse attached to your session.

```r
# The easiest way to get dplyr is to install the whole tidyverse:
install.packages("tidyverse")

# Alternatively, install just dplyr:
install.packages("dplyr")
```

For a bug fix or a feature that only exists on the development branch, the README documents installing from GitHub through pak.

```r
# install.packages("pak")
pak::pak("tidyverse/dplyr")
```

Once installed, load it and run a grouped summary. This is the README's own example, and it exercises four verbs in one pipeline.

```r
library(dplyr)

starwars |>
  group_by(species) |>
  summarise(
    n = n(),
    mass = mean(mass, na.rm = TRUE)
  ) |>
  filter(
    n > 1,
    mass > 50
  )
```

The output is a tibble with three columns, species, n and mass, and one row per species that has more than one character and an average mass above 50. The README's printed result shows nine such rows, including Droid, Gungan, Human, Kaminoan and Mirialan. If you see more rows than that, check whether na.rm is set on the mean, because a single missing mass value propagates through the whole group otherwise.

The README also links a printable data transformation cheat sheet, which is the fastest way to see the verb list on one page.

## Where dplyr is the wrong tool

The verbs are a front end. When you use arrow, dbplyr, dtplyr, duckplyr or sparklyr, the code you write is translated, and translation is not total. duckplyr's description in the README is the most explicit about this: it has an automatic R fallback when translation is not possible. The other backends do not advertise that behaviour in the README, which means an operation that cannot be translated is something you discover at runtime, not at review time.

That is the first real limitation. A pipeline that runs instantly on a tibble can fail or behave differently once the same verbs are pointed at a database through dbplyr, because SQL has different semantics for ordering, null handling and types. Nothing in the README promises identical results across backends.

The second limitation is memory. dplyr on a data frame holds the data in R's memory. The README's answer to larger-than-memory data is a separate package, not a dplyr feature, so if your dataset does not fit, you have not chosen dplyr, you have chosen arrow or DuckDB and are using dplyr as the syntax.

The third is that dplyr is not a data cleaning toolkit. It has no verb in the documented core set for deduplication, fuzzy matching, string parsing or date normalisation. Those live in other tidyverse packages. If your task is mostly reshaping messy text, the grammar gives you less than you would expect from the package's reputation.

## dplyr against base R and data.table

The honest alternative inside R is base R itself, and the difference is not capability but readability. Base R's bracket syntax does everything dplyr's verbs do. What it does not do is name the operation. A reader of df[df$species == "Droid", ] has to parse the condition to know it is a filter. A reader of filter(species == "Droid") does not. The trade is that base R has no dependency beyond R itself, while dplyr pulls in the tidyverse's shared infrastructure.

The more interesting comparison is data.table, and the README itself makes it concrete: dtplyr exists to translate dplyr code into high performance data.table code for large in-memory datasets. That is an unusual admission. Rather than argue that dplyr is fast enough, the project ships a translation layer to a competing package's engine. If your bottleneck is in-memory speed on tens of millions of rows, the documented path is to write dplyr and let dtplyr emit data.table, or to write data.table directly and skip the translation step. The second option removes a layer that can fail to translate.

For database work the comparison is less about alternatives and more about whether dbplyr's SQL translation covers your query. For out-of-core work, arrow and duckplyr are different engines behind the same verbs, and the choice between them is a storage decision, not a dplyr decision.

## Maintenance, releases and what the licence leaves open

The repository is not archived and the last push was on 2026-06-02, which is under four months before today's date. The release history is uneven in a way worth noting: v1.1.4 shipped on 2023-11-17, then v1.2.0 on 2026-02-04 and v1.2.1 on 2026-04-03. That is a gap of more than two years between minor releases, followed by two releases eight weeks apart. The README does not explain the gap and the release notes are not part of the repository's published documentation, so treat the cadence as something to watch rather than something to predict.

Upgrade cost is the practical question. dplyr has a large reverse-dependency surface, which is why the repository carries a revdep/ directory: the maintainers check downstream packages against changes before release. If you maintain a package that depends on dplyr, that directory is the signal that breaking changes are taken seriously, and it is also the reason a dplyr upgrade can ripple through your own dependency tree.

The licence field reports NOASSERTION, which means the automated classifier could not map the repository's LICENSE and LICENSE.md files to a recognised identifier. The README does not restate the licence terms. Read LICENSE.md directly before you redistribute dplyr or bundle it into a product, and do not treat the NOASSERTION value as either permissive or restrictive. This is a description of what the metadata says, not legal advice.

## Conclusion

Adopt dplyr if your data fits in memory, or if it does not and one of the documented backends (arrow, dbplyr, dtplyr, duckplyr, sparklyr) covers your storage. Do not adopt it as a general-purpose engine for out-of-core work on a backend that has no translation, because the verbs are only as fast as what sits underneath them. Before committing, verify three things: that install.packages("dplyr") resolves on your R version, that your grouping logic survives a group_by() followed by summarise() on a real subset of your data, and that your chosen backend can translate the specific verbs you rely on. The package documents no rollback path for a broken upgrade, so pin the version in your project's dependency record rather than tracking CRAN blindly.

## FAQ

### What does dplyr stand for?

The README does not expand the name. It describes dplyr only as a grammar of data manipulation, and the package's own documentation does not give an acronym or full form anywhere in the repository.

### What does %>% do in dplyr?

The README does not use %>% in any of its examples. Every pipeline shown uses the native R pipe, |>, which passes the left-hand value into the first argument of the right-hand call.

### What are the uses of dplyr?

It provides a consistent set of verbs for the most common data manipulation tasks: mutate() adds variables, select() picks variables by name, filter() picks cases by value, summarise() reduces many values to one summary, and arrange() changes row order. group_by() makes any of those operations run per group.

### Is dplyr the same as tidyverse?

No. dplyr is one package within the tidyverse, and the README offers both install.packages("tidyverse") for the whole collection and install.packages("dplyr") for dplyr alone.

### How do I install dplyr in R?

Run install.packages("dplyr") for the CRAN release, or install.packages("tidyverse") to get the whole collection. For the development version the README documents pak::pak("tidyverse/dplyr") after installing pak.

### How do I use dplyr filter in R?

filter() picks cases based on their values. The README example starwars |> filter(species == "Droid") returns a tibble containing only the rows whose species column equals "Droid", with all columns preserved.

## Sources

- [Issues](https://github.com/tidyverse/dplyr/issues)
- [Project website](https://dplyr.tidyverse.org/)
- [README](https://github.com/tidyverse/dplyr/blob/main/README.md)
- [Releases](https://github.com/tidyverse/dplyr/releases)
- [tidyverse/dplyr on GitHub](https://github.com/tidyverse/dplyr)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/tidyverse-dplyr
