# hadley/r4ds: the source repository behind R for Data Science

> The r4ds repository is the Quarto source of the R for Data Science book, not a package. It is useful if you want to read, translate or contribute to the text, and the second edition lives at r4ds.hadley.nz.

**hadley/r4ds** — R for data science: a book

- Repository: https://github.com/hadley/r4ds
- Website: https://r4ds.hadley.nz/
- Stars: 5,172 · Forks: 4,444
- Language: R
- License: NOASSERTION
- Published: 2026-09-23 · Updated: 2026-09-23 · Language: en
- Canonical page: https://hysenlabs.com/projects/hadley-r4ds

## What hadley/r4ds actually is, and who it is for

The repository holds the source of R for Data Science, and the README states plainly that "This repository contains the source of R for Data Science book." It is a book repository, not a library you install. The rendered second edition is published at r4ds.hadley.nz, and the homepage field points to the same address. Anyone who types r4ds hadley nz into a search box is looking for that site, and the repository is the machinery behind it rather than the reading experience itself.

The audience splits three ways. Readers who want the current text should use the website, because the repository is where the text is edited, not where it is served. Translators and instructors who want to fork a chapter and adapt it need the source, since each topic lives in its own file: data-visualize.qmd, data-transform.qmd, data-tidy.qmd, data-import.qmd, program.qmd, communicate.qmd and the rest. Contributors who spot a typo or an unclear paragraph need the repository too, because that is where a change can be proposed.

If you are looking for a package that performs the operations the book teaches, this is the wrong repository. The book teaches a workflow built on other packages; it does not ship one. The top-level entries are chapter files, diagram and screenshot directories, and build configuration, with no R/ directory and no exported functions.

## How the Quarto source is organised and rendered

Each chapter is a .qmd file at the repository root, and the naming is flat rather than nested: visualize.qmd, transform.qmd, import.qmd, strings.qmd, regexps.qmd, factors.qmd, datetimes.qmd, missing-values.qmd, joins.qmd, iteration.qmd, functions.qmd, quarto.qmd and others sit side by side. The README says the book is built using Quarto, and _quarto.yml is the configuration that ties the chapters into one document. A _freeze/ directory is present, which is the Quarto mechanism for caching computation so that a rebuild does not re-execute every code chunk.

Supporting material is separated by kind. diagrams/ holds drawing sources, screenshots/ holds interface captures, images/ holds other assets, and data/ holds datasets the chapters use. Two files at the root, _common.R and figures.R, are shared R code for the build, and r4ds.scss is the stylesheet. contributors.R and contributors.csv generate the contributor listing, and issues.json is present at the root as well.

The README documents an image pipeline in unusual detail, which tells you the maintainers care about how figures look in print and on the web. Omnigraffle drawings use 12pt Guardian Sans Condensed or Ubuntu mono, are exported as 300 dpi png, and are then included with a dpi argument that compensates for the difference between the website font size and the drawing font size. The README gives this example:

## Building the book locally with Quarto

The README does not give a step-by-step local build recipe, so what follows is derived from the repository layout and the stated tool, not from a documented procedure. The README names Quarto as the build system, and _quarto.yml is the project file, so the ordinary Quarto workflow applies: render the project from the repository root.

First clone the repository and move into it. The default branch is main.

```bash
git clone https://github.com/hadley/r4ds.git
cd r4ds
```

With Quarto installed, render the whole book from the project file. Because _freeze/ is committed, Quarto can reuse cached results for chapters whose code has not changed, which shortens a rebuild considerably.

```bash
quarto render
```

To work on a single chapter instead of the whole book, render that file directly. The chapter files are at the root, so the path is short.

```bash
quarto render data-visualize.qmd
```

R itself is needed for the code chunks, and the book's chapters depend on packages the text introduces rather than on a single declared dependency list in the README. The DESCRIPTION file at the root is the place to look for what the build expects. If a chunk fails, the error will name the missing package.

The README does document one specialised path, for producing the O'Reilly edition. It requires the htmlbook package, installed from GitHub, and it writes output into a sibling directory named r-for-data-science-2e before copying HTML and PNG files there. The README then says to commit and push to atlas. That path is specific to the publisher's workflow and is not something a reader needs.

## The dpi setting in figure includes, and why it is easy to get wrong

The README spends more space on image sizing than on anything else, and the reason is a real constraint: a figure that looks correct in the browser can look wrong in print, and vice versa. The README states that the website font is 18 px, which equals 13.5 pt, and that the drawing font is 12 pt. The export dpi is therefore scaled by 12 / 13.5, giving 270 from a 300 dpi source. The README notes the author also verified this empirically by screenshotting.

The include call passes that value explicitly, with echo and out.width set to suppress the code and the default width:

```r
#| echo: FALSE
#| out.width: NULL
knitr::include_graphics("diagrams/transform.png", dpi = 270)
```

Screenshots take a different route. The README says to use a light theme, to zoom in twice for small interface elements such as toolbars, and to capture with Cmd + Shift + 4. For screenshots the dpi argument is omitted entirely, and the include call is just the file path:

```r
#| echo: FALSE
#| out.width: NULL
knitr::include_graphics("screenshots/rstudio-wg.png")
```

This is the kind of detail that only matters if you are editing figures. If you are reading the book, it is invisible. If you are contributing a diagram, getting it wrong produces a figure that is subtly the wrong size, and the README's arithmetic is the only guidance offered.

## Where the repository is the wrong choice

The most common mismatch is expecting an artefact to install. There is no package to load, no function to call, and no CRAN entry implied by this repository. Someone searching for hadley r4ds pdf or hadley r4ds free download is looking for the book as a document, and the repository is not that either: it is source that has to be rendered. The website at r4ds.hadley.nz is the readable form.

The O'Reilly build path is a second trap. It depends on htmlbook, installed with pak::pak("hadley/htmlbook"), and it copies files into ../r-for-data-science-2e/, a directory that exists only in the maintainer's working setup. Running that code without the sibling directory will fail at the file.copy step. The README does not document a fallback, and it does not describe how to obtain the directory.

A third limitation is scope. The repository is a book, so it has no release cadence in the software sense. The only release listed is first-ed, dated 2021-02-04. That tag marks the first edition, not an ongoing series of versioned artefacts, and the second edition is tracked through commits to main rather than through releases. Anyone treating the release list as a versioning scheme will find nothing to upgrade between. The last push to the repository was on 2026-07-18, so work has continued since that tag, but the release list does not reflect it.

## How it compares with a maintained code-first alternative

The closest thing to an alternative in the R documentation world is a package that ships its own vignettes, such as the tidyverse packages themselves. The difference in approach is structural. A package vignette is versioned with the code, installed alongside it, and read with vignette() after installation; it documents the functions in that package and moves when the package moves. The r4ds repository is the inverse: the prose is the primary artefact, the code exists to demonstrate it, and the whole thing is assembled by Quarto into a book rather than attached to a library.

That difference decides which one you want. If you need reference material for a specific function, a package vignette is the right source, because it is pinned to the version you installed. If you need a teaching sequence that moves from visualisation through transformation to programming and communication, the book's chapter order is the point, and no single package vignette covers that arc. The book also covers material that is not owned by any one package, such as spreadsheets, databases, quarto and rectangling.

A second comparison is with the first edition. The repository carries a preface-2e.qmd file alongside a preface, and the release list names first-ed. The second edition reorganised the material, which is why chapter files such as logicals.qmd, numbers.qmd and missing-values.qmd exist as separate units. If you learned from the first edition, the file list is not a one-to-one mapping of what you remember.

## Licence, contribution and what to check before forking

The GitHub licence field for this repository reads NOASSERTION, which means the platform could not classify the LICENSE file at the root into a known licence. That is a statement about tooling, not about the terms. Before you reuse chapters in a course pack, translate the book, or republish figures, read the LICENSE file itself and the DESCRIPTION file, which in R projects conventionally carries a License field. The README does not summarise the terms, so there is nothing here to paraphrase.

Contribution terms are stated. The README points to a Contributor Code of Conduct at contributor-covenant.org, version 2/0, and says that by contributing to the book you agree to abide by its terms. The CODE_OF_CONDUCT.md file is present at the root. There is also a contribute.qmd chapter, which is where the book itself explains how to take part.

Upgrade cost is low in the usual sense, because there is nothing to upgrade. You pull main and re-render. The real cost is environmental: a working Quarto installation, the R packages the chapters use, and enough patience for a full render if the freeze cache is invalidated. If you only want to read the second edition, none of that applies, and the website is the cheaper path.

## Conclusion

Adopt hadley/r4ds if you want to read the second edition at r4ds.hadley.nz, translate it, or send a pull request against a specific .qmd chapter; the README asks contributors to follow the Contributor Code of Conduct. Do not clone it expecting an installable R package, and do not expect the O'Reilly build path to work without the htmlbook package and the local r-for-data-science-2e directory. Before relying on the repository, check the LICENSE file at the root and the DESCRIPTION file, since the GitHub licence field reads NOASSERTION.

## FAQ

### Is hadley/r4ds an R package I can install?

No. The README describes the repository as the source of the R for Data Science book, and the top-level entries are chapter files such as data-visualize.qmd and data-transform.qmd rather than an R/ directory. The readable output is published at r4ds.hadley.nz.

### How do I build the R for Data Science book from the hadley/r4ds source?

The README states the book is built using Quarto, and _quarto.yml is the project file at the root. Cloning the repository and running quarto render from the root is the natural workflow, with _freeze/ providing cached results for unchanged chapters.

### What does the NOASSERTION licence on hadley/r4ds mean for reuse?

It means GitHub could not classify the LICENSE file into a known licence, not that no terms exist. Read the LICENSE and DESCRIPTION files at the root before reusing chapters, translations or figures, since the README does not summarise the terms.

## Sources

- [hadley/r4ds on GitHub](https://github.com/hadley/r4ds)
- [Issues](https://github.com/hadley/r4ds/issues)
- [Project website](https://r4ds.hadley.nz/)
- [README](https://github.com/hadley/r4ds/blob/main/README.md)
- [Releases](https://github.com/hadley/r4ds/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/hadley-r4ds
