# sparklyr's README is knitted, which is why its examples print ?? for row counts

> The R interface to Apache Spark, distributed as a CRAN package and documented as a rendered R Markdown file. Reading it closely shows what that build step froze into the page: unknown row counts, output cut mid-record, a GAM smoothing formula printed as if it were the answer, and three example dependencies the install section never mentions.

**sparklyr/sparklyr** — R interface for Apache Spark

- Repository: https://github.com/sparklyr/sparklyr
- Website: https://spark.posit.co/
- Stars: 972 · Forks: 311
- Language: R
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/sparklyr-sparklyr

## The README is generated, and that is visible in its own output

The first line of the file is an HTML comment saying that README.md is generated from README.Rmd and that the source file is the one to edit. That single note explains most of what is odd about the page. Everything in it is knitted output, captured when the document was built, and three artifacts of that are still sitting in the published text. The row counts never resolved: the filter example prints its source header as `SQL [?? x 19]` and the window function example as `SQL [?? x 7]`, so the number of rows in a lazy Spark table was still a placeholder when the page was rendered and the placeholder shipped. Two examples also stop in the middle of a record, with the last printed row of the flights example reading `6  2013     1     1` and the last row of the batting example ending at a bare `1892`. And the plotting example prints a message rather than a result, `geom_smooth() using method = 'gam' and formula = 'y ~ s(x, bs = "cs")'`, which is ggplot2 reporting the smoothing it chose, frozen into the documentation as if it were output worth keeping.

## Three install paths, and the newest code is not the one on CRAN

There are three documented routes, and they do not all give you the same code. The CRAN route is one line and gets you a released version:

```
install.packages("sparklyr")
```

A local Spark install is a separate step, described as being for development purposes, and it is what the rest of the page assumes exists on your machine:

```
library(sparklyr)
spark_install()
```

The third route is the one to read carefully. To upgrade to what the documentation calls the latest version, you install devtools and then install from GitHub, and you restart your R session afterwards:

```
install.packages("devtools")
devtools::install_github("sparklyr/sparklyr")
```

So the current development state of the package reaches you through the default branch, not through CRAN, and the session restart is part of the instruction rather than an optional hint. The release history is separate from that: v1.9.5 from 2026-06-22, v1.9.4 from 2026-04-20, and v1.9.2 from 2025-10-06.

## The examples use four packages and the install section names two of them

The dplyr walkthrough opens by telling you that you may need two extra packages to execute the code, and gives the line:

```
install.packages(c("nycflights13", "Lahman"))
```

The plotting example then calls `library(ggplot2)` with no install line anywhere near it, and the SQL section calls `library(DBI)` the same way. So four packages are needed to run the front-page examples end to end, and only two are named. The data copy itself is worth reading for its defaults, because every call passes `overwrite = TRUE`:

```
library(dplyr)
iris_tbl <- copy_to(sc, iris, overwrite = TRUE)
flights_tbl <- copy_to(sc, nycflights13::flights, "flights", overwrite = TRUE)
batting_tbl <- copy_to(sc, Lahman::Batting, "batting", overwrite = TRUE)
src_tbls(sc)
#> [1] "batting" "flights" "iris"
```

The listing comes back alphabetically rather than in the order the tables were created, which matters if you are reading output to work out what happened. The overwrite default is what makes the example re-runnable, and it also means an existing table of the same name is replaced without asking.

## Four cluster managers are named and exactly one is demonstrated

The opening summary says the package installs and connects to Spark using YARN, Mesos, Livy or Kubernetes. The body then connects to a local instance and nothing else:

```
library(sparklyr)
sc <- spark_connect(master = "local")
```

The text says you can connect to both local instances and remote clusters, shows the local case, and sends remote deployments to a Deployment section on the project website rather than describing any of the four named managers here. The table of contents does reserve its own sections for two specific routes further down, one for connecting through Livy and one for connecting through Databricks Connect, so the four-way claim is a summary of routes rather than four documented paths. That Databricks entry is written across two lines in the table of contents, breaking after the word Databricks, and its anchor carries a version marker, `#connecting-through-databricks-connect-v2`, which is the only place in the file where a section identifier states which generation of an interface it documents.

## Fourteen promised sections, and a machine learning example that stops mid-sentence

The table of contents lists fourteen sections, running from installation and connecting through dplyr, SQL and machine learning to extensions, distributed R, table utilities, connection utilities, RStudio IDE, H2O, Livy and Databricks Connect. The feature summary at the top matches that breadth with four claims: connecting to Spark through those managers, using dplyr against datasets and streams and bringing results back into R, creating machine learning pipelines, and creating extensions that either call the full Spark API or run distributed R code. The machine learning section describes its functions as connecting to high-level APIs built on DataFrames, and points at Spark's own mllib guide. Its worked example starts with `ml_linear_regression` on the built-in `mtcars` dataset, asking whether fuel consumption can be predicted from weight and from the number of cylinders, and the sentence ends there. So the one part of the page a data scientist is most likely to copy starts and stops in the same paragraph.

## The package root carries release artifacts, a pkgdown config and a .claude directory

The repository root is a CRAN package with its publishing machinery checked in. `DESCRIPTION`, `NAMESPACE`, `man/`, `man-roxygen/`, `R/`, `inst/`, `tests/`, `utils/`, `tools/` and `java/` are the standard parts, with `java/` there because the package ships JVM-side code. Around them sit the release artifacts: `CRAN-SUBMISSION` and `cran-comments.md` record a submission, `revdep/` holds reverse dependency checks for a package that many others depend on, `_pkgdown.yml` configures the documentation site the README links to, `codecov.yml` and `appveyor.yml` add coverage and Windows build configuration, and `.Rbuildignore` keeps non-package files out of the tarball. Two working files are less usual. `README.Rmd` sits beside the `README.md` it generates, and the root also contains a `.claude/` directory, which is an agent configuration directory rather than anything R, and which neither the README nor the documentation links mention. The badge row above the summary points at two GitHub Actions workflows, one for R CMD check and one for Spark tests, plus the CRAN package page and a coverage report pinned to the main branch, so the branch that generates this page is also the branch whose test and coverage state is on display.

## Filtering stays lazy, window functions carry their ordering into SQL

The dplyr section claims all available verbs work against cluster tables, and the examples show the difference between a lazy filter and a collected result. The filter example never calls `collect()`, so what prints is a description of pending work, including the source and database lines and the unresolved row count. The window function example is more instructive, because it shows metadata travelling into the query:

```
batting_tbl %>%
  select(playerID, yearID, teamID, G, AB:H) %>%
  arrange(playerID, yearID, teamID) %>%
  group_by(playerID) %>%
  filter(min_rank(desc(H)) <= 2 & H > 0)
#> # Source:     SQL [?? x 7]
#> # Database:   spark_connection
#> # Groups:     playerID
#> # Ordered by: playerID, yearID, teamID
```

The `AB:H` range in the select is the tidy evaluation reaching into column positions, `min_rank` inside a grouped filter is the window function doing the ranking, and the printed header confirms that the grouping and the sort both survive translation into SQL. The comparison example then ends with `collect()` before handing the result to ggplot2, which is the point where data stops being a query and becomes an R data frame.

Two other routes complete the picture. SQL goes through the DBI interface and returns an R data frame immediately rather than a lazy table:

```
library(DBI)
iris_preview <- dbGetQuery(sc, "SELECT * FROM iris LIMIT 10")
iris_preview
```

And the aggregation example shows the whole pipeline, computing per-tail-number counts and means in the cluster, filtering the result there, and only then collecting:

```
delay <- flights_tbl %>%
  group_by(tailnum) %>%
  summarise(count = n(), dist = mean(distance), delay = mean(arr_delay)) %>%
  filter(count > 20, dist < 2000, !is.na(delay)) %>%
  collect()
```

The `count > 20` guard is also why the smoothing that follows is a generalised additive model: the per-group sample is small, and the cubic spline basis printed alongside it describes that compromise rather than a modelling decision anyone wrote down.

## Conclusion

sparklyr suits R users who want dplyr verbs, a DBI connection and mllib pipelines against a Spark cluster without leaving the R session, and the front-page examples are runnable once you install the two data packages they name. It is not a place to learn Spark deployment from, since only the local case is demonstrated and remote managers are deferred to the website. Before trusting anything in the documentation, remember the page is a build artifact: edit README.Rmd rather than README.md, expect the printed output to be stale rather than wrong, and install ggplot2 and DBI yourself, since the examples use both without saying so.

## FAQ

### How do I install sparklyr and connect to a local Spark instance?

Install the package from CRAN with `install.packages("sparklyr")`, then call `spark_install()` to put a local Spark in place for development. `sc <- spark_connect(master = "local")` returns the connection object that dplyr verbs run against.

### How do I upgrade sparklyr to the newest code?

The documented route is `install.packages("devtools")` followed by `devtools::install_github("sparklyr/sparklyr")`, and you restart your R session afterwards. That path comes from GitHub rather than CRAN, so it is ahead of the published release.

### Can sparklyr run SQL and dplyr window functions against a cluster?

Yes. The connection object implements the DBI interface, so `dbGetQuery()` returns results as an R data frame, and dplyr window functions are supported, with grouping and ordering carried into the generated SQL as shown by the printed Groups and Ordered by lines.

### Which Spark cluster managers can sparklyr connect to?

The feature summary names YARN, Mesos, Livy and Kubernetes. The worked example connects only to a local instance, and remote deployments are pointed at the Deployment section of the project website rather than configured here.

### How does sparklyr compare with PySpark?

This repository does not answer that. Nothing in its README mentions PySpark or compares the two, so the comparison has to come from Spark's own multi-language documentation rather than from the sparklyr docs.

## Sources

- [License: Apache-2.0](https://github.com/sparklyr/sparklyr/blob/main/LICENSE)
- [Project website](https://spark.posit.co/)
- [README](https://github.com/sparklyr/sparklyr/blob/main/README.md)
- [Releases](https://github.com/sparklyr/sparklyr/releases)
- [sparklyr/sparklyr on GitHub](https://github.com/sparklyr/sparklyr)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/sparklyr-sparklyr
