# stringi, four install targets and an ASCII gate on the sources

> stringi is the R package for string, text, and natural language processing, built on ICU so that collation, case mapping, transliteration, and normalisation behave the same in every locale. The README is short because the manual is a website. The build system is where the detail lives: four different install configurations, a source gate that rejects any non-ASCII file, and a phony target list that covers less than half of what the Makefile defines.

**gagolews/stringi** — Fast and Portable Character String Processing in R (with the Unicode ICU)

- Repository: https://github.com/gagolews/stringi
- Website: https://stringi.gagolewski.com/
- Stars: 318 · Forks: 48
- Language: C++
- License: NOASSERTION
- Published: 2026-09-18 · Updated: 2026-09-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/gagolews-stringi

## Four install targets, and they are not equivalent

The Makefile defines four ways to install the same package, and the differences are not cosmetic.

```make
r-icu-system:
	R CMD INSTALL .

r-icu-bundle:
	R CMD INSTALL . --configure-args='--disable-pkg-config'

r-icu-bundle55:
	R CMD INSTALL . --configure-args='--disable-cxx11 --disable-pkg-config'
```

The default target, which is simply called r, disables pkg-config and turns on two compiler diagnostics instead. The three above trade those diagnostics away for different ICU sources: the system ICU found through pkg-config, the ICU bundled in the repository, and the bundled ICU with C++11 support switched off.

That last one is the interesting case. Disabling the language standard is a compatibility valve for toolchains that predate it, and keeping it as a supported build variant means the package still installs on older compilers long after C++11 became the default everywhere else.

The consequence for a user is that which ICU you run against changes collation and case mapping results. Same R, same code, different sort order. That is the price of portability being delegated to a library with four build configurations, and it is why the package documentation spends more words on the build than a typical R package would.

The default target is the one you want if you are just installing.

## The check target turns off the CRAN incoming checks

The release check runs as cran but suppresses the part that would contact cran.

Three environment variables are set before the check command. STRINGI_DISABLE_PKG_CONFIG forces the bundled configuration, R_DEFAULT_INTERNET_TIMEOUT is raised to 240 seconds, and both _R_CHECK_CRAN_INCOMING_ and its REMOTE_ variant are set to false.

Raising the timeout is the clue. The incoming checks are the ones that fetch information from CRAN and compare against it, and a package that checks its own sources or downloads reference material needs more than the default allowance. Turning the checks off entirely, while keeping the as-cran flag, means the local rules still apply and the remote comparison does not.

The check also passes --no-manual, so the PDF manual is not built, and it selects the tarball to test with a timestamp sort rather than a fixed filename, which is what makes repeated local runs convenient.

The check target depends on both the ASCII gate and the build, so neither is optional. The order in the dependency list is the enforcement.

## A gate that stops the build on any non-ASCII source file

One target exists purely to fail, and it is the first thing that runs before a release check.

The command pipes a file type inspection of DESCRIPTION, configure, configure.win, NAMESPACE, cleanup, everything under R, src, man, inst, and tools through two greps, one selecting text files and one rejecting anything that is not us-ascii. It then asserts that the result is empty.

In practice, that means a single non-ASCII character in a single source file stops the release check before a compiler is invoked. For a package whose entire purpose is Unicode handling, the rule is not about the data. It is about the code staying portable across the encodings a source tree might be checked out under.

The list of inspected paths is where this gets slightly stale. tools/ does not appear in the repository's top-level listing, while R, src, man, inst, datasets, and docs all do. A glob that matches nothing is not an error, so the gate passes silently over a directory that is no longer there, which means any file that later moves into a new location has to be added to this list by hand.

That is the cost of an explicit gate, and the benefit is that the failure is a stopped build rather than a corrupted artifact.

## Eight targets are marked phony and the rest are not

The Makefile declares fewer phony targets than it defines, and one of the omissions can change behaviour.

The declaration covers eight names: r, check, build, clean, purge, html, docs, and test. The file also defines stop-on-utf8, autoconf, r-icu-system, r-icu-bundle, r-icu-bundle55, tinytest, check-revdep, and a default all target.

A target that is not declared phony is satisfied by a file of the same name. Make checks for that file first and skips the recipe if it exists. So if a file named tinytest, check-revdep, or autoconf ever appears in the working directory, running make on that target quietly does nothing.

autoconf is the sharp case, because a program of that name is installed on any machine that builds this package from source. A stray file or directory with that name in the checkout would make the autoconf target a no-op, and the recipe that regenerates the configure script is the one that runs before every install and every release build.

A second detail is deliberate and commented: the test target must not be merged with tinytest, because merging would remove the ability to run the test suite against the other three ICU configurations. That is a small piece of build hygiene that would be easy to break by accident.

## The repository ships a subset of ICU and a binary database

The licensing note is longer than the feature list, and it explains what is actually in the tree.

The README states that the package's own source code is distributed under an open source BSD-3-clause licence. It then says the same repository also contains a custom subset of ICU4C source code, copyrighted by Unicode, Inc. and others, and that a binary version of the Unicode Character Database is included.

So there are two licences in one repository. The ICU project is covered by the Unicode licence, described as a simple, permissive non-copyleft free software licence compatible with the GNU GPL. The README links the upstream explanation of why that licence is what it is: it exists so that ICU can be included in free software projects and also in proprietary and commercial products.

That is the reason the subset is bundled at all. Bundling means an install does not depend on finding a matching system ICU, which is what makes the package portable across platforms and locales in the first place, and the permissive upstream licence is what makes bundling into a GPL package and into commercial software both possible.

The tradeoff is a repository that carries a data file and a code subset whose updates are not the maintainer's to make. The requirement of ICU4C 61 or newer is the floor, and R 3.4 or newer is the interpreter floor.

## stringr became a wrapper here in 2015

The relationship with the tidyverse package is stated in one paragraph, and it explains the API you already know.

The stringi API was inspired by the early version of stringr, described as pre-tidyverse and version 0.6.2. And since stringr version 1.0.0 in 2015, stringr has been powered by stringi.

That second sentence is the important one for anyone choosing between them. The functions people call in stringr are wrappers over the calls in this package, so the tidyverse names are a compatibility surface rather than an independent implementation. When the two disagree, the difference is in the wrapper, not in the text handling underneath.

The package also points at stringx, a set of wrappers around stringi with a base R-compatible API, which is the same idea applied to base R rather than to tidyverse. So there are two convenience layers on offer and one engine.

The feature list behind that engine covers concatenation, padding, and wrapping, substring extraction, pattern searching with Java-like regular expressions, collation and sorting, random string generation, case mapping and folding, transliteration, Unicode normalisation, and date-time formatting and parsing. The citation is a 2022 paper in the Journal of Statistical Software, and the reference manual is a website rather than a document in the repository.

## Sixteen months between two patch releases

The release history is uneven, and one pair of releases landed a day apart.

The three most recent tags are v1.8.6 on 26 March 2025, v1.8.7 on 27 March 2025, and v1.8.8 on 28 July 2026. Two patch releases one day after each other, then a sixteen month gap to the next one.

The last push to the repository is 18 August 2026, a few weeks after the newest tag, and the project is not archived. So there is movement, just not on a schedule you can plan around.

For a package this old and this widely depended on, that is not alarming on its own. The changelog lives in a NEWS file in the repository root rather than on the package page, which is the older convention and means the release notes are only in the git tree.

What the gap does affect is pinning. A team that wants reproducible builds will want 1.8.8 pinned explicitly, and should expect the build to depend on which of the four targets was used when it was produced, since the ICU source changes collation behaviour.

## Conclusion

stringi is the default answer for text work in R, and for most users the CRAN install is all that is needed. Reach for the repository when you are building an R package that needs the C++ layer, or when you need to reproduce a specific ICU configuration. Check four things first. Which ICU you want, since the Makefile has four install targets and they are not equivalent. That your sources stay ASCII, because the build will stop. Whether your reverse dependency checks pass, since that target is a separate step with an unfinished cleanup note. And whether the version you pin is 1.8.8, given the long gap before it.

## FAQ

### stringi vs stringr: what is the difference?

stringi's API was inspired by the early pre-tidyverse version 0.6.2 of stringr. Since stringr version 1.0.0 in 2015, stringr has been powered by stringi, so stringr functions are wrappers over stringi calls rather than a separate implementation.

### Is there a stringi alternative with a simpler R API?

The project points at stringx, a set of wrappers around stringi with a base R-compatible API. It also publishes a tutorial and reference manual on the package homepage and links a free open-access textbook on R programming.

### What are the system requirements for stringi?

R 3.4 or newer and ICU4C 61 or newer. The repository contains a custom subset of ICU4C source code plus a binary version of the Unicode Character Database, and an INSTALL file in the repository root has more detail.

### How does stringi handle collation and case across locales?

Portability comes from ICU, the International Components for Unicode library. The package covers collation and sorting, case mapping and folding, transliteration, Unicode normalisation, and pattern searching with Java-like regular expressions.

## Sources

- [gagolews/stringi on GitHub](https://github.com/gagolews/stringi)
- [Issues](https://github.com/gagolews/stringi/issues)
- [Project website](https://stringi.gagolewski.com/)
- [README](https://github.com/gagolews/stringi/blob/master/README.md)
- [Releases](https://github.com/gagolews/stringi/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/gagolews-stringi
