Library / SDK
gagolews/stringi avatar
gagolews/stringi

stringi: the ICU layer under string processing in R

Fast and Portable Character String Processing in R (with the Unicode ICU)

318 stars48 forksC++NOASSERTION

At a glance

What is it?
gagolews/stringi is an R package for string and text processing built directly on the ICU library, maintained by Marek Gagolewski and published in the Journal of Statistical Software. It is the engine stringr has used since 2015, and installing it means choosing between your system ICU and the bundled subset.
Who is it for?
Use stringi if you process text in R and care that the result is the same on every platform and in every locale, which is what ICU buys you for collation, normalisation, case mapping and transliteration. Most people should simply install from CRAN and let stringr call it.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 44 days ago.
What is it written in?
Mainly C++, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Where stringi sits in R

stringi is an R package for string, text and natural language processing, implemented in C++ on top of ICU, the International Components for Unicode library. The README states the system requirements as R 3.4 or newer and ICU4C 61 or newer, with details in the INSTALL file.

Its position in the ecosystem is unusual. The README explains that stringi's API was inspired by an early pre-tidyverse version of Hadley Wickham's stringr, specifically v0.6.2, and that since stringr 1.0.0 in 2015, stringr has been powered by stringi. So the dependency arrow points the way people often assume it does not: the tidyverse-facing package is the wrapper, and this is the engine.

The maintainer is Marek Gagolewski, with contributions from Bartlomiej Tartanus and others. The package has a published reference: the README cites the Journal of Statistical Software, volume 103 issue 2, 2022, pages 1 to 59, doi 10.18637/jss.v103.i02.

If you write R code that touches text, you are very likely already running this package without importing it directly.

Why ICU is the whole argument

The README's feature list is long, and nearly every line is something ICU provides rather than something reimplemented in R: string concatenation, padding and wrapping, substring extraction, pattern searching with Java-like regular expressions, collation and sorting, random string generation, case mapping and folding, transliteration, Unicode normalisation, and date-time formatting and parsing.

Collation is the one that decides most real arguments. Sorting strings by code point is not sorting them the way a human in a given locale expects, and base R's behaviour depends on the session locale. ICU provides locale-aware collation, so the same code produces the same order on a colleague's machine in a different country.

Case mapping and Unicode normalisation matter for the same reason. Two strings that look identical can differ in composition, and uppercasing is not a per-character operation in every script. Transliteration lets you fold characters across scripts, which is what you want before comparing user input against a reference list.

Portability across locales and platforms is the claim in the README, and it is a claim about ICU's data rather than about R.

Installing stringi and choosing an ICU

The package is on CRAN, and the README links to its CRAN entry. Behind a normal install is a decision the repository exposes directly. The Makefile in the root defines several install targets, and two of them matter. Building against your system ICU is the plain install:

bash
R CMD INSTALL .

That path relies on pkg-config finding a system ICU4C. The alternative is to build the bundled ICU subset shipped in the repository, which is what happens when pkg-config is disabled:

bash
R CMD INSTALL . --configure-args='--disable-pkg-config'

The bundled route is slower to compile and produces a larger library, but it does not depend on what the host happens to have installed, which is why it is the safer choice on a minimal container or an old distribution. The system route is faster and keeps you on the ICU version your platform patches.

Either way this is a compiled install with a configure step, so on a machine without a toolchain expect a source build to fail rather than to fall back gracefully.

What the build actually costs

The root listing shows a full autotools-style setup: configure.ac, configure, configure.win, a cleanup script and an INSTALL file. That is not typical for an R package, and it exists because of ICU.

The cost is real. A source build compiles C++ against ICU, so it takes minutes rather than seconds, and it is the reason stringi is often the slowest package in a fresh R environment. On managed platforms a binary package avoids that entirely, and using one is the practical advice for anyone who does not have a reason to build from source.

The project tracks this. Release v1.8.8, published on 2026-07-28, lists two changes: fixed build and runtime warnings, and a configure script that no longer tries to fall back to C++11. That second item removes a compatibility path, so hosts with only an older C++ standard are now out, which is consistent with the Makefile still carrying an r-icu-bundle55 target that passes --disable-cxx11.

The repository also carries a check target that runs R CMD check --as-cran with STRINGI_DISABLE_PKG_CONFIG set, and a check-revdep target that runs reverse dependency checks, so upstream breakage is watched deliberately.

Three licences, one package

The README is explicit that this is not a single-licence package, which is why GitHub's detector reports NOASSERTION instead of naming one.

The source code of stringi itself is distributed under the BSD-3-clause licence. That is permissive and compatible with essentially everything.

Layered on top, the git repository contains a custom subset of ICU4C source code copyrighted by Unicode, Inc. and others, along with a binary version of the Unicode Character Database. The README states the ICU project is covered by the Unicode license, which it describes as a simple permissive non-copyleft licence compatible with the GNU GPL and intended to allow ICU in both free and proprietary products.

For most users none of this changes anything, since you are consuming a CRAN package rather than redistributing it. If you redistribute a product that embeds R and stringi, or you vendor the source, you are dealing with BSD-3-clause for the R and C++ code and the Unicode licence for the bundled data. That is a description of what the README declares, not legal advice.

stringr, stringx and base R

The three alternatives are all closer than they look, because two of them are this package wearing different clothes.

stringr is the tidyverse-facing wrapper. Since version 1.0.0 in 2015 it has been powered by stringi, so it offers a smaller, more consistent API over the same engine. Choose stringr for readability and for code that fits tidyverse style, and drop to stringi when you need something stringr does not expose, such as collation control or transliteration.

stringx is a separate project by the same author, described in the README as a set of wrappers around stringi with a base R compatible API. It exists for people who want ICU correctness behind the function names they already know from base R.

Base R itself is the real alternative, and the difference is locale dependence. Base R string functions behave differently across platforms and locales, and lack ICU-grade collation, normalisation and transliteration. If your scripts only ever run on your laptop with ASCII input, you will not notice. If they run in CI in a container with a C locale, or process names in multiple scripts, you will.

Maintenance and stability

The last push to the repository was on 2026-08-18. Releases are infrequent and small: v1.8.8 on 2026-07-28, v1.8.7 on 2025-03-27 with an empty change list, and v1.8.6 on 2025-03-26 fixing build warnings plus a PROTECT stack imbalance in stri_encode_from_marked.

That release pattern is what a mature dependency looks like. There is no feature churn, and the changes are build fixes and correctness fixes in the C layer. A PROTECT stack imbalance is a memory-management bug in R's C interface, which is the category of defect that causes crashes far from the call that triggered it, and the fact that it shipped and was fixed is a reasonable signal that the C code is exercised rather than dormant.

The practical maintenance question is ICU versions rather than stringi releases. If you build against a system ICU, an ICU upgrade on the host changes collation and normalisation behaviour underneath you, including potentially the sort order of data you have already written. Pinning the ICU version, or using the bundled subset, is the way to keep that stable.

Editorial conclusion

Use stringi if you process text in R and care that the result is the same on every platform and in every locale, which is what ICU buys you for collation, normalisation, case mapping and transliteration. Most people should simply install from CRAN and let stringr call it. Reach for the bundled ICU build with --disable-pkg-config on minimal containers or old distributions where the system ICU4C is missing or older than 61, and accept the longer compile. Do not expect frequent feature releases, because the recent versions are build and correctness fixes. Verify one thing before relying on it in production: whether your deployment pins the ICU version, since a host ICU upgrade can change sort order under data you have already stored.

Frequently asked questions

How do I install stringi in R?

install.packages("stringi") from CRAN is the normal route. When building from source, R CMD INSTALL . uses the system ICU and R CMD INSTALL . --configure-args='--disable-pkg-config' builds the bundled ICU subset instead.

What is the relationship between stringi and stringr?

The README states that stringi's API was inspired by an early stringr version, v0.6.2, and that since stringr 1.0.0 in 2015 stringr has been powered by stringi.

Which licence does stringi use?

The source code is BSD-3-clause. The repository also bundles a subset of ICU4C source and a binary Unicode Character Database covered by the Unicode license, which is why GitHub reports the repository licence as NOASSERTION.

What system requirements does stringi have?

The README lists R 3.4 or newer and ICU4C 61 or newer, with further detail in the INSTALL file. The bundled ICU subset is the fallback when the system library is unsuitable.

Why does installing stringi take so long?

It compiles C++ against ICU through a configure step, which is why the repository carries configure.ac, configure and configure.win. Using a binary package from your platform avoids the compile.

Official sources

  1. gagolews/stringi on GitHub
  2. Issues
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/gagolews-stringi.svg)](https://hysenlabs.com/projects/gagolews-stringi)