ICU: the Unicode and locale library behind C++ and Java text handling
The home of the ICU project source code.
At a glance
- What is it?
- The unicode-org/icu repository holds ICU4C and ICU4J, the Unicode Consortium's libraries for collation, formatting and character properties. It is infrastructure, not an application, and its cost is measured in build time and data size.
- Who is it for?
- Adopt ICU when you need locale-correct collation, number and date formatting, or Unicode character properties that your platform's runtime does not provide, and when you can accept a large data payload and a real build step. Do not adopt it for a small utility that only needs UTF-8 validation, and do not assume the JDK's bundled ICU4J matches the current release.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly C++, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What ICU4C and ICU4J actually provide
ICU is the Unicode Consortium's implementation of Unicode and locale behaviour, split into two parallel libraries that live in the same repository. The README points at icu4c/ for C and C++ and icu4j/ for Java, and those two directories are where essentially all of the code is. Everything else in the tree is scaffolding: tools/ for build and data tooling, vendor/ for vendored dependencies, docs/ for documentation sources.
The problems it solves are the ones that look trivial until a product ships in more than one language. Sorting strings so that Swedish and German agree with their own readers. Formatting a date, a currency amount or a number according to a locale's conventions rather than the C locale's. Asking whether a code point is a letter, a digit, or a combining mark. Segmenting text into words and sentences for languages that do not use spaces. Converting between Unicode encodings and legacy character sets. Each of these has edge cases that number in the thousands, and ICU is the artifact where that accumulated knowledge is stored.
The intended audience is library and platform authors rather than application developers. If you are writing an application in Java, you are almost certainly already calling ICU through the JDK, which bundles a version of ICU4J for its text, date and number formatting. If you are writing C or C++, you are the one who has to decide whether to depend on ICU4C or on whatever locale support your C library offers. That decision is the interesting one, and the repository does not make it for you.
How the data pipeline shapes the architecture
The part of ICU that surprises newcomers is that the libraries are not primarily code. They are code plus a large body of locale and Unicode data, and the repository is organised around producing that data. The tools/ directory exists because the data has to be compiled: source data is transformed into binary resource bundles that the runtime libraries load.
That pipeline is why ICU4C and ICU4J can stay behaviourally aligned despite being written in different languages. They consume the same conceptual data set, and the project's own quality reporting treats API comparison as a first-class concern. The README links an API Comparison report for ICU4J alongside Errorprone and JaCoCo reports, which tells you the maintainers track surface changes deliberately rather than letting the two libraries drift.
The consequence for an adopter is that upgrading ICU is not just a code change. New Unicode versions bring new characters, new collation rules and revised locale data, and the binary payload grows. The repository layout reflects this: the data is a build input, and the build produces artifacts whose size you control only by choosing how much data to include. The README does not document a supported way to trim that payload; the icu4c/ and icu4j/ readmes are where you would look, and they are linked rather than reproduced at the top level.
Building ICU4C from source: a first real use
The repository does not present a single install command, because ICU is a source tree with two independent build systems. The README's job is to route you to icu4c/ or icu4j/; the build instructions live in those subdirectories and in the user guide at unicode-org.github.io/icu.
For ICU4C, the repository ships a top-level .bazelrc and WORKSPACE, so a Bazel build is present in the tree, and the icu4c/ subdirectory is where the C and C++ library is built. The README itself does not print a configure or make sequence; check icu4c/readme.html before running anything, because that is the file the README links for ICU for C/C++.
The README gives no code sample for either library, so there is nothing to copy here. What it does give is the set of entry points: the source at github.com/unicode-org/icu, bugs at unicode-org.atlassian.net/projects/ICU, API documentation at unicode-org.github.io/icu-docs, and the user guide at unicode-org.github.io/icu. For Java, the repository contains a top-level pom.xml and a .mvn/ directory, so ICU4J is a Maven project rooted at icu4j/, and icu4j/readme.html is the file to read before assuming an invocation.
The practical first step is therefore not a snippet but a build. Confirm which of the two libraries you need, read the readme in that subdirectory, and only then write code. If you are working from the source tree rather than a published artifact, the subdirectory readme is the only place in this repository that describes how the tree is meant to be built.
Where ICU is the wrong dependency
ICU is large, and that is a design outcome rather than an oversight. A binary that links ICU4C carries a data payload that can dwarf the application code around it. For a command-line tool that needs to uppercase a handful of ASCII strings, or a service that only ever handles one locale, the dependency is disproportionate. The C standard library and the platform's own locale support are the right answer there, warts and all.
The second limitation is version skew. Because so much software already links an ICU that ships with the operating system, adding ICU4C to a project can produce two copies in one process, and the one you compiled against may not be the one that loads first. That is a packaging problem rather than a bug in ICU, and the repository does not solve it for you.
The third is that the project is upstream infrastructure with a long release cadence and a formal contribution process. The README states that a CLA is required to contribute. If your organisation needs a behaviour change in collation or locale data, you are participating in a standards-adjacent process, not patching a small library. Teams that want to move fast on i18n behaviour and are not prepared for that process should look elsewhere, or should consume ICU through a runtime that already owns the upgrade path.
Finally, ICU is not a translation system. It handles the mechanics of text and locale data. It does not manage message catalogs, translator workflows or string extraction, and nothing in the repository suggests otherwise.
ICU4X and the alternatives worth weighing
The releases list in this repository includes tags under icu4x/, which points to a genuinely different approach to the same problem. ICU4X is a separate implementation of Unicode and locale functionality, and the versioning scheme here (icu4x/2026-08-31/79.x) is distinct from the numbered ICU releases such as release-78.3. Where ICU4C and ICU4J are mature, data-heavy runtimes with long histories, ICU4X is a newer design aimed at modularity and at environments where shipping the full data set is not acceptable. If your concern is payload size or a language outside C++ and Java, that difference in approach is the one to investigate rather than assuming the two are interchangeable.
If you are in Java specifically, the JDK's own java.text and java.time classes are backed by a bundled ICU4J. That is a real alternative for many applications: you get locale-aware formatting without adding a dependency, at the cost of being tied to whatever ICU version your JDK shipped with. The difference is control. Depending on ICU4J directly lets you pin the version and get newer data; relying on the JDK means you inherit the JDK's upgrade schedule.
For C and C++, the alternative is the platform's locale machinery. It is smaller and always present, and it is also inconsistent across operating systems and limited in collation and segmentation. The honest framing is that ICU exists because that machinery was not enough, and the trade you are making is size and build complexity in exchange for behaviour that matches the Unicode standard.
Maintenance, licensing and what to check before pinning
The repository is not archived, and the last push was on 2026-09-22, so work is ongoing. The most recent numbered release listed is ICU 78.3 from 2026-03-17, with ICU4X tags appearing more frequently through mid-2026. That cadence matters for planning: if you depend on ICU4C or ICU4J, your upgrade cost is dominated by data changes and by revalidating formatting and collation output, not by API churn.
Licensing needs care rather than assumption. The repository reports NOASSERTION for its license, and the README points to the Unicode Terms of Use at unicode.org/copyright.html together with a LICENSE file in the repository root. The README also notes that Unicode and the Unicode Logo are registered trademarks of Unicode, Inc. Because the license identifier is not machine-classified, treat the LICENSE file and the Terms of Use as the authoritative sources and have your own process confirm they fit your distribution model. This is a description of what the repository states, not legal advice.
Upgrade cost has one more component the README does not address: the data you ship. Moving to a new ICU release generally means regenerating and redistributing the data payload, which affects package size and, for embedded targets, may affect whether the upgrade is feasible at all. Verify that before you pin, not after.
Editorial conclusion
Adopt ICU when you need locale-correct collation, number and date formatting, or Unicode character properties that your platform's runtime does not provide, and when you can accept a large data payload and a real build step. Do not adopt it for a small utility that only needs UTF-8 validation, and do not assume the JDK's bundled ICU4J matches the current release. Before committing, verify which release tag you are pinning, whether your platform already ships an ICU whose version you must match, and how much of the data you can trim.
Frequently asked questions
What is Unicode ICU?
ICU is the International Components for Unicode, a project under the stewardship of the Unicode Consortium. This repository holds its source, split into ICU4C for C and C++ and ICU4J for Java, providing Unicode and locale behaviour such as collation, formatting and character properties.
What is ICU4?
The repository does not use the name ICU4. It contains ICU4C, the C and C++ library, and ICU4J, the Java library, plus ICU4X tags in its release list. If you have seen ICU4 written elsewhere, the repository itself does not define it.
What is ICU in programming?
In programming, ICU is the library that supplies Unicode and locale data: sorting rules, number and date formatting, character properties and text segmentation. The README directs users to the user guide at unicode-org.github.io/icu for documentation.
How can I use Unicode characters in C++ with ICU?
The README does not print a code sample, so the route it gives is documentation rather than a snippet: API docs at unicode-org.github.io/icu-docs and the user guide at unicode-org.github.io/icu. For the C and C++ library specifically, the README links icu4c/readme.html, which is where the build and usage details for that subdirectory live.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/unicode-org-icu)