Open-source project
infinilabs/analysis-pinyin avatar
infinilabs/analysis-pinyin

analysis-pinyin: Chinese to Pinyin Analysis for Elasticsearch and OpenSearch

🛵 This Pinyin Analysis plugin is used to do conversion between Chinese characters and Pinyin.

3,095 stars550 forksJavaApache-2.0

At a glance

What is it?
A Java analysis plugin that adds a pinyin analyzer, tokenizer and token filter to Elasticsearch and OpenSearch, so Chinese names and text can be matched by romanised input. The install path is version-specific, and the default settings trade offset accuracy for token overlap.
Who is it for?
Adopt analysis-pinyin if you run Elasticsearch or OpenSearch and need Chinese names or short text matched by pinyin, initials such as ldh, or mixed Latin and Chinese input. Skip it if you need accurate highlighting and position-based queries on the same field, because the default ignore_pinyin_offset setting permits overlapping tokens and the README states that position-related queries and highlights become incorrect.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 142 days ago.
What is it written in?
Mainly Java, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The matching problem analysis-pinyin solves

Chinese text in a search index is not searchable by the way users actually type. A user looking for 刘德华 may type liu, ldh, or liudehua, and a standard analyzer treats those as unrelated strings. analysis-pinyin closes that gap by producing Pinyin tokens alongside the original characters at index time. The README describes the plugin as facilitating conversion between Chinese characters and Pinyin, and it ships three components under one name: an analyzer named pinyin, a tokenizer named pinyin, and a token filter named pinyin. The audience is narrow and specific: engineers running Elasticsearch or OpenSearch clusters that hold Chinese content, who need romanised search without adding a separate romanisation service in front of the cluster. The topics list confirms the supported engines: elasticsearch, opensearch, and easysearch.

What the pinyin tokenizer actually emits

The plugin works at the analysis chain level, not as a query rewriter. You register a tokenizer of type pinyin, point an analyzer at it, and the tokens it produces are what the inverted index stores. The README's worked example analyzes 刘德华 and returns five tokens in order: liu, de, hua, the original 刘德华, and ldh. That ordering matters, because each token carries a position, and the offsets in the example map back to the source characters: liu covers offsets 0 to 1, de covers 1 to 2, hua covers 2 to 3, while the original and the initials both span 0 to 3. The token set is controlled by the optional parameters, which are all booleans or small integers. keep_full_pinyin defaults to true and yields the per-character Pinyin; keep_first_letter defaults to true and yields the collapsed initials; keep_original defaults to false and must be switched on to retain 刘德华 itself. Because several tokens can cover the same character range, the plugin exposes ignore_pinyin_offset, which defaults to true. The README is direct about the cost: after version 6.0 offsets are strictly constrained and overlapped tokens are not allowed, so this parameter allows the overlap by ignoring the offset, and all position-related queries or highlights become incorrect. That is the central design trade-off in the plugin, and it is not hidden.

Installing analysis-pinyin and running a first query

The README points to https://release.infinilabs.com/ for packaged downloads, and also gives the plugin CLI route. The install URL embeds the target engine version, so the version segment must match your cluster. For Elasticsearch the README shows:

bash
bin/elasticsearch-plugin install https://get.infini.cloud/elasticsearch/analysis-pinyin/8.4.1

For OpenSearch the equivalent command is:

bash
bin/opensearch-plugin install https://get.infini.cloud/opensearch/analysis-pinyin/2.12.0

The README adds a tip to replace the version number with the one matching your Elasticsearch or OpenSearch installation. After a restart, create an index with a custom tokenizer. This snippet is the README's own configuration, trimmed to the tokenizer block:

json
PUT /medcl/
{
    "settings" : {
        "analysis" : {
            "analyzer" : {
                "pinyin_analyzer" : {
                    "tokenizer" : "my_pinyin"
                    }
            },
            "tokenizer" : {
                "my_pinyin" : {
                    "type" : "pinyin",
                    "keep_separate_first_letter" : false,
                    "keep_full_pinyin" : true,
                    "keep_original" : true,
                    "limit_first_letter_length" : 16,
                    "lowercase" : true,
                    "remove_duplicated_term" : true
                }
            }
        }
    }
}

Note that remove_duplicated_term is set to true here even though its documented default is false, and the README warns that it can influence position-related queries. To confirm the tokens, call the analyze endpoint on the same text the README uses:

json
GET /medcl/_analyze
{
  "text": ["刘德华"],
  "analyzer": "pinyin_analyzer"
}

The response should contain the five tokens described earlier: liu, de, hua, the original characters, and ldh. From there the README maps a name field as a keyword with a pinyin sub-field using term_vector with_offsets, indexes a document, and searches with queries such as name.pinyin:liu or name.pinyin:ldh.

Non-Chinese text, duplicated terms and the position trap

Mixed input is where the parameter set gets fiddly. keep_none_chinese defaults to true and retains Latin letters and digits; keep_none_chinese_together defaults to true and keeps them as one unit, so DJ音乐家 becomes DJ, yin, yue, jia. Set it to false and the same input becomes D, J, yin, yue, jia. The README notes that keep_none_chinese must be enabled first for that flag to mean anything. A separate parameter, none_chinese_pinyin_tokenize, defaults to true and splits Latin strings that happen to look like Pinyin, so liudehuaalibaba13zhuanghan becomes liu, de, hua, a, li, ba, ba, 13, zhuang, han. That behaviour is helpful for pasted usernames and surprising for product codes. Two flags carry explicit warnings. remove_duplicated_term, off by default, collapses de的 into de to save index space, with the README noting position-related queries may be influenced. ignore_pinyin_offset, on by default, is the one to think hardest about: the README says to use multi-fields with different settings for different query purposes, and to set it to false if you need offsets. In practice that means one sub-field for recall-oriented matching and another for highlighting, rather than one field trying to do both.

Where analysis-pinyin is the wrong tool

The plugin is an analysis component, not a Chinese segmentation library. It does not decide where words begin and end in running prose; it romanises what the tokenizer is handed. The README's own token filter example pairs the pinyin filter with a whitespace tokenizer over space-separated names, which is a realistic hint about intended use: names, tags, and short fields, not long documents. If your requirement is phrase search with accurate highlights, the default ignore_pinyin_offset setting works against you, and the README says so plainly. If you need tone marks or tone-aware ranking, nothing in the documented parameter list provides them; the flags govern token shape, not phonetic detail. And if your cluster is not Elasticsearch, OpenSearch or Easysearch, this plugin does not apply at all, since it installs through the engine's own plugin CLI. The repository layout, with separate elasticsearch/, opensearch/ and pinyin-core/ directories, reflects that the plugin is built per engine rather than as a standalone library.

Alternatives and how they differ

The related searches point at analysis-stconvert, a sibling INFINI Labs plugin, and the difference is the transformation itself. analysis-stconvert handles Chinese character conversion between Simplified and Traditional forms; analysis-pinyin handles character-to-romanisation. They solve adjacent problems and are often used together, but neither substitutes for the other. Outside that family, the common alternative is to romanise upstream: convert Chinese fields to Pinyin in your application or an ingest pipeline, store the romanised text in a plain text field, and let the built-in analyzers do the rest. That approach keeps offsets and highlighting predictable because the index only ever sees Latin characters, but it moves the conversion out of the cluster, which means reindexing whenever the conversion rules change and keeping the converter in sync across every writer. The plugin's advantage is that conversion happens inside the analysis chain, so the same index accepts Chinese input and Pinyin queries without a pre-processing step. Its disadvantage is exactly the offset behaviour described above. Choose based on whether you would rather own the conversion code or accept the plugin's token overlap.

Maintenance, versions and licence

The repository is not archived, and the last push was on 2026-05-11. The most recent release listed is dated 2024-05-06, so the gap between the last release and the last commit is roughly two years. The README states the plugin supports major versions of Elasticsearch and OpenSearch, and the install examples reference 8.4.1 and 2.12.0, which tells you the maintainers publish builds per engine version rather than one universal artifact. That is the upgrade cost: a cluster major-version upgrade likely means a different plugin URL and a rebuild or reinstall, and index mappings that reference the pinyin tokenizer will need the plugin present before the index can be opened. The project is licensed Apache-2.0, which permits commercial use and modification with the usual attribution and notice requirements; the LICENSE.txt file sits at the repository root. This is a statement of what the licence is, not legal advice, and anyone embedding the plugin in a distributed product should read the licence text and their own obligations. The README does not document a rollback procedure for the plugin, and it does not list a compatibility matrix beyond the per-version download URLs.

Editorial conclusion

Adopt analysis-pinyin if you run Elasticsearch or OpenSearch and need Chinese names or short text matched by pinyin, initials such as ldh, or mixed Latin and Chinese input. Skip it if you need accurate highlighting and position-based queries on the same field, because the default ignore_pinyin_offset setting permits overlapping tokens and the README states that position-related queries and highlights become incorrect. Before installing, confirm the plugin build matches your exact Elasticsearch or OpenSearch version, since the README warns you to replace the version number in the install URL, and decide per field whether to leave ignore_pinyin_offset at its default or set it to false and use multi-fields.

Frequently asked questions

What does pinyin mean in Chinese?

In this project the term names the romanisation the plugin produces: the README converts Chinese characters such as 刘德华 into tokens like liu, de, hua and the initials ldh. The plugin's job is that conversion between Chinese characters and Pinyin, exposed as an analyzer, tokenizer and token filter.

How do I convert Mandarin to Pinyin?

With analysis-pinyin you register a tokenizer of type pinyin in the index settings and point an analyzer at it, then the analysis chain emits the Pinyin tokens. The README's example analyzes 刘德华 and returns liu, de, hua, the original characters, and ldh.

What does the pinyin analyzer return for a Chinese name?

The README analyzes 刘德华 and gets five tokens: liu, de, hua, the original characters 刘德华, and the initials ldh. Which of those appear depends on the parameters, since keep_full_pinyin and keep_first_letter default to true while keep_original defaults to false.

Why are highlights or position queries wrong with analysis-pinyin?

The ignore_pinyin_offset parameter defaults to true, which allows tokens that overlap the same character range by discarding offset information. The README states that all position-related queries or highlights become incorrect as a result, and recommends multi-fields with different settings for different query purposes, or setting it to false if you need offsets.

Official sources

  1. infinilabs/analysis-pinyin on GitHub
  2. Issues
  3. License: Apache-2.0
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/infinilabs-analysis-pinyin.svg)](https://hysenlabs.com/projects/infinilabs-analysis-pinyin)