# IK Analysis for Elasticsearch and OpenSearch: Chinese Segmentation With a Custom Dictionary

> The IK Analysis plugin wraps the Lucene IK analyzer and exposes it as ik_smart and ik_max_word in Elasticsearch and OpenSearch. It is the standard way to get usable Chinese tokenization plus a dictionary you control.

**infinilabs/analysis-ik** — 🚌 The IK Analysis plugin integrates Lucene IK analyzer into Elasticsearch and OpenSearch, support customized dictionary.

- Repository: https://github.com/infinilabs/analysis-ik
- Stars: 17,518 · Forks: 3,268
- Language: Java
- License: Apache-2.0
- Published: 2026-09-21 · Updated: 2026-09-21 · Language: en
- Canonical page: https://hysenlabs.com/projects/infinilabs-analysis-ik

## The problem analysis-ik solves, and who has it

Chinese text has no spaces between words. A general-purpose tokenizer that splits on whitespace and punctuation treats a whole Chinese sentence as one token, which makes a term query match nothing and makes relevance scoring meaningless. The IK Analysis plugin solves this by integrating the Lucene IK analyzer into Elasticsearch and OpenSearch, so Chinese text is segmented into words at index time and at query time. The audience is narrow and specific: engineers running an Elasticsearch or OpenSearch cluster that stores Chinese content and who need control over how that content is split. The plugin ships two analyzers and two tokenizers, both named ik_smart and ik_max_word, which means the same names appear in mapping definitions and in analyzer configuration. The repository also lists easysearch under its topics, and the README states that major versions of Elasticsearch and OpenSearch are supported. The plugin is Apache-2.0 licensed and maintained by INFINI Labs.

## ik_max_word versus ik_smart, and why the difference matters at query time

The two analyzers differ in granularity, and the README is explicit about the trade-off. ik_max_word performs the finest-grained segmentation, and the README's own example splits 中华人民共和国国歌 into 中华人民共和国, 中华人民, 中华, 华人, 人民共和国, 人民, 人, 民, 共和国, 共和, 和, 国国, 国歌. That exhaustive output is the point: it generates every plausible combination, which suits term queries. ik_smart performs the coarsest-grained segmentation, and the same input becomes 中华人民共和国, 国歌, which suits phrase queries. The README adds a warning that is easy to miss: ik_smart is not a subset of ik_max_word. You cannot assume that anything ik_smart produces will also be present in ik_max_word output. The conventional mapping in the README uses ik_max_word as the index analyzer and ik_smart as the search analyzer, which is a deliberate asymmetry: index broadly so terms exist, then query narrowly so the query does not over-match. Changing either side later requires reindexing, because the tokens stored in the index were produced by the analyzer configured at write time.

## Installing the plugin and running a first query

The README gives two routes. You can download packaged plugins from the release site at release.infinilabs.com, or install through the plugin CLI against a versioned URL. The CLI form for Elasticsearch, exactly as the README shows it, is:

```bash
bin/elasticsearch-plugin install https://get.infini.cloud/elasticsearch/analysis-ik/9.1.4
```

For OpenSearch the equivalent command uses the OpenSearch plugin binary and a different version path:

```bash
bin/opensearch-plugin install https://get.infini.cloud/opensearch/analysis-ik/2.12.0
```

The README carries a tip that the version number in the URL must match your Elasticsearch or OpenSearch version. That is not decoration. Treat the two URLs above as templates, not as the versions you should install.

Once installed, create an index and give the field an analyzer pair. The README's mapping example is:

```bash
curl -XPOST http://localhost:9200/index/_mapping -H 'Content-Type:application/json' -d'
{
        "properties": {
            "content": {
                "type": "text",
                "analyzer": "ik_max_word",
                "search_analyzer": "ik_smart"
            }
        }

}'
```

The README then indexes documents with curl and runs a match query on 中国 with a highlight block using pre_tags and post_tags. In the README's result, two documents match and the highlight wraps the matched term in the configured tags. If you run the same query and get no highlight, the likely cause is that the field was mapped before the plugin was installed, so the tokens in the index were produced by a different analyzer.

## Dictionary configuration and hot reload over HTTP

Custom dictionaries are the reason many teams pick this plugin over a stock analyzer. The configuration file is IKAnalyzer.cfg.xml, and the README places it at {conf}/analysis-ik/IKAnalyzer.cfg.xml or at {plugins}/elasticsearch-analysis-ik-*/config/IKAnalyzer.cfg.xml. It holds four keys: ext_dict and ext_stopwords for local files, and remote_ext_dict and remote_ext_stopwords for URLs. The README's example points ext_dict at custom/mydict.dic and custom/single_word_low_freq.dic.

The hot reload mechanism is the part worth understanding before you design around it. The README states that the location value is a URL such as http://yoursite.com/getCustomDict, and that the HTTP response must carry two headers: Last-Modified and ETag, both strings. If either changes, the plugin fetches the new word list. The body must be one word per line, with \n as the newline. The README suggests placing a UTF-8 .txt file under nginx or another simple HTTP server, since those servers emit Last-Modified and ETag when the file changes. Two constraints follow. First, a static file server that does not emit both headers will not trigger an update, so the reload silently does nothing. Second, the reload depends on a reachable HTTP endpoint, which is an availability dependency inside your cluster's startup and refresh path.

## Where analysis-ik is the wrong choice

The plugin is built for Chinese. If your corpus is English, Spanish or another space-delimited language, the standard analyzers already tokenize correctly, and adding IK gives you a second analyzer to reason about for no gain. The README does not claim general multilingual coverage, and nothing in the repository layout suggests language-specific resources beyond the IK dictionaries.

The dictionary reload path is the second limitation. Hot reload requires an HTTP service that returns Last-Modified and ETag, and the README does not document a fallback when that service is unreachable or when a proxy strips those headers. The README also does not document rollback: there is no described procedure for reverting a dictionary change that produced bad tokens, and because the index stores tokens produced at write time, a dictionary change does not retroactively fix existing documents. You reindex, or you live with a split between old and new segmentation.

The third issue is version coupling. The install URLs are versioned, and the README's tip says the version must match your Elasticsearch or OpenSearch version. A cluster upgrade therefore implies a plugin upgrade, and the README does not describe a mixed-version or rolling upgrade path for the plugin itself. The latest release listed on the repository is dated 2024-05-06, while the last push to the default branch was on 2026-06-30, so the codebase moves between releases. Plan the upgrade as a cluster-level event, not a background task.

## How it compares with smartcn, stconvert and analysis-icu

The alternatives people search alongside this plugin take different approaches. Analysis smartcn is the Lucene/Solr Chinese analyzer: it segments Chinese using an HMM-based model trained on a corpus, with no dictionary file for you to edit. That is the core difference. smartcn gives you no ext_dict key and no remote_ext_dict hot reload, so domain terms that the model does not know stay unsplit. If your relevance problems come from product names, internal jargon or new terms, smartcn offers no place to put them, while analysis-ik does.

Analysis stconvert is not a segmentation alternative at all. It converts between Simplified and Traditional Chinese, so it addresses a different problem: matching text written in one script against a query written in the other. You would run it alongside a segmenter, not instead of one.

Analysis-icu is the ICU-based plugin, which handles Unicode normalization, collation and script-aware segmentation across many languages. It is the better fit when Chinese is one of several languages in the same index and you want consistent Unicode handling rather than a tunable Chinese dictionary. The trade-off is the same one as with smartcn: no user-editable dictionary and no documented HTTP hot reload. Choose analysis-ik when dictionary control is the requirement; choose ICU when breadth is.

## Licence, maintenance and upgrade cost

The plugin is licensed under Apache-2.0, and the README states that you may not use the files except in compliance with the License, with the full text at apache.org/licenses/LICENSE-2.0. Apache-2.0 is a permissive licence that includes a patent grant, and it carries no copyleft obligation on your own code. The repository also contains a licenses/ directory and a LICENSE.txt at the top level. This is a description of what the repository states, not legal advice; if you redistribute the plugin inside a product, read LICENSE.txt yourself.

On maintenance: the repository is not archived, and the last push to the default branch was on 2026-06-30. The most recent release listed is dated 2024-05-06. Those two dates are far apart, so the release cadence is slower than the commit activity, and anyone pinning to a published artifact should check whether a build matching their cluster version exists before planning an upgrade. The upgrade cost itself is dominated by the version match requirement and by reindexing when analyzer configuration changes. A dictionary-only change through remote_ext_dict avoids reindexing for new documents, but it does not rewrite tokens already stored.

## Conclusion

Adopt analysis-ik if you index Chinese text in Elasticsearch or OpenSearch and need dictionary control, because the default analyzers will not split Chinese into useful terms. Skip it if your corpus is English or another space-delimited language, where the built-in analyzers already do the work. Before rolling it out, verify that the plugin version matches your cluster version exactly, and confirm that your custom .dic files are UTF-8 encoded, since a wrong encoding is the documented reason a custom dictionary silently fails to take effect.

## FAQ

### How do I install analysis-ik for Elasticsearch?

Use the plugin CLI with a versioned URL, for example bin/elasticsearch-plugin install https://get.infini.cloud/elasticsearch/analysis-ik/9.1.4, or download a packaged plugin from release.infinilabs.com. The README warns that the version number must match your Elasticsearch version.

### What is the difference between ik_max_word and ik_smart in analysis-ik?

ik_max_word performs the finest-grained segmentation and generates every plausible combination, which suits term queries. ik_smart performs the coarsest-grained segmentation, which suits phrase queries, and the README notes that ik_smart is not a subset of ik_max_word.

### Why is my custom dictionary in analysis-ik not taking effect?

The README's first FAQ answer is to check that the custom dictionary text is UTF-8 encoded. The configuration lives in IKAnalyzer.cfg.xml under the ext_dict and ext_stopwords keys.

### How does analysis-ik hot reload a dictionary without restarting Elasticsearch?

Set remote_ext_dict or remote_ext_stopwords to a URL. The HTTP response must include Last-Modified and ETag headers, and if either changes the plugin fetches the new word list, which must be one word per line separated by \n.

### Does analysis-ik work with OpenSearch?

Yes. The README gives a separate install command using bin/opensearch-plugin install with an OpenSearch-specific version path, and states that major versions of Elasticsearch and OpenSearch are supported.

## Sources

- [infinilabs/analysis-ik on GitHub](https://github.com/infinilabs/analysis-ik)
- [Issues](https://github.com/infinilabs/analysis-ik/issues)
- [License: Apache-2.0](https://github.com/infinilabs/analysis-ik/blob/master/LICENSE)
- [README](https://github.com/infinilabs/analysis-ik/blob/master/README.md)
- [Releases](https://github.com/infinilabs/analysis-ik/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/infinilabs-analysis-ik
