# Juman++: An RNNLM-Powered Morphological Analyzer for Japanese

> Juman++ (jumanpp) is a C++ morphological analyzer from Kyoto University that uses a recurrent neural network language model to score the semantic plausibility of word segmentation candidates. Version 2 delivers more than 250 times the analysis speed of the original. It is aimed at NLP researchers and engineers who process Japanese text and need accurate word boundary detection and part-of-speech tagging.

**ku-nlp/jumanpp** — Juman++ (a Morphological Analyzer Toolkit)

- Repository: https://github.com/ku-nlp/jumanpp
- Website: https://nlp.ist.i.kyoto-u.ac.jp/index.php?JUMAN%2B%2B
- Stars: 413 · Forks: 47
- Language: C++
- License: Apache-2.0
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/ku-nlp-jumanpp

## What Juman++ Does and Who It Is For

Morphological analysis is the process of segmenting a sentence into its constituent words and labeling each word with its dictionary form, reading, part of speech, and subtype. Japanese text presents a specific challenge because word boundaries are not marked by spaces, so the analyzer must infer them from context.

Juman++ addresses this by combining a dictionary-based lattice of possible word segmentations with a recurrent neural network language model that scores each segmentation path by its semantic plausibility. The model favors segmentations where the resulting word sequence makes linguistic sense, not just sequences that fit the dictionary.

The tool is aimed at NLP researchers and engineers who build Japanese text processing pipelines. Its output format is compatible with KNP, a dependency structure analyzer from the same Kyoto University NLP laboratory, and Python example scripts for both tools are included in the sample/ directory.

Version 2, described in a 2018 ANLP paper, improved accuracy over the original and raised analysis speed by more than 250 times. The latest formal release is v2.0.0-rc4, published on 2023-10-03.

## Building and Installing Juman++

The recommended installation path is from a release package rather than from the git repository. The git version does not include the pretrained model; only the package distribution does. The README warns that a package download should be around 300 MB and that a smaller download indicates a source snapshot without the model.

On Ubuntu 22.04, install the required system packages first:

```bash
sudo apt install libprotobuf-dev protobuf-compiler
```

For other Linux distributions and CentOS/RHEL, the docs/building.md file provides additional instructions. Then extract the package and build:

```bash
tar xf jumanpp-<version>.tar.xz
cd jumanpp-<version>
cmake -S . -B build \
    -DCMAKE_BUILD_TYPE=Release \
    -DCMAKE_INSTALL_PREFIX=<prefix>
cmake --build build -j<parallelism>
cmake --install build
```

For best performance on local hardware, the README recommends building with extended CPU instruction sets:

```bash
cmake -S . -B build -DCMAKE_CXX_FLAGS="-march=native"
```

This enables FMA and BMI extensions and is documented as working best on Intel Haswell and newer processors. The Docker path uses a multi-stage Alpine build; the Dockerfile in the repository root handles the full compile and produces a minimal runner image with jumanpp as the entrypoint.

System requirements are a C++14-compatible compiler (gcc 5.1 or later, clang 3.4 or later, or MSVC 2017), and CMake 3.13 or later. The project builds and tests on Linux and macOS using GCC and clang, and on Windows using MinGW64-gcc and MSVC2017.

## Running the Analyzer and Reading the Output

Juman++ reads UTF-8 encoded text from stdin, one sentence per line. Lines beginning with '# ' are treated as comments and passed through. The analyzer writes one word per line followed by a closing EOS marker.

A typical run from the README:

```bash
echo "魅力がたっぷりと詰まっている" | jumanpp
```

This produces a tab-separated output with the surface form, reading, dictionary form, part of speech, conjugation information, and a semantic category or auto-recognition flag for each word. The EOS line marks the end of the sentence.

The main command-line options include `--beam` to set the local beam width used in analysis, `--specifics` for lattice format output, and `--model` to specify a non-default model file. The full option list is available via `--help`.

Two Python sample files are included: sample/python_juman.py for parsing Juman++ output and sample/python_knp.py for KNP output. These demonstrate how to integrate the analyzer's output into a Python processing pipeline without reimplementing the output parser.

## Beam Configuration, Path Diffs, and Partial Annotation

The beam width controls how many segmentation candidates the analyzer maintains during analysis. A wider beam increases the chance of finding the best segmentation but also increases computation time. The default beam width is 5.

For researchers who want to study where different beam configurations produce different analyses, the repository builds a binary called jpp_jumandic_pathdiff. The README describes it as finding sentences where two different beam configurations produce different results. Use it as:

```bash
jpp_jumandic_pathdiff <model> <input> > <output>
```

The output is in partial annotation format. The full beam result is written as actual tags and the trimmed beam result is written as comments, with a score header showing both path scores.

A partial annotation tool is also available at a separate repository linked from the README. Partial annotation lets researchers mark correct analyses for ambiguous segments without annotating the full sentence, which reduces the annotation cost for building training data.

## Using Juman++ as a General Morphological Analyzer

The README states explicitly that Juman++ is a general tool that does not depend on Jumandic or the Japanese language, though some Japanese-specific functionality exists. The framework can be adapted to any language or encoding problem where word boundaries are not explicit in the input text.

The repository links to a tutorial project (jumanpp-t9) that implements something similar to T9 predictive text input using the Juman++ framework. T9 input involves disambiguating sequences of numeric key presses into words, which shares the same lattice-and-scoring structure as morphological analysis.

Adapting Juman++ to a new language requires training a new model and building a new dictionary. The training scripts for the Jumandic model are in a separate repository (jumanpp-jumandic). Retraining the Jumandic model specifically requires access to the Mainichi Shinbun corpus for 1995, which has restricted access. A custom language does not need that corpus but does need its own annotated training data.

## Limitations and Constraints

Juman++ handles only UTF-8 encoded text. Any input in a different encoding must be converted before processing. There is no built-in encoding detection or conversion.

The pretrained model is tied to a specific release. The README warns that the current git version is not compatible with the models from v2.0.0-rc1 and v2.0.0-rc2. Teams upgrading from an old release need to verify model compatibility before replacing their binary.

The training corpus for Jumandic (the standard Japanese dictionary model) requires the Mainichi Shinbun corpus for 1995 through the Kyoto University corpus. This corpus is not freely available; access requires an agreement with the rights holder. Teams that cannot obtain the corpus cannot retrain the standard model.

Building from the git repository produces the binary but not the analysis model. This is a common point of confusion: the git clone is useful only for developing the analyzer itself, not for running it on text.

The latest formal release is v2.0.0-rc4, dated 2023-10-03. The last push to the repository was on 2026-04-17, which is within six months of today.

## Research Background and Licence

Juman++ is backed by published NLP research. The core RNNLM approach was introduced in a 2015 EMNLP paper by Morita, Kawahara, and Kurohashi. Version 2 improvements were described at ANLP 2018 and then in an EMNLP 2018 paper on handling scriptio continua. A later journal paper in the Journal of Natural Language Processing covers the design and structure of the toolkit. These publications establish the method's academic basis and provide detail on the model architecture that the README summarizes.

The repository is licensed under Apache-2.0. This permits use in commercial products, modification, and redistribution, provided attribution and the licence notice are preserved. The upstream dependency on protobuf carries its own licence; teams should verify compatibility with their project's licence requirements.

The project is maintained by the KU-NLP group at Kyoto University. The repository includes a CITATION.cff file for academic citation, which is the correct file to reference when citing the tool in a paper.

## Conclusion

Juman++ is the right tool for NLP practitioners working on Japanese text who need RNNLM-based morphological analysis backed by published research. It is a poor fit for teams working in languages other than Japanese unless they have the resources to train a new model and dictionary from scratch: the README describes that path as possible but requires a custom corpus. Verify before building that protobuf is installed on Ubuntu 22.04 (`sudo apt install libprotobuf-dev protobuf-compiler`), and download the package release rather than the git source, since only the package contains the pretrained model.

## FAQ

### Why is the Juman++ package download 300 MB when the repository source is much smaller?

The package from the Releases page includes the pretrained RNNLM model, which is the main analysis artifact. The git repository contains only source code and build scripts; it does not include the model. The README warns that a download significantly smaller than 300 MB indicates a source snapshot without the model, which cannot be used for analysis.

### Can Juman++ be used for languages other than Japanese?

According to the README, Juman++ is a general tool that does not depend on Jumandic or the Japanese language. The T9 tutorial project linked from the README demonstrates adapting the framework to a numeric-key prediction problem. Using it for a new language requires building a new dictionary and training a model on annotated data for that language.

### What corpus is needed to retrain the Jumandic model for Juman++?

The README states that retraining the Jumandic model requires access to the Mainichi Shinbun newspaper corpus for 1995, used through the Kyoto University corpus. Training scripts are available in a separate repository (jumanpp-jumandic). This corpus has restricted access; teams without it cannot retrain the standard Japanese dictionary model.

## Sources

- [ku-nlp/jumanpp on GitHub](https://github.com/ku-nlp/jumanpp)
- [License: Apache-2.0](https://github.com/ku-nlp/jumanpp/blob/master/LICENSE)
- [Project website](https://nlp.ist.i.kyoto-u.ac.jp/index.php?JUMAN%2B%2B)
- [README](https://github.com/ku-nlp/jumanpp/blob/master/README.md)
- [Releases](https://github.com/ku-nlp/jumanpp/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/ku-nlp-jumanpp
