Library / SDK
shineware/KOMORAN avatar
shineware/KOMORAN

KOMORAN, a Korean analyzer whose one real distinction is the space

Korean Morphological Analyzer by shineware

320 stars66 forksJavaApache-2.0

At a glance

What is it?
KOMORAN is a Korean morphological analyser written entirely in Java with no external library dependencies, and the feature it actually leads with is the ability to return morphemes that include the following space. The rest of the repository is a Gradle build with an Elasticsearch plugin in it, continuous integration config for a discontinued service, and a README that is nearly half a bibliography.
Who is it for?
KOMORAN suits a Java service that needs Korean tokenisation without taking a third-party dependency tree, especially one that wants to edit its dictionary in version control as text. Check four things first.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Activity is slowing. The repository last received commits 6 months ago.
What is it written in?
Mainly Java, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The one real distinction is the space inside a morpheme

Five of the six claimed features are table stakes for a Java library. The sixth is not.

Pure Java with no external libraries, running in about fifty megabytes of memory, configured as plain text files you can edit directly, and callable by adding one line of source after adding the library: those are all real and all unremarkable.

The sixth is that KOMORAN can analyse at the level of morphemes that include a trailing space. Other Korean analysers return tokens with the space attached or stripped depending on configuration; this one treats the space as part of the unit it returns.

The practical consequence is that two analysers will hand you different token boundaries from the same sentence, and any downstream component that assumes conventional boundaries will behave differently without an obvious error. It also means you do not have to reconstruct spacing, which is a real convenience if your downstream task cares about the surface form.

The demo sentence the README uses is the standard one for this: the country is a democratic republic, which contains a particle that attaches to the preceding word and is the usual place tokenisers disagree.

So the differentiator is real, and it is also a compatibility decision rather than a quality claim.

The newest tag is from 2020 and the branch moved in 2026

The release history and the commit history are six years apart.

The three most recent tags are 3.3.8 in November 2019, described as a dictionary and grammar probability model update, then 3.3.9 in January 2020 described as an irregular processing bug fix, then a 3.4.0 beta in May 2020.

The last push to the default branch is 30 March 2026, and the repository is not archived. So work has continued on master for roughly six years without a tag being cut.

That has two consequences and neither is fatal. For a user installing from a package registry, the newest available artefact is six years old, and none of the model improvements made since are in it. For a user building from the repository, the code is current but the version string says beta, and there is no release note describing what changed between the last tag and the current branch.

The tag messages are also more informative than the version numbers. Both of the last two describe a specific class of change rather than a number, which is the convention a project keeps when releases track model updates rather than API changes.

The coverage badge points at a third account

The repository belongs to an organisation, and the coverage badge belongs to somebody else.

The badge link at the top of the readme targets a personal account with a different name and the same repository name. Everything else in the file, the clone instructions, the citation entry, the documentation links, and the repository listing itself, points at the organisation.

The likely explanation is that the project started under an individual account and moved, which is the same pattern visible in other long-lived open source projects. What survives a move is every hard-coded link that nobody revisited.

For a badge that reports test coverage, the consequence is that clicking it takes you to a different project's coverage report, or to nothing. The badge is the one element of the readme a reader is most likely to click, because it is the only one carrying information about the project's health.

It is also the kind of detail that tells you something about attention. The documentation site, the Slack invite, the Python wrapper, the third-party Python port, and roughly thirty-five academic references are all current and correctly linked. One badge is not.

Continuous integration config for a retired service

The root listing contains both a GitHub directory and a configuration file for a continuous integration service that most projects have moved away from.

The GitHub directory is where workflows and issue templates live. The other file is configuration for a hosted service that many projects stopped using years ago, and the badge that would report its status is not in the readme.

Two configurations for continuous integration can coexist for a long time, especially when the second one is a single file and the migration is not forced by anything breaking. What usually happens is that the newer one is the one being edited and the older one is a historical artefact nobody reads.

What makes it worth mentioning is what the file tells you about the build. Its presence means the project was tested automatically on a service that provisions its own virtual machines, which is a slower and more heavyweight setup than the hosted runners a GitHub workflow would use. Combined with a Gradle wrapper committed to the repository, the picture is of a Java project set up to build identically on any machine, which was the point of the wrapper and remains a reasonable choice.

Half the readme is a list of papers that used it

The readme has six feature bullets and about thirty-five academic references, and the references are longer.

The bibliography is split into domestic and international sections, each organised by year, covering 2019 and 2020. It includes conference papers, journal articles, a doctoral thesis, a preprint, and a working paper. The subjects range widely: automatic petition topic analysis, a sequence-to-sequence analyser built to tolerate neologisms and spacing errors, video recommendation from extracted morphemes, sentiment analysis with a bidirectional recurrent model, named entity recognition for economic analysis, Korean word embeddings with and without Hanja, and a paper arguing that meaning-sound systematicity is also present in Korean.

Two things are notable. The list is bounded at 2020, which is the same year the tags stop, so the bibliography is measuring adoption at the moment of the last release rather than today. And the list includes papers that argue about the analyser itself, such as the one on dictionary expansion from use cases and the one comparing analysers by sentence type, which means this is a bibliography of influence and of argument, not only of use.

For a library whose distinguishing feature is contested token boundaries, having published comparisons is worth more than having many.

Zero external dependencies and a one-line integration

The dependency claim is specific and unusual enough to be worth taking at face value for a moment.

The stated reason there is no dependency problem with external libraries is that only the project's own libraries are used. Combined with the pure Java claim, the effect is that adding the analyser to a Java service brings in a self-contained set of jars with no transitive tree to audit.

The integration claim follows from that: after adding the library, one line of source code is enough to use the analyser. For a component whose output shape differs from its competitors, one line is a strong promise, since most integration cost in text processing is in shape conversion rather than in the call.

The memory figure is the third claim and the most falsifiable: about fifty megabytes, achieved by processing at the character level and by using a trie for the dictionary. Fifty megabytes is a number you can check, and it is the one to check first if you are embedding this in something with a small heap.

The dictionary claim is the one with the longest tail. Dictionaries as plain text files that can be edited directly is exactly what you want for version control, and it is also what makes it obvious that the quality of the output depends entirely on which dictionary files your build ships.

An Elasticsearch plugin in the same Gradle build

The module list is short and one of the modules is not a library.

At the top level there is a core module, an admin module, an Elasticsearch plugin, a documentation directory, and a directory of reStructuredText sources, alongside the Gradle build files and the wrapper for both shell and Windows.

Shipping an Elasticsearch plugin in the same build as the analyser is the interesting decision. An Elasticsearch plugin has its own packaging, its own descriptor, and its own version compatibility with the search engine it plugs into, so the two artefacts have to move at different speeds. Building them together means one command produces both, and one version number covers both, whether or not that is accurate.

The reStructuredText directory suggests the hosted documentation site is generated from this repository rather than maintained separately, which is the other reason to expect the build to carry more than the library.

Two readme files sit at the root, the default one in Korean and the English one with a suffix. That is the reverse of the usual convention, and it matches the language of the citation entry and the documentation links, all of which are Korean or English as the site requires rather than as the readme defaults to.

Editorial conclusion

KOMORAN suits a Java service that needs Korean tokenisation without taking a third-party dependency tree, especially one that wants to edit its dictionary in version control as text. Check four things first. Whether space-inclusive morphemes are what you want, since that output differs from other analysers and you will need to handle the boundaries. Where you get the library from, because the repository's newest tag is six years old and the branch has moved since. Whether you need the Elasticsearch plugin, which lives in the same Gradle build. And whether the dictionary files in your build are current, because the model files are what make the analyser good and the README defers all of that to hosted documentation.

Frequently asked questions

What is KOMORAN and what makes it different from other Korean analysers?

KOMORAN is a Korean morphological analyser written in pure Java with no external library dependencies. Its distinguishing feature is that it can return morphemes that include the trailing space, where other analysers return tokens with different boundary conventions.

How much memory does KOMORAN need?

The README states it can run in about 50MB of memory, using character-level processing and a trie for its dictionary. It is described as usable anywhere Java is installed.

How do I install and use KOMORAN?

The README defers to hosted documentation for both, with separate pages for installation and a three-minute tutorial, plus example pages for analysis, model training, and Spark with Scala. After adding the library, one line of source code is enough to call the analyser.

Is there a Python version of KOMORAN?

The README lists an official Python wrapper maintained by the same authors, and separately a community Python port published by another developer. The official one is the one to prefer, since it is maintained alongside the Java release.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. README
  4. Releases
  5. shineware/KOMORAN on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/shineware-komoran.svg)](https://hysenlabs.com/projects/shineware-komoran)