Model or dataset
togatoga/karukan avatar
togatoga/karukan

Karukan: a Rust IME workspace with a Swift frontend sitting outside it

Japanese Input Method System for Linux, macOS, Neural Kana-Kanji Conversion Engine

734 stars57 forksRustApache-2.0

At a glance

What is it?
A Japanese input method for Linux and macOS that turns kana into kanji with a GPT-2 or Qwen3 model run through llama.cpp. The layout is unusually legible for an IME: four Rust crates in one workspace, a JSON-RPC server for macOS, an fcitx5 addon with a C FFI, and a CLI that also builds dictionaries and serves HTTP.
Who is it for?
Karukan is a real input method rather than a demo, and its documentation is unusually specific about the parts that usually go undocumented, namely chunk boundaries, persona keywords and user dictionary formats. The gaps are in packaging and licensing clarity: one tag from February, install instructions that live in two other READMEs, and three licence layers while the repository metadata names only one.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly Rust, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Four crates in the workspace, and a comment that still says three

The Rust workspace declares four members: karukan-engine, karukan-cli, karukan-im/core and karukan-im/fcitx5. The core crate is the shared engine, holding the state machine, romaji conversion and karukan-imserver, a JSON-RPC server for macOS. The fcitx5 crate is the Linux frontend, an fcitx5 addon behind a C FFI. karukan-engine is the conversion library, romaji to hiragana plus neural conversion through llama.cpp. karukan-cli holds the dictionary build, Sudachi dictionary generation, a dictionary viewer, AJIMEE-Bench and an HTTP server.

The macOS frontend is not a member. It lives at karukan-im/macos as Swift and InputMethodKit, so it is built with a different toolchain, and the four CI workflows match that split: karukan-engine-ci, karukan-im-ci, karukan-fcitx5-ci and karukan-macos-ci.

One detail is worth noting. The comment above the shared dependency block reads Shared across all 3 crates, above serde, anyhow, tracing and tracing-subscriber, while the members list at the top of the same file declares four. Either the comment predates the fourth member or one member does not use those four crates.

The first launch downloads a model while the keyboard keeps working

Neural conversion needs a model, and the model comes from Hugging Face on first launch. The project is explicit about the failure shape here, and it is a good one: the download runs in the background, kana input and dictionary conversion remain usable while it is in flight, and neural conversion switches itself on once the model finishes loading. After that first download the network is not used again, and the local model is what answers.

The models are GPT-2 and Qwen3 based, run through llama.cpp, which keeps the inference path in the same binary stack as the rest of the toolchain. The system dictionary is built from SudachiDict data rather than shipped precompiled, and docs/dictionary.md covers installing it, user dictionaries and candidate priorities.

So the first run of this IME is a wait, but not a blank screen, and the degraded mode during that wait is a plain dictionary converter rather than an error.

Live conversion removes the space bar, and chunking puts the boundary back

Live conversion shows candidates while you type, without pressing Space, and Ctrl+Shift+L toggles it. That single toggle changes the feel of the keyboard more than any other setting, because it changes when conversion happens rather than what conversion produces.

The more interesting documentation is the one that admits conversion has a seam. docs/chunking.md explains where a conversion gets split into chunks and describes how to split it yourself so a fixed display boundary is respected. For a neural converter that reads surrounding context, the chunk boundary decides how much left context the model actually sees, and the project treats that as a knob rather than an implementation detail. docs/symbols.md covers the same kind of ground on the output side, covering punctuation and bracket styles, the width of symbols, digits and letters, and space behaviour. config.toml holds the settings for live conversion, the conversion strategy and the learning cache.

Two separate mechanisms bias the output, one learned and one typed by hand

Conversion learning records the candidates you pick and promotes them in later conversions, and it works mid-typing as well as at commit: predictive conversion on forward match offers already learned candidates before you finish the word.

The second mechanism is the conversion persona. You write keywords related to the sentences you often produce into config.toml, and conversions skew toward that vocabulary. docs/persona.md covers it. Between them, one bias comes from your own history and the other from a file you edit by hand, which is a useful distinction: the learned cache is data you did not write, and the persona is a vocabulary you chose.

A third path bypasses both. docs/user-dictionary.md documents user dictionaries in Mozc and Google IME TSV form and in binary form, so an existing dictionary can be carried over instead of relearned.

A Mozc port that annotates every candidate it generates

The candidate rewriter is ported from Mozc and generates the variants a Japanese IME is expected to produce: half-width katakana, letter case changes, full-width and half-width forms, related symbol candidates, and number notations including kanji numerals, 大字, Roman numerals, circled digits and base 16, 8 and 2.

What is unusual is that each generated candidate keeps the annotation it inherited from Mozc, so a candidate that came from the half-width katakana rule is labelled as such and one produced by the base 16 rule says so. A user can therefore see why a candidate is on the list rather than guessing at it.

Emoji are handled on two routes: kana readings resolve to pictographs, and Slack style trigger queries work as well. The documentation gives examples of both, where a reading such as ぴえん and a query such as :smile each produce a pictograph.

Three licence layers while the metadata names one

The project offers a dual licence, MIT OR Apache-2.0, and the root carries both LICENSE-MIT and LICENSE-APACHE. The repository metadata names Apache-2.0 on its own, so the recorded licence is one of the two offered choices rather than the choice itself, and a consumer reading only that field would not know that MIT is also available.

The third layer is narrower and matters most for anyone redistributing. Data under karukan-engine/data/ is derived from Mozc and is distributed under BSD 3-Clause. The provenance of each derived file and Mozc's copyright notices are collected in THIRD_PARTY_LICENSES, which is where a distributor starts.

So there is no contradiction to resolve here, only three answers to three different questions: the code is MIT or Apache-2.0 by your choice, the metadata names Apache-2.0, and the derived dictionary data is BSD 3-Clause whatever you pick for the code.

Install is two links and no command, and the only tag is seven months old

The root README carries no install command of any kind. Installation is two pointers: Linux users are sent to the karukan-fcitx5 README, specifically its install anchor, and macOS users to the karukan-macos README. Everything a new user needs is therefore one level down, split by platform, and the root document shows none of it.

The rest of the root is automation rather than documentation: renovate.json for dependency updates, four CI workflows, .dockerignore and a scripts/ directory, a tests/ directory, and a CLAUDE.md at the top level alongside Cargo.lock and Cargo.toml. The workspace pins edition 2024, rust-version 1.92 and resolver 3, with thin LTO in the release profile.

Releases are thinner still. One tag exists, v0.1.0, published 2026-02-23, while the default branch was pushed on 2026-10-02. Anyone installing from a release therefore gets a February snapshot unless they build from the branch.

Editorial conclusion

Karukan is a real input method rather than a demo, and its documentation is unusually specific about the parts that usually go undocumented, namely chunk boundaries, persona keywords and user dictionary formats. The gaps are in packaging and licensing clarity: one tag from February, install instructions that live in two other READMEs, and three licence layers while the repository metadata names only one. Before adopting it, read the platform README for the install path, check which tagged release you actually get, and confirm which licence you need for the Mozc derived data.

Frequently asked questions

What happens the first time I start Karukan?

It downloads a GPT-2 or Qwen3 based model from Hugging Face in the background. Kana input and dictionary conversion stay usable during the download, neural conversion switches on automatically once the model is loaded, and the network is used only for that first download.

How do I turn Karukan's live conversion on and off?

Live conversion displays candidates while you type without pressing Space, and Ctrl+Shift+L toggles it. config.toml holds the settings for live conversion, the conversion strategy and the learning cache, as described in docs/configuration.md.

Which platforms does Karukan support and how do I install it?

Linux through an fcitx5 addon behind a C FFI, and macOS through a Swift and InputMethodKit frontend. The root README has no install command; it links to karukan-im/fcitx5/README.md for Linux and karukan-im/macos/README.md for macOS.

Which licence covers Karukan and its dictionary data?

The project code is dual licensed MIT OR Apache-2.0, with both licence files at the repository root, while the repository side of the metadata names Apache-2.0 only. Data under karukan-engine/data/ is derived from Mozc and is distributed under BSD 3-Clause, with per-file provenance in THIRD_PARTY_LICENSES.

Can I bring my own dictionary into Karukan?

Yes. docs/user-dictionary.md covers registering user dictionaries in Mozc or Google IME TSV form and in binary form. The system dictionary is built from SudachiDict data, and docs/dictionary.md covers installing it and setting candidate priorities.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. README
  4. Releases
  5. togatoga/karukan on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/togatoga-karukan.svg)](https://hysenlabs.com/projects/togatoga-karukan)