Open-source project
daac-tools/vibrato avatar
daac-tools/vibrato

daac-tools/vibrato is a Viterbi tokenizer for Japanese that only matches MeCab after -S and -M 24

🎤 vibrato: Viterbi-based accelerated tokenizer

424 stars25 forksRustApache-2.0

At a glance

What is it?
Vibrato is a Rust reimplementation of the MeCab morphological analyzer, tuned for very large dictionaries. Its defaults emit whitespace tokens MeCab drops, its library API refuses the zstd dictionaries its own CLI loads, and the quickstart runs from a clone because no install command exists.
Who is it for?
Choose vibrato when you are tokenizing large Japanese corpora in Rust and need MeCab-shaped output, and only after you have budgeted for building from source, since the quickstart is a clone plus cargo run and no install line exists. Verify the tokenizer flags against your own gold standard, because acknowledged cost tiebreakers can still shift boundaries, and remember the released tag trails the source, so a dictionary built for one is not guaranteed to load in the other.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly Rust, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Whitespace is a token here until you ask for it not to be

MeCab treats ASCII spaces as invisible during analysis, because SPACE is defined in char.def. Vibrato, running on the same ipadic dictionary, hands back a token for each one under its default settings:

code
$ echo 'mens second bag' | cargo run --release -p tokenize -- -i ipadic-mecab-2_7_0/system.dic.zst
mens	名詞,固有名詞,組織,*,*,*,*
  	記号,空白,*,*,*,*,*
second	名詞,固有名詞,組織,*,*,*,*
  	記号,空白,*,*,*,*,*
bag	名詞,固有名詞,組織,*,*,*,*
EOS

The flags that restore MeCab behavior are -S, which indicates if spaces are ignored, and -M 24, which sets the maximum grouping length for unknown words. Two separate knobs, because the whitespace token is not the only difference. In the default run above, second comes back tagged 名詞,固有名詞,組織. MeCab tags that same word 名詞,一般, and with the two flags applied vibrato agrees:

code
$ echo 'mens second bag' | cargo run --release -p tokenize -- -i ipadic-mecab-2_7_0/system.dic.zst -S -M 24
mens	名詞,固有名詞,組織,*,*,*,*
second	名詞,一般,*,*,*,*,*
bag	名詞,固有名詞,組織,*,*,*,*
EOS

Anyone diffing vibrato output against MeCab output on Latin-script text inside Japanese documents will hit both effects at once, and only the pair of flags corrects them together. The value 24 is given without justification in the walkthrough, so treat it as the known-good setting for this dictionary rather than a general constant.

Even with the flags applied, the file concedes a residual difference: there are corner cases where tokenization results in different outcomes due to cost tiebreakers, and that is described as not an essential problem. The word not is doing work there. It is an acknowledgment that the tie-breaking rule is not identical to MeCab's, so a downstream pipeline that depends on exact agreement still needs a spot check rather than an assumption.

The CLI reads a compressed dictionary that the library API will not open

Every dictionary distributed on the Releases page is compressed with zstd, and the tokenize binary takes the compressed file directly:

code
$ echo '本とカレーの街神保町へようこそ。' | cargo run --release -p tokenize -- -i ipadic-mecab-2_7_0/system.dic.zst

The vibrato library API has no such path. Compressed models must be decompressed outside the API before the dictionary type will accept them:

rust
// Requires zstd crate or ruzstd crate
let reader = zstd::Decoder::new(File::open("path/to/system.dic.zst")?)?;
let dict = Dictionary::read(reader)?;

That is a real difference between the two surfaces. A library consumer following the Releases page will land on a .dic.zst file, pass it to Dictionary::read, and get a decode error, while the same file works fine through the command line example. The decompression crate is also left to the caller, with either zstd or ruzstd accepted. This also means a web or service deployment has to stage the dictionary on disk before the first request instead of streaming it, which is precisely the situation the browser demo sits in: the Wasm demo at vibrato-demo.pages.dev is noted as taking a little time to load the model.

Token output format is a second thing the API does not decide for you. The command line binary writes the MeCab field layout, ending every run with a literal EOS line, and the field count varies with the entry: some tokens carry nine comma-separated fields, while user dictionary entries print only the features the user supplied. A second output mode drops the field layout altogether, since -O wakati separates tokens by spaces and is what most downstream text pipelines want:

code
$ echo '本とカレーの街神保町へようこそ。' | cargo run --release -p tokenize -- -i ipadic-mecab-2_7_0/system.dic.zst -O wakati
本 と カレー の 街 神保 町 へ ようこそ 。

There is no install command, only a clone and a build

The quickstart never says how to install vibrato. It opens with a reminder to install rustc and cargo following the official instructions at rust-lang.org, then goes straight to fetching a dictionary and running the tokenize binary from inside the checkout:

code
$ wget https://github.com/daac-tools/vibrato/releases/download/VERSION/ipadic-mecab-2_7_0.tar.xz
$ tar xf ipadic-mecab-2_7_0.tar.xz

The literal word VERSION in that URL is the catch. The walkthrough says to specify an appropriate Vibrato release tag to VERSION such as v0.5.0, so the person following along has to know the tag list and pick a dictionary built for the release their code expects. Nothing in the command derives the tag for you.

The release history makes that choice load bearing. Tags are v0.5.0 from 2023-02-22, v0.5.1 from 2023-05-12, and v0.5.2 from 2025-03-19, with a gap of about two years between the last two. The repository itself was last pushed on 2026-09-19, roughly six months after the newest tag, so the source tree a clone gives you is ahead of any precompiled dictionary you can download.

The stated entry point is therefore a source build rather than a package add. No step ever shows vibrato being added as a dependency of another Cargo project, which leaves a consumer to point at the git repository or at a local path and to work out for itself whether the library crate or the tokenize binary is the one it needs. A precompiled dictionary is called an easy start, and the Releases page is said to distribute several precompiled dictionaries from different resources, with mecab-ipadic v2.7.0 used as the worked example. System dictionaries can also be compiled or trained from your own resources, but that path is delegated to a docs directory rather than shown inline.

Nine workspace members, and the examples directory is explicitly excluded

The repository is a Cargo workspace whose member list names the whole toolchain and then excludes the examples:

toml
[workspace]
members = [
    "vibrato",
    "compile",
    "map",
    "tokenize",
    "benchmark",
    "train",
    "dictgen",
    "evaluate",
]

exclude = [
    "examples",
]

Only vibrato and tokenize appear in the quickstart, which means the build you run with cargo run is compiling one of seven crates you were not told about. The names describe a full lifecycle: dictgen generates dictionaries, train fits them, compile turns system dictionaries into vibrato format, map handles character conversion, benchmark measures speed, and evaluate scores output. Each of those has a home page in the docs directory, which the walkthrough describes as providing descriptions of more advanced usages such as training or benchmarking.

Training costs are part of the feature set rather than an afterthought, and the file that covers them is separate: the training parameters, called costs in dictionaries, are supported for corpora you supply yourself, with the detailed description placed in docs/train.md. Additional checks to avoid empty embeddings appear as a line item in the changelog-style list of other changes.

Two example projects sit outside the workspace, in examples/mecab_smalldic and examples/wasm, and the exclude entry keeps them from being built by default workspace commands. The wasm example is what the hosted demo is built from, and that demo is hosted at vibrato-demo.pages.dev. The Python binding is not in this repository either; it lives in a separate project at github.com/daac-tools/python-vibrato, and so does any discussion of using vibrato outside Rust.

User dictionary cost decides the match, and the docs reach for negative numbers

A user dictionary is layered on top of the system dictionary and pointed at with -u. The format is CSV with a fixed shape:

code
<surface>,<left-id>,<right-id>,<cost>,<features...>

The first four columns are always required and everything after them is optional, which means an entry can be a bare surface with identifiers and a number, and the feature columns you leave out simply are not printed later. The worked example leans on the cost column hard:

code
$ cat user.csv
神保町,1293,1293,334,カスタム名詞,ジンボチョウ
本とカレーの街,1293,1293,0,カスタム名詞,ホントカレーノマチ
ようこそ,3,3,-1000,感動詞,ヨーコソ,Welcome,欢迎欢迎,Benvenuto,Willkommen

With -u user.csv attached, the run splits the input in ways the system dictionary would not have chosen. 本とカレーの街, which carries a cost of 0, is emitted as a single custom noun token and displaces the ordinary segmentation of that span, and ようこそ with a cost of -1000 is pulled out ahead of the default interjection entry. The rule of thumb in these numbers is that a negative cost is a hammer for forcing a span, and that whatever wins locally changes the boundaries of its neighbours as well, which is why the outputs on both sides of the custom span differ from the plain run.

Because features are free-form, the same surface can carry several comma-separated values in one row, as ようこそ does above, and those values are echoed in the output instead of the standard field padding.

Two license files on disk against one license in the metadata

The licensing is a choice, stated as a choice. Contributors are told the crate is licensed under either of the Apache License, Version 2.0 or the MIT license at your option, and both files are present at the top of the repository as LICENSE-APACHE and LICENSE-MIT alongside the manifest. The license recorded for the repository is only the Apache-2.0 side.

That gap matters for anyone who needs a single license for a compliance checklist, since the metadata field and the files on disk disagree in kind rather than in version. Choosing MIT is the option that the file listing proves is available.

Release timing follows the same pattern of uneven surfaces. Three tags are visible in the history, and the intervals are wide: v0.5.0 on 2023-02-22, v0.5.1 on 2023-05-12, v0.5.1 to v0.5.2 on 2025-03-19, about twenty-two months later. No tag has appeared since, while the last push to the branch was on 2026-09-19, so anyone pinning to a tag is pinning to code that predates the current tree by half a year.

Support runs through a Slack workspace rather than an issue tracker narrative: daac-tools.slack.com with an invitation link posted in the file for both developers and users to ask questions and discuss a variety of topics. Contribution rules are pointed at by a single link to CONTRIBUTING.md, and the technical background is cited as three outside write ups, two papers from the 29th annual meeting of the Japanese NLP society in 2023 and one 2022 engineering blog post.

The speed claim is a figure and a wiki page, not a number

Performance is the project's headline, and it is argued rather than measured in the file. Vibrato is described as a Rust reimplementation of the MeCab tokenizer whose implementation has been simplified and optimized for even faster tokenization. The specific mechanism named is cache-efficient id mappings, and the specific case named is language resources with a large matrix, with unidic-cwj-3.1.1 given as an example at a matrix of 459 MiB. The claim is that on resources of that size vibrato runs faster as a result.

No timing numbers appear in the file itself. What is there is a figure showing an experimental result of tokenization time with MeCab and its reimplementations, plus a pointer to a Speed-Comparison wiki page said to hold the detailed experimental settings and other results. So the evidence is separated from the assertion, and the experimental conditions live outside the repository. The image itself is not in the repository either: the top level holds a figures/ directory, while the settings and the remaining results sit on the wiki page.

The surrounding design choices are consistent with a memory-mapped, large-dictionary story. Dynamic id mapping for arbitrary static dictionaries is listed as a feature, and the workspace carries a map crate for it alongside the one that loads dictionaries. The evaluate crate is the other half of that pair, since the benchmarking and scoring steps a tokenizer comparison needs are a first-class workspace member rather than a script someone keeps locally. The technical write ups cited in the file explain where the speed comes from, covering CPU cache efficiency in minimum-cost morphological analysis and the trade off between model size and parsing speed when splitting the score computation of a CRF-based analyzer. Neither paper's numbers are reproduced in the documentation.

One practical caveat for judging the comparison: the MeCab compatibility flags, the user dictionary, and the unknown word grouping setting all change the work the analyzer does, so a benchmark run under different flags is not measuring the same path.

Editorial conclusion

Choose vibrato when you are tokenizing large Japanese corpora in Rust and need MeCab-shaped output, and only after you have budgeted for building from source, since the quickstart is a clone plus cargo run and no install line exists. Verify the tokenizer flags against your own gold standard, because acknowledged cost tiebreakers can still shift boundaries, and remember the released tag trails the source, so a dictionary built for one is not guaranteed to load in the other. Wrappers in other languages, the browser build, and the advanced training and benchmarking paths are all documented outside this repository.

Frequently asked questions

Does vibrato produce the same tokens as MeCab with default settings?

No. MeCab ignores spaces because SPACE is defined in char.def, while vibrato emits each space as a 記号,空白 token. Passing -S and -M 24 makes the output identical on the worked example.

Can the vibrato library API load a compressed system dictionary?

Not directly. The distributed models are zstd compressed and must be decompressed outside the API, with a reader from the zstd crate or the ruzstd crate, before Dictionary::read will accept them.

What is the -M 24 argument in vibrato tokenization?

-M sets the maximum grouping length for unknown words, and the walkthrough pairs it with -S to obtain the same results as MeCab. It is presented as the value to use with the ipadic example dictionary.

How is a vibrato user dictionary file formatted?

It is CSV with the shape <surface>,<left-id>,<right-id>,<cost>,<features...>, where the first four columns are always required and the features are optional. It is passed to the tokenizer with the -u argument.

Which crates does the vibrato workspace build?

The workspace members are vibrato, compile, map, tokenize, benchmark, train, dictgen, and evaluate, with the examples directory explicitly excluded.

How is vibrato licensed?

It is offered under either the Apache License, Version 2.0 or the MIT license at your option, with LICENSE-APACHE and LICENSE-MIT both present in the repository, while the recorded license is only Apache-2.0.

Official sources

  1. daac-tools/vibrato on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/daac-tools-vibrato.svg)](https://hysenlabs.com/projects/daac-tools-vibrato)