Model or dataset
nanaism/yomiyasu avatar
nanaism/yomiyasu

yomiyasu: an agent skill that edits AI-generated Japanese prose, and a linter that admits what it cannot judge

Agent Skill for Refining AI-Generated Japanese into Natural Japanese.

1,800 stars40 forksPythonMIT

At a glance

What is it?
Seven transformation rules, a stdlib-only Python checker and a 160 document corpus, shipped in a ZIP whose advertised file count changed twice in three releases.
Who is it for?
The strongest thing in this repository is the refusal to guess. Rule six forbids adding subjects, causes, conditions or numbers that the source does not contain, and the release notes take the same position about their own verification, stating that a small sample of comparisons does not guarantee improvement on new material or that nothing regresses.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 9, 2026, and from our analysis. They are not legal advice.

Editorial analysis

A skill whose job is deleting the model's own habits

yomiyasu, written as a Japanese word meaning easy to read, is an MIT licensed agent skill with about 1,590 stars and 35 forks, no open issues, and its last push on 2026-10-06. The repository language is Python, which matters because the code here is not the transformation itself. The transformation is done by whatever model your agent is already running. The Python is a checker.

The target audience is stated precisely: technical articles, design and specification documents, pull request descriptions and internal reports. The README frames the problem as AI-generated Japanese that reads oddly in a specific way, and it names four failure modes of earlier attempts. Banning surface words fails because a vague synonym replaces the vague word. Giving too many rhetorical rules fails because the model over-applies them and invents grandiose coinages. Omitting the subject and object while attaching concepts to metaphorical verbs forces the reader to reconstruct the context. And piling up bold text and bullet lists hollows out the information density, so a long document carries less substance than a short one.

The installation surface is wide because the skill is a single file at the repository root. The README states that the root `SKILL.md` is the only entry point, and the Claude Code plugin manifest points at it directly. Two installers are documented:

bash
# 新規インストール
npx skills add nanaism/yomiyasu

# 最新版へのアップデート
npx skills update yomiyasu

The README also flags the update path it verified, saying it confirmed the normal repository route on skills CLI 1.7.0 and that you do not need to uninstall and reinstall.

Seven rules, and the one that constrains the rest

The transformation principles are numbered one to seven, and they read like a review checklist rather than a style guide.

Rule one asks the model to pair up subjects with predicates, modifiers with what they modify, and pronouns with what they point at, then verify that conditions, exceptions, negations, parallels, quantities and ordering are still traceable after the rewrite. Rule two separates what each sentence is doing, explanation, request, suggestion or plan, and fixes only the endings that do not match, while preserving the plain or polite form you started with and never converting everything to polite because the topic is technical. Rule three tidies personification, fixing expressions that give tools or concepts feelings and will, and keeping inanimate subjects that simply describe how a system behaves.

Rule four is the metaphor rule, and it names its targets directly: verbs for breaking, toppling, working and dissolving get replaced with plainer wording while the connotation and the range of meaning are preserved. The rule also requires checking that subject, object and modifier attachment did not drift after the swap, and that word order still lets a reader follow. Rule five allows cutting a preamble only when the claim and its weight do not change, keeping evaluation, necessary contrast and the connectors that move from comparison to decision.

Rule six is the one that constrains everything else. The model must not add subjects, causes, conditions, numbers or emotions that are not in the original. Where a rule calls something a contract or a canonical reference, the model should pick a word that fits the role instead. Where meaning cannot be determined, the output keeps the original's unresolved state as a provisional sentence in the body, and the reasoning plus any open questions go after the body rather than inside it.

Rule seven sets numeric targets: average sentence length of 30 to 45 characters and zero to two commas per sentence, while explicitly refusing to strip commas from prose that already reads well just to satisfy a count.

The After samples break the project's own sentence length rule

This is the most checkable contradiction in the repository, and it is worth doing yourself rather than taking anyone's word for it.

Rule seven sets an average sentence length of 30 to 45 characters. The two official After samples in the README run considerably past that. The first sentence of the business and specification example, the one explaining why a design system gets introduced, measures 69 characters. The sentence about component appearance and accessibility requirements measures 73. In the technical example, the opening sentence about asynchronous queue processing measures 55.

So both facts stand. The rule states a target, and the samples the maintainer considers exemplary output sit 20 to 60 percent above it. There is no stated exemption for explanatory prose.

There is a reasonable reading where both are right, and it is worth spelling out rather than picking a winner. A 30 to 45 character average works for Japanese when you count characters the way a Japanese reader does, and Japanese readers do count characters. But a rewritten sentence that merges three short claims into one explicit subject-verb-object structure will naturally run longer than the vague original it replaces. Rule one pushes toward naming the actor and the object, and every name you add costs characters. Rule six pushes against deleting anything that carries meaning, which also costs characters. The length target and the precision rules pull in opposite directions, and in the samples precision won.

What to do with that as a user: treat 30 to 45 as a guideline the model may exceed when the sentence needs an explicit subject, and do not accept a rewrite that shortens a sentence by dropping the actor. If the length number matters to you for a reason such as a mobile reading column, say so in your prompt rather than assuming the default will hold.

A linter that tells you what a PASS does not mean

The bundled tools are Python scripts that inspect wording and Markdown bold usage, sharing a module named `markdown_visibility.py`. The README states they need no external libraries and run on the standard library alone, and that a shared helper exists so the bold rules stay consistent across scripts.

What is more unusual is the disclaimer, and it comes before the tool list finishes rather than at the end. The README says the tool is a simple check aimed at specific syntax, not a complete Markdown renderer, and that whether bold actually displays correctly depends on the viewing environment. It says detection results are candidates for review, and that a `[PASS]` means only that no configured rule fired. It explicitly does not guarantee naturalness, meaning preservation, correct display in every environment, or the absence of false positives. It tells you not to mechanically fix items until the count reaches zero, but to compare each against the original sentence and context.

That last instruction is the operational point. A linter that reports a score invites optimization against the score, and a prose checker that reports severity invites argument about severity. Both invitations lead to worse prose, which is the exact failure mode rule two warns about when the model over-applies rhetorical rules.

The scripts named in the README start with a static lint checker for AI-flavored patterns, described as unnatural metaphorical verbs, excessive bold and bullet lists, emoji, and sentence-ending patterns. Both release v1.0.6 and v1.0.7 report the same verification shape: 18 unit tests including 50 bold regression cases, and identical linter output across 160 stored documents, with v1.0.7 changing only the guidance text attached to warnings while detection rules, lines, severity and scores stayed the same. The repository tree backs this up with `tests/`, `evals/`, `references/` and a `scripts/` directory, plus an `articles/` folder that holds the explanatory writing.

Nine files in the release notes, eleven in the README

If you install from source you never hit this. If you upload to the Claude web interface you hit it immediately, because that path requires a purpose-built archive rather than the repository ZIP.

The README is specific about why. Claude's custom skill registration rejects a full repository ZIP because it carries plugin configuration and development files, so the project publishes a separate `yomiyasu.zip` attached to the latest release. The README says that archive contains exactly 11 files counted from the repository root: the skill body, five reference documents, the inspection scripts, and the license. You download it and upload it without unzipping.

Now read release v1.0.7, published the day before the latest one. Its notes say the ZIP is limited to 9 files, naming the skill body, reference documents, inspection tools and LICENSE. Release v1.0.6, two days earlier, describes the same archive as containing the skill body, reference files, inspection scripts and LICENSE, without giving a count.

So the numbers do not line up across three documents describing the same artifact, and the README's breakdown is the only one that itemizes, which makes it the most likely to be current. It is also the number you would check a download against.

The fix is cheap and worth doing. After downloading, count the entries and compare them with what your agent reports as installed. A mismatch is not a broken skill, it means the archive and the documentation are describing different sets, and you want to know which side moved before you trust the file list for anything. Note also that v1.0.6 removed the Claude Code specific `argument-hint` field from the shared skill frontmatter because the web upload path does not allow it, so the two surfaces genuinely do differ in what they accept.

One more piece of context on packaging: the topics list includes `antigravity`, `gemini`, `cursor`, `codex` and `claude-code`, while the README documents install paths for Claude Code, Codex, Cursor and the Claude web interface. Antigravity and Gemini appear as topics without a documented path in the README.

Performance claims that depend on which number you read

Release v1.0.8 leads with a 99.1 percent reduction in diff inspection time, and the release notes immediately explain that the figure is the maximum across the conditions measured, not a general speedup.

Three rows are given. Diff self-comparison on 7 documents went from 8,473.40 ms to 73.08 ms. Diff on 48 real before and after pairs went from 184.30 ms to 78.44 ms. Linting 160 existing documents went from 20.41 ms to 15.42 ms.

Read them in order and the shape is clear. The headline number comes from feeding documents identical content to the diff tool, which is a self-comparison and the easiest possible case for an algorithm change. The realistic case, 48 real pairs, improved 57.4 percent. The lint path, which is what you would run most often on a single document, improved 24.4 percent and was already measured in tens of milliseconds.

The measurement conditions are stated with unusual care: public version v1.0.7 at a named commit against the new version, one machine, Python 3.14.5, one warm-up, seven alternating runs, median, total time. The note then says the 99.1 percent maximum excludes CLI startup and does not include LLM generation time, and that the 7 document self-comparison used the README and 6 guides as identical input, which is separate from the 48 real pairs.

This is a good model for how to read a performance claim from a small project. The author has told you exactly which input produced the headline figure and what it excludes. If you only ever lint single documents, the number you care about is 24.4 percent on a 20 millisecond baseline, which you would never have noticed. If you diff large documents in a batch loop, the self-comparison figure is closer to your workload than the 48 pair figure is, since batch diffing is also repetitive work over similar input.

Separately, v1.0.7 describes its prose changes as based on human A and B evaluation, cross-checking 32 Codex outputs from fixed inputs and 5 Claude outputs under identical conditions across 22 unique conditions after deduplication. It states plainly that this does not guarantee improvement on new topics or that nothing regresses anywhere. That sentence is worth more than the counts around it.

Editorial conclusion

The strongest thing in this repository is the refusal to guess. Rule six forbids adding subjects, causes, conditions or numbers that the source does not contain, and the release notes take the same position about their own verification, stating that a small sample of comparisons does not guarantee improvement on new material or that nothing regresses. That is rarer than it should be. The cost is that the skill leaves sentences unresolved, appends its reasoning after the prose instead of inside it, and can produce output you have to read twice, which is a real trade for anyone processing thousands of documents. On the tooling side, a `[PASS]` line means only that no configured rule fired, and the README says so in plain language before you get the chance to read it as a quality grade. Check the file count in the ZIP you download against the number your agent reports, since the release notes gave 9 files and the README says 11, and treat the sentence length target as a guideline rather than a gate, since the official After samples run well past it. Nothing here needs a network call or an API key, which makes the whole thing cheap to try on a document you already care about and easy to throw away if the register is wrong.

Frequently asked questions

Does yomiyasu need an API key or a network call to work?

The bundled inspection scripts need no external libraries and run on the Python standard library alone. The rewriting itself is done by the model already running in your agent, so there is no separate service to configure or pay for.

How do I install yomiyasu into an agent other than Claude Code?

The README documents a skills CLI route with `npx skills add nanaism/yomiyasu` and an openskills route with `npx openskills install nanaism/yomiyasu` followed by `npx openskills sync`, which makes the skill reachable through `AGENTS.md` for agents such as Codex and Cursor.

What does a PASS result from the bundled linter actually tell me?

Only that none of the configured rules fired. The README says explicitly that a pass does not guarantee naturalness, meaning preservation, correct bold rendering in every environment, or freedom from false positives, so treat each finding as a candidate to check against the original sentence.

Why does the README tell me to download a separate ZIP instead of the repository archive?

Claude's custom skill registration rejects a full repository ZIP because it includes plugin configuration and development files. The project publishes a registration-only archive on the latest release, and the README says to upload it without unzipping it first.

Can yomiyasu be used together with another Japanese proofreading skill?

The README warns that running it alongside another Japanese proofreading skill can produce conflicting instructions, and suggests temporarily disabling a similar skill if the output becomes inconsistent. There is no documented merging or priority mechanism.

Official sources

  1. Issues
  2. License: MIT
  3. nanaism/yomiyasu on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/nanaism-yomiyasu.svg)](https://hysenlabs.com/projects/nanaism-yomiyasu)