CLI tool
MrGeDiao/shuorenhua avatar
MrGeDiao/shuorenhua

shuorenhua: a Chinese-first rewrite skill that strips AI tone without losing the numbers

AI skill Chinese-first rewrite skill for Codex / Claude Code / Cursor / ChatGPT, removes AI tone, preserves facts.

1,877 stars79 forksPythonMIT

At a glance

What is it?
MrGeDiao/shuorenhua is a rewrite skill for Codex, Claude Code, Cursor and ChatGPT that targets Chinese AI writing patterns. Its selling point is a fidelity contract: the phrasing changes, the facts do not.
Who is it for?
Adopt shuorenhua if you draft Chinese with a model and publish the result to other people: README files, release notes, issue replies, forum posts. Skip it if your text is English-first, or if you want a general style linter across languages, since the rule layer is built around Chinese phrasing and English coverage is the smaller half.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 6 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem shuorenhua targets: Chinese drafts that sound like a model wrote them

The repository is explicit about the failure it is aiming at. Models drafting Chinese produce a recognizable register: opening pleasantries, summary prompts, business jargon that inflates ordinary actions into strategy, engineer posturing, translationese, nominalization, and authority claims with no source behind them. The README lists the affected surfaces as chat, progress updates, README files, release notes, forum posts, issue replies and long-form Chinese writing.

The intended user is a developer, maintainer or writer who drafts in Chinese with an AI assistant and then publishes the result. That is a narrower audience than a general proofreading tool. The skill is not trying to improve prose for its own sake; it is trying to remove a specific set of marks that make readers suspect a machine wrote the text.

The interesting part is what the project refuses to do. The README's own example takes a sentence about an interface optimization and shows two rewrites. The naive one drops the jargon and produces a shorter sentence, but also drops `p95`, `480ms` and `160ms`. The project treats that as a failure, not a success, and the example maps to a hard-constraint case in its benchmark. The claim the project makes for itself is that the rendering words should go and the evidence should stay.

How the rewrite pipeline works: scene, protected spans, tier, then a two-way read-back

The processing order is fixed and documented rather than left to the model's judgement. First the text is classified into a main scene: `chat`, `status`, `docs` or `public-writing`. Then the skill marks out what it calls protected spans: numbers, versions, commands, quotes, responsible parties and factual relations. Next it decides hit strength, `Tier 1 / 2 / 3`, and how aggressively to rewrite.

Only after that does it touch sentence and paragraph patterns. Phrase lists are described as a fallback layer, not the main mechanism, which is a deliberate position against find-and-replace humanizers. After rewriting, the text goes through a fidelity read-back in two directions: every fact in the input has to be findable in the output, and every new relation in the output has to point back to something in the input. If residue remains, a lighter Residual Audit runs once more.

The fidelity contract is the most concrete thing in the repository. Numbers keep the object they modify, so a latency figure cannot be summarized as a vague improvement. Relations are not rewritten: a sentence saying something shows the potential of an architecture cannot become a sentence saying the architecture was adopted. Scope, conditions, negation, modality, completion, direction and intensity all count as facts. Abstract claims are not made concrete: if the source only says efficiency improved, the rewrite cannot add that time or cost was saved. When information is missing, the skill is supposed to point at the gap rather than fill it.

Scene packs adjust the emphasis. A README gets judged on whether the first screen says what the thing is, who it is for and what problem it solves. A release note lists changes, verification and limits instead of a launch declaration. An issue reply leads with reproducibility, current assessment and next step. An API reference protects endpoints, methods, fields, status codes, constraints and recovery actions. For long-form writing, a separate `scope` setting governs how much can be cut: `structural` allows deleting and reordering sentences, `bounded` is the default for long text and moves whole empty sentences into a suggested-deletion list, and `in-place` forbids deleting sentences at all and only lowers the tone inside them.

Installing shuorenhua in Claude Code, Codex and other agents

The README offers a no-install path first: a ChatGPT custom GPT for people with Plus or Pro, where you paste text and get a rewrite. For anything repeatable, there are three entry sizes. The mini entry is a single self-contained file under 1,500 characters, suited to one-off sessions, custom instructions or a tight context window. The lite entry is `SKILL.md` alone, for ad hoc rewrites and light review. The full entry is `SKILL.md` plus the `references/` directory, and the README recommends it for long-running projects, public text, technical documentation and false-positive protection. The Claude Code plugin ships the full set.

In Claude Code the install is two plugin commands run inside the conversation.

text
/plugin marketplace add MrGeDiao/shuorenhua
/plugin install shuorenhua@shuorenhua

After that, the README says you can simply ask it to remove the AI tone from a passage. For Codex, the documented route is a clone plus a single execution that points the agent at the skill file.

bash
git clone https://github.com/MrGeDiao/shuorenhua.git && cd shuorenhua
codex exec -C . "读取 ./SKILL.md,按其中规则改写以下文本:……"

Agents that support a `skills` command have a shorter path.

bash
npx skills add MrGeDiao/shuorenhua

If you only want the problems flagged rather than a new draft, the README says to add a line asking for annotation mode, which marks issues without rewriting. For a project that is used repeatedly, the README suggests a trigger block in `AGENTS.md` that routes any request about removing AI tone or sounding less like a template to `shuorenhua/SKILL.md`, and explicitly excludes code, logs, configuration and command output from the skill. That exclusion is worth copying: the skill is written for prose, and applying it to a stack trace is a category error.

Where shuorenhua gets it wrong, and when it is the wrong tool

The project's own release notes are the best source on its limits. The v2.3.1 entry states that the release closed on a single-seat basis with Opus, and that DeepSeek V4 Pro was withdrawn from the formal seat because a real L1 case plus a same-condition rerun showed run-to-run variance. A second model, Grok 4.6, was run as a replacement, and the notes record its rewrites and hard judgements as clean but its scoring as covering only 1 of 7 batches, filed as supporting evidence rather than a second formal seat. In plain terms: the headline pass numbers come from one model, and the project says so.

The same notes record a known gap in the human corpus. The direct human samples still lack `docs` and `status` categories, and the repository's `check_repo` reports this as a known gap that does not block CI, to be closed once 12 pieces are collected. So the false-positive evidence is thinner for documentation and status text than for other scenes.

The benchmark structure also sets a boundary on what the published numbers mean. The 111-case main benchmark and the 20 scene samples are separate suites and are not additive. The 8 human long-form residual comparisons, drawn from 7 author groups, are used only to observe false positives; they do not enter the benchmark, rewrite or judge denominators, and no human-likeness threshold is set from them. Anyone reading the numbers as a general quality score is reading them wrong.

There is also a language boundary. The rule layer is described as covering 210+ Chinese phrases, 96 English phrases and 25 structural anti-patterns. English is supported but is the smaller half, and the scene packs, examples and fidelity discussion are all framed around Chinese. If your output is English-first, most of the repository's tuning does not apply to you. If you want a linter that checks style across a mixed-language codebase, this is not that either.

How shuorenhua differs from generic humanizer tools and from a style linter

The obvious comparison is a word-swap humanizer: a list of AI-flavored phrases on one side and plain replacements on the other. The difference is structural. shuorenhua's phrase tables are explicitly the fallback layer, applied after scene classification, protected-span marking and tier judgement. The README's own worked example is the clearest statement of the split: a phrase-swap pass removes the jargon and also removes the latency numbers, while the rule-based pass keeps them. A humanizer optimizes for the sentence reading naturally; this project optimizes for the sentence reading naturally while the evidence survives.

A second comparison is a prose linter or style checker. Those report violations and stop. shuorenhua produces a rewrite by default, and its annotation mode is an option you have to ask for. The trade-off is that a rewrite is harder to review than a list of flagged lines, which is presumably why the project invests in the two-way read-back and the hard-constraint gate.

The third comparison is doing nothing and editing by hand. That is a real option, and for a single paragraph it is probably faster. The case for the skill is volume and consistency: many release notes, many issue replies, the same set of marks each time. The project's answer to the consistency problem is the layered release gate, where L1 hard constraints (fabricated facts, protected-span drift, changed attribution, scope violations) must fail zero times and the false-positive rate on text that should not be changed must stay under 10 percent, while style goals are only reported as trends per model and style observations are recorded without blocking.

Evaluating shuorenhua before you trust it on your own writing

The evaluation apparatus is unusually inspectable for a prompt-based tool, and it is also where the maintenance cost sits. The main benchmark is 111 cases: 61 SF cases where the text should be changed and the main problems caught, and 50 SNF cases where normal text should be passed through or only lightly flagged. The 20 scene samples are scored separately on naturalness, fidelity and whether the result is publishable as-is, with length and rhythm judged for long text. Inside the main benchmark there are 16 scene-pack positive and negative cases, 4 long-form in-place cases and 3 bounded cases.

Models under test see only `evals/benchmark-blind.md`, which is anonymized, shuffled and carries no expected answers; the judge scores against a separate mapping table. The models run, the benchmark version and the results are logged in `evals/run-manifest.md`. A zero-dependency script, `python3 automation/eval/hard_metrics.py --run <batch-dir>/`, performs a rough check of word retention, dash density and protected spans. The latest release-ready evidence is the v2.3.1 results file under the single-seat Opus basis.

The cost side is real. The full entry is `SKILL.md` plus `references/`, and the reference set includes separate documents for protected spans, positive style, scene packs and structures. The release notes for v2.3.1 mention a fix to a counting-unit contradiction in the dash-density criterion in `references/structures.md` section 20, which tells you the reference documents are edited and versioned, not static. Following the project means tracking those files, not just the skill file.

On licensing: the repository is MIT, and the README states that the human corpus text and its adaptations retain their own licences and are not covered by the repository-root MIT. The scene packs, benchmark and reference files sit under the MIT terms. That distinction matters if you plan to redistribute the evaluation corpus rather than the skill.

Editorial conclusion

Adopt shuorenhua if you draft Chinese with a model and publish the result to other people: README files, release notes, issue replies, forum posts. Skip it if your text is English-first, or if you want a general style linter across languages, since the rule layer is built around Chinese phrasing and English coverage is the smaller half. Before trusting it on your own material, run the mini entry on one real paragraph that contains a number and a version string, and check that both survive the rewrite.

Frequently asked questions

Does shuorenhua work with Claude Code?

Yes. The README documents a plugin install using two commands, `/plugin marketplace add MrGeDiao/shuorenhua` and `/plugin install shuorenhua@shuorenhua`, and the Claude Code plugin ships the full entry with `SKILL.md` plus `references/`.

Will shuorenhua delete the numbers and version strings in my text?

The project's stated design is the opposite: numbers keep the object they modify, and protected spans covering numbers, versions, commands, quotes and attribution are marked before any rewriting happens. The README's worked example shows a phrase-swap rewrite losing `p95`, `480ms` and `160ms`, and presents the rule-based rewrite as the correct outcome.

Can I use shuorenhua without installing anything?

The README lists a ChatGPT custom GPT as the no-install route, requiring Plus or Pro, where you paste a passage and get a rewrite. The mini entry, a single self-contained file under 1,500 characters, is the smallest installable alternative for one-off sessions or custom instructions.

Does shuorenhua work on English text?

It is Chinese-first. The rule layer covers 210+ Chinese phrases against 96 English phrases, and the scene packs, examples and fidelity discussion are framed around Chinese writing. English is supported but is the smaller half of the rule set.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
  4. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/mrgediao-shuorenhua.svg)](https://hysenlabs.com/projects/mrgediao-shuorenhua)