Model or dataset
Raymondhou0917/speak-human-tw avatar
Raymondhou0917/speak-human-tw

speak-human-tw will not touch your draft until you approve its list, and its own counts disagree

「說人話」:繁體中文的去 AI 味改寫 skill。抓 38 種 AI 寫作痕跡,順手校正中國用語與半形標點,給 Claude Code / Codex / Cursor 用。

1,022 stars113 forksPythonMIT

At a glance

What is it?
A Traditional Chinese proofreading skill for AI agents, built around a refusal: two rounds by default, a protection list that prices and promo codes cannot be rewritten, a self score that blocks delivery below 35 of 50, and an eval suite that openly does not test whether the result sounds human.
Who is it for?
speak-human-tw is worth trying if you write in Traditional Chinese for Taiwanese readers and already have a draft you suspect reads like a press release, because the two round default and the protection list make it safe to run over work you care about. It is not a writing skill and it will not become one.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The default is two rounds, and the first one does not touch the file

The behaviour that surprises most people is the default round count. The first pass produces a numbered list only: where the line is, what the original sentence was, which writing trace it matches, and what a change would look like. It then asks one question, in the form of asking which of the listed items need changing, and stops. The second pass happens only after you answer, and only then does it edit and write to your files.

The reasoning is stated rather than assumed. The danger of a revision tool is not failing to catch something, it is covering your original before you have seen what the tool intends to do. Anything you do not tick stays untouched.

There is a way out, and it is a phrase: say that it should skip the list and edit directly. That matters for the second installation path, where the command line prompt itself contains the instruction to apply the changes this time without asking first.

Non interactive runs detect that nobody can answer a question, apply the changes anyway, and hand back a summary afterwards listing the original sentence, the reason for the change and what it became, so the result can be traced through a `git diff`. There is also an annotation mode, triggered by asking it to mark problems without changing anything, which lists between one and five points.

So the confirmation gate is a deliberate friction point, not a limitation. It is also the difference between a tool you can point at a finished newsletter and a tool you can only point at a draft.

Prices, promo codes, links and real names are on a list nothing may rewrite

The rewrite procedure is a fixed six steps, and step two is a lock. Before any editing, a protection list is built from prices, discount codes, proper nouns, links, real names, quoted wording and refund promises, and nothing in that list is touched for the rest of the run. Step five walks the same list again item by item to confirm each one survived, and states the rule that no fact absent from the original may be added.

Step three sets the other boundary. Long text, defined as roughly a thousand characters or more, is not shortened on its own initiative, and line by line correspondence is preferred. A sentence that is pure padding is not deleted silently either: it gets an entry in the list explaining why removing it loses no information, and the decision stays with you. That is why the first round can contain a proposal to delete something rather than only a proposal to reword it.

Step four does the actual work, processing the trace categories one at a time with pattern matching taking priority and a word list as the backstop, while the Taiwan usage check runs alongside it in the same pass rather than afterwards.

The practical reading: the tool is told to be conservative with information and aggressive with phrasing, and the two round default exists so you can audit the aggressive half before it lands.

A 50 point self score, and it will not deliver below 35

Step six is a gate rather than a report. Long text is scored on five dimensions out of 50, and below 35 the text is not handed over. This is the one place where the skill refuses to produce output, and it is a fixed number rather than a soft target.

The five dimensions are not named individually in the visible part of the documentation, so what a score of 40 is made of cannot be read off directly. What can be read is the shape: a self assessment that the tool runs on itself, with a threshold, before delivery.

Step one is the other place where configuration matters. There are five contexts, and each gets a different intensity. Social posts are edited lightly with the colloquial register kept. Sales pages are edited hard, but the call to action must not lose its force. Newsletters, customer service replies and office documents sit between those two. So the same sentence can be flagged in one context and left alone in another, and the context is not something you pass as a flag; it is something the skill has to work out from the text.

Combined with the protection list and the two round gate, the design is consistent: the tool refuses to guess on facts, refuses to guess on register, and refuses to ship when its own score is low.

The trace table adds up to 38, and the prose says 35 or more

The detection side is organised as a table of five categories, and the counts are worth adding up. Content traces number 9, language and syntax traces 12, style and layout 8, communication residue 7, and tool traces 2. That totals 38.

The prose in the same document says the skill catches more than 35 such traces, and the repository description says 38. So the table and the description agree with each other and the sentence in the introduction does not. Nothing breaks, since more than 35 is technically consistent with 38, but anyone comparing versions will find the headline number is the softest figure in the file.

The examples give the character of each category. Content traces cover inflated significance, vague attribution, invented citations, formulaic forward looking paragraphs and the stance vacuum phrases that concede everything. Language and syntax, the largest group, covers stacked not A but B constructions, three part parallelism, explanatory lead ins, fake inference, scolding rhetorical questions at the end, and aphorism formulas. Style and layout covers excessive dashes, bold overload, stacked emoji, paragraph fragmentation with first and next and finally, and misused tables. Communication residue is the leftovers: hope this helps, canned closings, knowledge cutoff disclaimers, template placeholder text. Tool traces, only two, are the most specific thing in the table: a `utm_source=chatgpt.com` parameter, a `turn0search` placeholder code, and Markdown markers left in the wrong place.

The eval set is 42 cases, and none of them test whether it sounds human

The benchmark file holds 42 cases, split into 27 that must be caught and 15 that must not. The second group is the interesting one, because it exists to prevent the tool from over reaching. It protects parallelism that is backed by facts, sourced data, standard payment terms, rhythm sentences in long text, words that are debated as AI flavour, and conditional choices of the form use A or use B.

Scoring includes a rule against changing the words without changing the sentence. Deleting one inflated marker and replacing it with a different inflated marker counts as a failure, which is the only defence against a tool that satisfies a pattern by swapping synonyms.

The limitation is stated with unusual clarity. What those 42 cases verify is that the things which should change did change, the things which should not change were not touched, no synonym swap happened, and no fact drifted. None of them tests whether the result reads more human, because that cannot be scored automatically, and scoring it would just create another formula. The project declines to claim that part is measurable and points instead at the separate humanize reference and at your own judgement.

For an evaluation-driven audience that sentence is the most useful line in the file: it separates the guarantee from the marketing.

Human flavour is declared to belong to the author, not the tool

One hard line runs through the whole design: the human part is the author's, not the tool's. Stories, positions and turning points the author never said are not to be invented on their behalf. What the skill may leave behind is a marker saying the author needs to supply it. The reasoning given is that a fabricated I was wrong once is worse than the empty sentence it replaced, because the empty sentence is only boring and the invented one is a lie.

A separate reference file defines where the text may then be pulled. Ending without a conclusion is allowed. A position is allowed to change over time. Digression is allowed. The stated reason is that these are things an AI cannot produce, since it has no past and will not admit uncertainty.

The FAQ is equally blunt on the adjacent question. No, the output does not acquire somebody's voice: the skill cleans the draft and does not install a persona, and removing AI flavour and having a personal style are treated as two different jobs. No, this is not for evading AI detectors; the stated goal is text that is more specific, more honest and more like a particular person talking.

It also states what the tool is not for. The comparison it offers is a handwritten newsletter from 2024, written before any of this existed, which produced no findings, against an AI drafted reply which produced six. The document is careful about what that proves: it shows the skill can tell the difference, and it explicitly does not mean you can use it as a general purpose writing skill.

Install is a clone into a skills directory, and the small version is one file

For Claude Code the whole install is one command into the skills directory:

bash
git clone https://github.com/Raymondhou0917/speak-human-tw.git ~/.claude/skills/speak-human-tw

Afterwards, asking it in conversation to remove the AI flavour from a passage triggers it. For a single Codex run, the clone is followed by an inline prompt that tells the tool to read SKILL.md and apply the rules, with the instruction to apply directly rather than ask first:

bash
git clone https://github.com/Raymondhou0917/speak-human-tw.git && cd speak-human-tw
codex exec -C . "讀取 ./SKILL.md,按規則改寫以下文字,這次直接套用不用先問我:(貼上你的文字)"

Separate installation guides exist for Claude Code, Codex, Cursor and OpenCode, including a note about symlinks and following updates. The core is a single SKILL.md file; the fuller version adds the references directory with the trace list, the Taiwan usage table and the false positive protections, which is what the category table above is counting.

One thing to settle before you install: the language scope. The skill is calibrated from scratch for Traditional Chinese as used in Taiwan, with a built in check for mainland vocabulary, 60 or more term pairs, full width punctuation and quotation rules, and the file says explicitly that this is not a simplified to traditional conversion. For Simplified Chinese scenarios the rules are described as mostly applicable, but the localisation layer is built for Taiwan readers.

The knowledge base is a Wikipedia cleanup page and one sentence pattern talk

Where the rules come from is stated in the same file that documents the limits. The primary sources are a Chinese Wikipedia page on the characteristics of AI generated text, maintained as a first hand observation by the WikiProject AI Cleanup community, and a sentence pattern analysis of what is called AI speech register by 朱宥勳. The second source is a video, not a paper.

The stated practical validation is different in kind: the core rules come from more than three years of real revision records, meaning the habits human editors repeatedly flagged, and those went into both the rules and the eval set. That is why the benchmark can be described as 42 cases rather than a general claim.

The scope is also narrow on purpose. Five scenarios are claimed, newsletters, social posts, sales pages, customer service replies and office documents, and the file positions itself as filling a different gap from existing Chinese tools rather than competing with them.

The placement guidance is the part to keep. AI writing splits into two columns: structure, first drafts, data tidying and outlines on one side, and deciding a position, choosing tone, producing metaphors and writing punchlines on the other. The stated failure is that a draft from AI hands you both columns at once, so the output looks clean and still reads like boilerplate, because it contains none of your views. This skill scrapes the right hand column out of the draft.

The repository is MIT licensed, has a single release, v1.4.0 on 10 July 2026, and received a push on 2 October 2026. The tree holds CHANGELOG.md, CONTRIBUTING.md, assets, evals, install, references and a Python scripts directory.

Editorial conclusion

speak-human-tw is worth trying if you write in Traditional Chinese for Taiwanese readers and already have a draft you suspect reads like a press release, because the two round default and the protection list make it safe to run over work you care about. It is not a writing skill and it will not become one. Before you use it, read the difference between its arithmetic and its prose, since the trace categories sum to 38 while the document body says 35 or more. Decide whether you want the annotation pass, because that is the feature that stops it overwriting you and also the feature that costs you a round trip. And if your problem is that your writing has no opinion in it, this project says plainly that it cannot supply one.

Frequently asked questions

Does speak-human-tw rewrite my draft in one pass?

Not by default. The first pass produces a numbered list with the line, the original sentence, the trace it matches and a suggested change, then asks which items you want changed and stops. The second pass edits only what you approved. You can ask it to skip the list and edit directly, and non interactive runs detect that nobody can answer and apply changes with an after the fact summary.

How many AI writing traces does speak-human-tw detect?

The category table lists 9 content traces, 12 language and syntax traces, 8 style and layout traces, 7 communication residue traces and 2 tool traces, which sums to 38 and matches the repository description. The introductory sentence in the same document says more than 35.

Is speak-human-tw meant to fool AI detectors?

The project says no. The stated goal is text that genuinely reads better, more specific, more honest and more like a particular person talking, and it does not inject a persona or a personal style. It also declines to claim that the result is more human in any measurable way, since that cannot be scored automatically.

Does speak-human-tw work with Simplified Chinese?

The rules are described as mostly applicable, but the localisation layer is built for Traditional Chinese readers in Taiwan. That layer covers 60 or more mainland vocabulary pairs, full width punctuation conventions and quotation rules, and the project states it is not a simplified to traditional conversion.

How do I install speak-human-tw for a single Codex run?

Clone the repository, change into it, and run codex exec with a prompt that tells it to read ./SKILL.md, rewrite the pasted text, and apply the changes directly without asking first. Separate installation guides exist for Claude Code, Codex, Cursor and OpenCode, and the core of the skill is a single SKILL.md file.

Official sources

  1. Issues
  2. License: MIT
  3. Raymondhou0917/speak-human-tw on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/raymondhou0917-speak-human-tw.svg)](https://hysenlabs.com/projects/raymondhou0917-speak-human-tw)