Cangjie Skill: one release tag, two different packages, and the compile modes behind the output
GitHub describes it as 把书、长视频、播客等高价值内容蒸馏成可执行的 Agent Skills. The repository metadata lists Python as its primary language. The metadata lists the AGPL-3.0 license. This article stays within the project description and details documented in the GitHub repository README.
At a glance
- What is it?
- Cangjie Skill turns books, transcripts and podcasts into callable Agent Skill packs, but its packaging is unusual: the v2.5.0 tag was refreshed weeks after it was cut, output comes in either a single router Skill or a pack, and the whole toolchain lives in one script. Here is what each of those choices means before you wire it into an agent.
- Who is it for?
- Cangjie Skill fits teams that already have source material in text form and want a pack they can call from an agent, and it fits DeepSeek Harness users who install from the release tarball rather than from git. It does not fit anyone who expects git clone or a GitHub source archive to give them the current build, and it will not fetch a video for you.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 4 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 5, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The v2.5.0 tag and the current v2.5.0 package are not the same bytes
The v2.5.0 release was published on 2026-08-30. On 2026-09-13 the same version was refreshed, still labelled v2.5.0, and the changes were real: task-first validation now retains complete procedures and formulas explained in a single source location, output scoring counts missing runs and checks numeric values and units, and compiled Skills can carry declared scripts and text templates. A refreshed generic Skill ZIP and a SHA256 file were uploaded under the existing tag. The tag itself was not moved, so the automatic source archives GitHub builds from a tag do not contain the refresh. Existing users are told to download the refreshed package and to check `BUILD_INFO.json` for the source commit and the refresh date. The practical consequence is that a team pinning the git tag, or anyone who cloned the archive, has a build that predates the scoring and validation work, and nothing in that build tells them so. Tag history is not a reliable guide either: v1.0.0 was published on 2026-08-09 as a historical backfill, the same day as v2.0.0.
single or pack decides how many entry points your agent gets
Compilation ends in one of two deterministic delivery modes. `single` compiles one router-style Skill. `pack` compiles a compact pack consisting of a router plus promoted standalone Skills. Under both modes, extraction happens first: the capability bundle is the single source of truth, and extraction produces stable capability cards and metadata before any installable output is compiled. That ordering is the point, because the output shape is decided from the capability bundle rather than from the raw source text. In `pack` mode the promoted Skills become independently callable and composable, so the agent gets more direct entry points and you get more files to keep in step with each other. Registry v2 makes the output mode and the capability counts visible, and the project states that Registry v1 entries keep working. A consumer still reading the old registry therefore sees no difference between the two modes, which matters if you compare your own entries across an upgrade.
The DeepSeek Harness route installs a tarball, not a checkout
The one installation route the project spells out command by command is DeepSeek Harness, and it installs from the release package rather than from the repository. Download the package and its checksum, verify it locally, add the tarball to the web profile, then start it:
mkdir -p ~/.dsh/packages
curl -fL "https://github.com/kangarooking/cangjie-skill/releases/download/v2.5.0/dsh-cangjie-skill-2.5.0.tgz" \
-o ~/.dsh/packages/dsh-cangjie-skill-2.5.0.tgz
curl -fL "https://github.com/kangarooking/cangjie-skill/releases/download/v2.5.0/dsh-cangjie-skill-2.5.0.tgz.sha256" \
-o ~/.dsh/packages/dsh-cangjie-skill-2.5.0.tgz.sha256
(cd ~/.dsh/packages && shasum -a 256 -c dsh-cangjie-skill-2.5.0.tgz.sha256)
dsh plugin --profile web add ~/.dsh/packages/dsh-cangjie-skill-2.5.0.tgz
dsh webTwo things about this route deserve attention. The adapter layer is bundled inside the release package, so no platform-specific wrapper files are added to this repository, which also means you cannot read the adapter before you install it. And every step pins v2.5.0, so the command set as written reproduces the August build unless you change the version. The generic Skill route is separate: it is a ZIP with a checksum that you extract and install as the complete `cangjie-skill/` directory.
scripts/cangjie.py carries diagnostics, compilation and rollback in one file
The local toolchain is a single script, `scripts/cangjie.py`, which covers diagnostics, compilation, output replanning, incremental updates, repair, rollback, evaluation and benchmarking. The repository layout follows that shape: `extractors/`, `methodology/`, `templates/`, `schemas/`, `registry/`, `benchmarks/`, `books/`, `tests/`, `dist/`, `docs/` and `website/`, with `SKILL.md`, `CHANGELOG.md`, `CONTRIBUTING.md`, `GITHUB_REPO.md`, a LICENSE and three READMEs for Simplified Chinese, English and Japanese at the top level. Two consequences follow. The README gives no pip install command and names no package index entry for the script, so it is something you run from a checkout rather than something you add to a requirements file. And because compile, repair, replan and rollback are all operations of the same script against the same templates, a script older than the templates it consumes is the first thing to rule out when output looks wrong; `BUILD_INFO.json` is where you check which commit you actually have.
Content-addressed preprocessing is what makes an incremental repack safe
The evolution machinery in v2.5.0 is a list of specific mechanisms: content-addressed preprocessing, source diffs, impact analysis, transactional patches, edit detection, snapshots and rollback. Together they address one scenario, a pack that already exists and whose source has since changed. Content addressing is what lets the toolchain tell that a source is untouched, source diffs and impact analysis decide which compiled Skills are affected, transactional patches apply the change, and snapshots are what a rollback restores. What this list does not tell you is how many snapshots are kept, where they are stored, or what happens when the snapshot itself came from a bad extraction. Treat the rollback history as finite and check the build info before relying on it. Edit detection leaves a second gap: the project does not state which file changes count as semantic, so if a source is rewritten in place you cannot tell from the release tag alone whether a repack was ever needed.
Scoring counts missing runs and checks the numbers
Output scoring in the refreshed v2.5.0 package is described in three rules: it counts missing runs, it checks numeric values and units, and validation is task-first and retains complete procedures and formulas explained in a single source location. Those rules shape what survives compilation. A methodology a reader could follow from scattered notes fails when its procedure and formulas are not gathered in one place. A methodology with silent numeric defaults scores badly, because a number without a unit is treated as something to check rather than something to keep. The refresh also lets compiled Skills carry declared scripts and text templates, which moves part of a methodology out of prose and into files the Skill declares. What the README does not publish is the score that counts as a pass, the list of units compared against, or the tolerance on a numeric check. You cannot predict from the outside whether a given extraction will ship, so run the diagnostics and evaluation paths on one source before processing a whole shelf.
Media never enters the pipeline directly, and that caps the output quality
The stated inputs are books, videos that have subtitles or transcripts, podcasts, interviews, talks, courses, long-form articles and document collections. Video is the exception in practice. The project recommends running the `video-downloader` skill from the kangarooking-skills repository first: download the video, extract subtitles or audio transcripts, collect key materials, and hand the resulting text to Cangjie Skill for methodology extraction, skill construction and pressure testing. The trigger is a plain sentence:
Use cangjie-skill to distill this book into a set of executable Agent Skills: <file path>So Cangjie Skill does not fetch media, and the quality ceiling of a video derived pack is the transcript it was handed. A mis-transcribed number becomes a methodology that scores badly or, worse, one that passes. The README lists no supported file formats, no maximum input size and no limit on source length, so those are questions to answer with your own material rather than by reading the docs. Its scope sits next to `nuwa-skill`, which builds human skills such as an Elon Musk or Warren Buffett skill, and `darwin-skill`, which handles automatic skill evolution. Cangjie Skill works on what a person expressed systematically in a source, not on the person.
The website presents the packs, the repository stays the source
The official site at cangjie-skill.com offers visual Skill Pack browsing, a beginner-friendly usage guide, Skill detail pages and a contribution submission entry. The division of labour is stated plainly: this GitHub repository remains the sole source for the code, the methodology and the templates, while the website provides presentation, navigation and usage guidance. Registry v2 support is what lets the site show output mode and capability counts per pack. The consequence for contributors is that two submission paths exist, one through the website entry and one through the repository's own `CONTRIBUTING.md`, and the README does not say which governs a new pack or how a website submission reaches the repository. Version information lives in `CHANGELOG.md` and in `docs/releases/v2.5.0.md`. The last push to the repository was on 2026-10-02 while the newest release tag is v2.5.0 from 2026-08-30, so the default branch moves faster than the tags do, which is the same reason the refreshed package has to be fetched rather than rebuilt from the tag.
Editorial conclusion
Cangjie Skill fits teams that already have source material in text form and want a pack they can call from an agent, and it fits DeepSeek Harness users who install from the release tarball rather than from git. It does not fit anyone who expects git clone or a GitHub source archive to give them the current build, and it will not fetch a video for you. Before committing, download the refreshed v2.5.0 package and its checksum, open BUILD_INFO.json to confirm the source commit and refresh date, and run scripts/cangjie.py diagnostics on the source you plan to distil so you learn what the scoring rejects before you trust it with a whole library.
Frequently asked questions
What kinds of source material can Cangjie Skill distil?
Books, videos that have subtitles or transcripts, podcasts, interviews, talks, courses, long-form articles and document collections. Video material has to reach the tool as text, so the project recommends the video-downloader skill from kangarooking-skills first.
How do I install Cangjie Skill into DeepSeek Harness?
Download the v2.5.0 tarball and its SHA256 file into `~/.dsh/packages`, verify it with `shasum -a 256 -c`, then add it with `dsh plugin --profile web add` and start the web profile with `dsh web`. The adapter layer ships inside the release package, so no platform-specific wrapper files are added to the repository.
What is the difference between the single and pack output modes?
The single mode compiles one router-style Skill. The pack mode compiles a compact pack with a router plus promoted standalone Skills, which become independently callable and composable. Extraction produces the capability bundle before either output is compiled, and Registry v2 shows the mode and capability counts.
Do I need to download Cangjie Skill again after the September refresh?
Yes. The refreshed package was uploaded under the unchanged v2.5.0 tag, so GitHub's automatic source archives for that tag do not contain it. Existing users are told to fetch the new ZIP with its SHA256 file and to check `BUILD_INFO.json` for the source commit and refresh date.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/kangarooking-cangjie-skill)