Open-source project
fanbuz/codesucker avatar
fanbuz/codesucker

CodeSucker: an offline Electron tool that turns a local repo into a 60-page software copyright source document

拖入项目,生成 60 页软著源程序文档。离线抽取、注释清洗、自动排版,导出前先帮你过一遍审查——代码不出本机。

416 stars84 forksTypeScriptApache-2.0

At a glance

What is it?
CodeSucker is a macOS and Windows desktop app that scans a local project, strips comments, paginates to the Chinese software copyright filing rules and exports a docx. Nothing about the code leaves the machine, but the rules it encodes are specific to one filing regime.
Who is it for?
Adopt CodeSucker if you are preparing a Chinese software copyright registration and want the 60-page, 50-line-per-page, header-and-signature rules applied by a local pipeline instead of by hand. Do not adopt it if you need a signed macOS build, a headless CLI, or a tool that decides licence and authorship questions for you; the repository states that core is plain TypeScript with no Electron dependency and could be reused as a CLI later, but no CLI exists today.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 10 days ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The filing format CodeSucker automates

Chinese software copyright registration asks for source program identification material with a shape that is easy to get wrong: 30 consecutive pages from the front of the program and 30 from the back, at least 50 lines on every page, a header carrying the full software name plus version number, and continuous page numbers. The README describes the manual version of this as a job that takes hours, and says tools in this space either concatenate files without applying the rules or run as an online service, which means uploading source. CodeSucker is aimed at the person who already has a repository on disk and needs a preparation draft of that document, not a general code documentation generator. The whole application is a five-step wizard: import a folder, choose and order files, clean and lay out, preview pagination, then audit and export. It is a Windows and macOS desktop app built on Electron 43 with React 18 and zustand, and the README states that scanning, cleaning, layout and export make zero network requests. The only outbound call it describes is a startup check against public GitHub Release metadata for a newer version, and that check failing does not affect core processing.

How the pipeline is split between core and shell

The repository is an npm workspace with two packages. packages/core is described as a pure TypeScript pipeline with zero Electron dependencies, organised as discover, third-party clues, clean, select, render, audit. packages/app holds the Electron shell, which the README says does IO and windows only. That split matters if you ever want to script the pipeline, because the stage names are the pipeline: discovery walks the tree and classifies files, the third-party clue stage parses dependency manifests locally, clean runs the comment stripper, select applies the 1500 plus 1500 line cut, render produces pages, and audit checks them. The README lists four technical decisions behind this. Pagination is explicit page breaks rather than layout-driven page filling, with fixed line spacing only as a fallback so that changing fonts does not shift pages. Comment stripping uses a character-by-character state machine rather than regular expressions, on the argument that a comment marker inside a string literal such as a URL is a case regex approaches get wrong. Truncation is anchored so the first page begins at the first line of the first selected file and the last page ends at the last line of the last selected file. And business logic stays in core rather than in the Electron layer.

Installing CodeSucker from Releases, and the first export

There is no package manager install. The README points at GitHub Releases and lists three installers for v0.5.1: CodeSucker-0.5.1-mac-arm64.dmg, CodeSucker-0.5.1-mac-x64.dmg and CodeSucker-0.5.1-win-x64.exe. Each release also ships a SHA256SUMS.txt for verifying the download. The macOS builds are not signed or notarised with an Apple Developer ID, and the README says signing and notarisation will come in a later version. If Gatekeeper blocks the first launch, the documented path is to open it once, then go to System Settings, Privacy and Security and choose Open Anyway next to the warning. If macOS instead reports the app is damaged or should be moved to the trash, the README gives this command, which removes only the download quarantine flag from that one application bundle:

bash
xattr -rd com.apple.quarantine /Applications/CodeSucker.app
open /Applications/CodeSucker.app

To run from source instead, the README gives the clone and dev commands, and notes a mirror workaround for Electron binary downloads on slow networks:

bash
git clone https://github.com/fanbuz/codesucker.git
cd codesucker
npm install
npm run dev        # 启动桌面应用
npm test           # core 流水线冒烟测试
npm run verify     # 版本一致性 + 测试 + 完整构建

The package.json in the repository root requires Node 22.12.0 or newer, and its verify script chains version:check, lockfile:check, icons:check, licenses:check, test, build and test:integration. For a first real use, drag a project folder into the app, tick the files to include in step two and drag them into order with the entry file first, enter the full software name and version in step three, review the before-and-after diff, walk the 60-page preview in step four, then read the audit report in step five before exporting docx or txt.

What the cleaner does to comments, strings and long lines

The cleaning stage is the part most likely to surprise you. It recognises comment and string boundaries character by character, so the README's example is that the double slash inside a URL string is not deleted. Default suffix support covers more than 50 extensions, including Pascal, PowerShell, Visual Basic, R, HCL and Terraform, Groovy and Gradle, and Windows Batch, alongside the mainstream languages. Blank lines are removed and tabs converted to spaces, both of which the README marks as switchable. Lines longer than 78 columns are hard-wrapped. A separate pass replaces API keys, passwords, internal IP addresses and phone numbers with placeholders. The README is explicit that suffixes which could belong to several languages, such as .m, .inc, .cls and .v, are not auto-classified by language family, and that .m is still treated as Objective-C. That is a deliberate refusal to guess, and it means a MATLAB or Mathematica file with a .m extension will be cleaned under the wrong grammar. If your project is mostly such files, this tool is the wrong one, and no setting in the README changes it.

Truncation, pagination and the audit report

The 60-page rule is implemented as a fixed cut. Projects over 3000 lines are reduced to the first 1500 and last 1500 lines, then split into 50-line blocks in memory with explicit page breaks, so every page is guaranteed 50 lines rather than filled by layout. Page one is the start of the first selected file and page 60 is the end of the last, which is why file ordering in step two changes the result. The docx export writes the header as software name plus version number, adds automatic page numbering in the top right through a PAGE field, and uses SimSun at 10.5pt with fixed line spacing. A txt copy is exported alongside, plus a JSON summary of third-party code clues that the README says contains no source body and no absolute paths. The audit checks effective content, lines per page, whether the last page is at least two thirds full, header consistency, first and last page boundaries, and conflicts between @author or Copyright notices and the copyright holder you entered. It returns one of three levels: pass, warning, or rejection risk. The README's own FAQ says the docx is a preparation draft and that you should clear every rejection risk and check the registration authority's current requirements before submitting. Treat the three-level verdict as a checklist, not a guarantee.

Third-party clues are hints, not conclusions

The third-party stage parses Node.js, Java and Kotlin, Go, Rust and Python dependency manifests entirely on the machine, then combines that with third-party directories, licence declarations and generated-code markers to produce evidence, a confidence level and a filtering suggestion. The README is careful about what this is not: it does not decide code ownership, and it gives no legal conclusion about copyright or licence compliance. A dependency manifest only proves the project uses a dependency, and an SPDX tag may be the project's own open source declaration. Ordinary @author or Copyright notices do not trigger this stage at all; they are carried through to the pre-submission audit, where they are compared against the copyright holder you typed, and only for the files that end up in the final pagination. The statistics area shows a count of items awaiting verification, and the README states that files are never excluded automatically without confirmation. That is the right default, but it also means the clue list is work for you, not a finished answer.

Encoding, size limits and how files disappear from a scan

The scanner identifies encoding before deciding whether a file is binary. Supported encodings are UTF-8, UTF-8 with BOM, GBK, GB18030 and UTF-16LE and UTF-16BE; LF, CRLF and CR are all treated with the same line semantics. A single source file is scanned for content only if it is 2 MiB or smaller, and anything larger is listed as over the size limit with its actual size and a suggested action. Empty files, real binaries, unrecognised encodings and read or decode failures all stay in the scan report with a reason rather than vanishing. The README also explains a subtlety in the summary count: the candidate figure means files that survive the scan exclusion rules and carry a supported suffix, and files pruned early, for example inside an excluded node_modules, are not invented as per-file candidates. The interface shows how many exclusion rules were applied, and files matched by .gitignore are listed with that exclusion reason. Default scan exclusions are maintained per machine in the settings page and stack independently on top of each project's .gitignore. If you need a per-file account of everything on disk, this summary is not that.

Project state, licence and what a rescan invalidates

The last push to the repository was on 2026-08-23, and v0.5.1 was released the same day, following v0.5.0 on 2026-08-22 and v0.4.5 on 2026-08-21. The repository is not archived. The licence is Apache-2.0, declared in package.json and in the LICENSE file, with a NOTICE file, a THIRD_PARTY_NOTICES.txt and a licenses directory present at the top level. Apache-2.0 is permissive and includes a patent grant; it also carries notice and attribution obligations for redistributed derivative works. Whether those obligations apply to your distribution is a question for your own counsel, and the README states plainly that the tool is not legal advice. On upgrade cost, the README describes a rescan action for when source changes outside the application: it keeps the current project configuration and unsaved edits, and invalidates old processing, pagination, audit and export results immediately. That is the honest behaviour, but it means a rescan discards the pagination you already reviewed, so plan to rescan before the audit step rather than after. Project selections and export configuration are stored in .codesucker.json; application-level rules, recent projects and window state live in the local config directory. Pinning and removing recent projects affects only the list, not the folders on disk.

Editorial conclusion

Adopt CodeSucker if you are preparing a Chinese software copyright registration and want the 60-page, 50-line-per-page, header-and-signature rules applied by a local pipeline instead of by hand. Do not adopt it if you need a signed macOS build, a headless CLI, or a tool that decides licence and authorship questions for you; the repository states that core is plain TypeScript with no Electron dependency and could be reused as a CLI later, but no CLI exists today. Before relying on it, run npm run verify against the tag you intend to build, confirm the SHA256SUMS.txt entry for your installer, and read the step 5 audit report rather than assuming the docx is submission-ready.

Frequently asked questions

Does CodeSucker upload my source code anywhere?

The README states that scanning, cleaning, layout and export make zero network requests, so the source never leaves the machine. The only outbound request it describes is a startup check for the latest public GitHub Release metadata, which sends no source, project paths, configuration or identity data, and a failure there does not affect core processing.

Can I submit the docx that CodeSucker generates directly?

The README's FAQ describes the docx as a preparation draft for the source program identification material, not a submission. It says you should still read the step 5 report, clear every rejection risk, and check the registration authority's latest requirements and your own applicant details.

Which encodings, line endings and file sizes can CodeSucker scan?

It supports UTF-8, UTF-8 with BOM, GBK, GB18030, UTF-16LE and UTF-16BE, and treats LF, CRLF and CR with the same line semantics. A source file is content-scanned only when it is 2 MiB or smaller; larger files, empty files, binaries, unrecognised encodings and read or decode failures are listed in the scan report with a reason.

Why does macOS say CodeSucker cannot be verified or is damaged?

The README states that the current macOS installers are not signed with an Apple Developer ID or notarised, and that signing and notarisation will be added in a later version. It documents opening the app once and then using System Settings, Privacy and Security to allow it, and provides an xattr command to remove the download quarantine flag if macOS still reports the app as damaged.

Official sources

  1. fanbuz/codesucker on GitHub
  2. Issues
  3. License: Apache-2.0
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/fanbuz-codesucker.svg)](https://hysenlabs.com/projects/fanbuz-codesucker)