# CodeSucker paginates 60 pages itself, and never counts lines in a laid-out doc

> CodeSucker turns a local project into the source listing Chinese software copyright registration asks for, fully offline and with no language model involved. Pagination is explicit rather than counted out of a rendered document, comment stripping is a character state machine rather than a regex, and the macOS build is unsigned.

**fanbuz/codesucker** — 拖入项目，生成 60 页软著源程序文档。离线抽取、注释清洗、自动排版，导出前先帮你过一遍审查——代码不出本机。

- Repository: https://github.com/fanbuz/codesucker
- Stars: 416 · Forks: 84
- Language: TypeScript
- License: Apache-2.0
- Published: 2026-09-20 · Updated: 2026-09-20 · Language: en
- Canonical page: https://hysenlabs.com/projects/fanbuz-codesucker

## Pagination is written into the text, not measured out of a layout

The filing rule the tool is built around is a specific one: front and back 30 pages, at least 50 lines per page, page numbers 1 to 60, page 1 showing the start of the program and page 60 the end. The usual way to hit that is to lay out a document and count what falls where, which is why the README calls out not relying on typesetting to fill pages.

CodeSucker cuts in memory at 50 lines per block and inserts an explicit page break character at each boundary, so the count is correct before the file is ever rendered. Fixed line spacing in the exported docx, 宋体 at 10.5pt, is described as a fallback so that changing fonts does not shift the breaks.

The head and tail are chosen rather than sampled. A file longer than 3000 lines contributes its first 1500 and last 1500. Page 1 is anchored to the first line of the first selected file, and page 60 closes on the last line of the last one, so the opening and closing of the listing match the opening and closing of the program.

That anchoring is why the file selection step exists. Reorder the files, put the entry point at the top, and the whole document changes, so ordering is an input rather than a detail.

## Comment removal is a state machine, because // inside a URL breaks regexes

The cleaning stage strips comments, deletes blank lines, converts tabs to spaces, and hard-wraps long lines at 78 columns. The interesting part is how it finds comment boundaries.

It scans character by character with a state machine rather than a pattern match. The reason given is that comment markers inside string literals are a known failure of the regex approach, and the README uses a URL in a string as the example: `//` in `"https://..."` must not be treated as the start of a comment.

That is the correct diagnosis. A pattern that matches `//` to end of line will happily eat half a line of code whenever a string contains a protocol, and in a real project with URLs in config constants that happens often enough to corrupt the listing.

Language coverage is listed by suffix rather than by parser, more than 50 of them covering Java, Kotlin, Python, JavaScript, TypeScript, Go, Rust, C, C++, C#, Swift, PHP, Ruby, Vue, HTML, CSS, SQL, Pascal, PowerShell, Visual Basic, R, HCL, Groovy and Windows Batch. Suffixes that could belong to several languages are not guessed: `.m`, `.inc`, `.cls` and `.v` are left alone, and `.m` is handled as Objective-C.

Sensitive data masking runs in the same stage, replacing API keys, passwords, internal IPs and phone numbers with placeholders.

## Third-party clues point at files and never exclude them

The tool parses dependency manifests for Node.js, Java and Kotlin, Go, Rust and Python on the local machine, then combines what it finds with third-party directories, licence headers and generated-code markers to give you evidence, a confidence level and a filtering suggestion.

What it refuses to do is the important part. Clues only prompt, and no file is excluded without confirmation. Unprocessed clues do not block export, and the set of files you tick is the only scope that reaches preview, pagination, audit and export.

The README is also careful about what a clue proves. A dependency manifest shows that a project uses something, and an SPDX identifier may be the project's own declaration rather than a third-party license file. Local workspace and path dependencies are not treated as third-party by default. An ordinary `@author` or `Copyright` line does not on its own trigger a third-party clue, because that declaration is what the submission audit checks separately.

The stated position is that the tool gives explainable hints and no legal conclusion about authorship or licence compliance, and it points to `docs/third-party-code-risk.md` for formats, limits and known blind spots.

## The pre-submission audit has three verdicts and checks the header

Before it exports, the tool audits and reports through three levels: pass, warning, and a return-for-correction risk.

The checks are specific. Valid content, lines per page, whether the last page is at least two thirds full, header consistency, the first and last page boundaries, and conflicts between `@author` and `Copyright` attributions and the copyright holder you entered in step three. A document missing the version number in its header produces a warning.

The attribution check is the one that can actually block a filing, and it works by scanning the whole text for `@author` and `Copyright` and comparing them against what you typed, scoped to the files that end up in the paginated output. That is a plausible source of friction in a project with more than one author string, which is most real projects.

The export itself writes the header with the software name and version, a page number field in the top right, and 宋体 at 10.5pt with fixed line spacing, plus a plain text copy for checking. A third-party clue JSON summary is attached without source text and without absolute paths, which is a deliberate privacy choice given the tool's offline claim.

## No language model runs at any stage, by explicit statement

For a tool that rewrites your source code, this is the claim worth checking rather than assuming, and the README addresses it directly.

Scanning, third-party clue analysis, file selection, cleaning, pagination, audit and export involve no local or cloud language model. None of the source, source fragments, project paths, dependency manifests or project configuration is sent to any model service, and the results come from built-in deterministic local rules.

It also states two things it will not do: it does not judge whether your source was generated by AI, and it does not reach automatic legal conclusions about code provenance, ownership or licence compliance.

That third point has a concrete consequence for your filing. If you have submitted AI-assisted code, this tool will not flag it, so the obligation to disclose remains yours. The zero-network claim is qualified in one place: the startup update check queries GitHub Releases for public version metadata, and the README says a failure there does not affect core function.

For the record on encoding, scanning covers UTF-8, UTF-8 with BOM, GBK, GB18030, UTF-16LE and UTF-16BE, treats LF, CRLF and CR identically, and caps files at 2 MiB, with anything larger, empty, genuinely binary or undecodable kept visible in the report with a reason instead of silently vanishing.

## The macOS build is unsigned, and the README tells you to override Gatekeeper

The macOS package has no Apple Developer ID signature and no notarization. The README is direct about the consequence and about the workaround: if Gatekeeper blocks the first open, open it once, then go to System Settings, Privacy and Security and choose to open anyway.

If the system still calls the app damaged, the documented route is to verify the download came from the project Release, check the SHA256, and remove the quarantine attribute:

```bash
xattr -rd com.apple.quarantine /Applications/CodeSucker.app
open /Applications/CodeSucker.app
```

The README says that command only removes the quarantine mark for this app, and tells you not to run it on an app from an unknown source. That caution is correct, because removing quarantine is exactly the step that turns an unverified binary into a running one.

One inconsistency is worth knowing before you download. The download table lists `CodeSucker-0.5.1-mac-arm64.dmg`, `CodeSucker-0.5.1-mac-x64.dmg` and `CodeSucker-0.5.1-win-x64.exe`, while the newest release is v0.5.2 and `package.json` reads 0.5.2. The filenames in the table are one version behind. Each release does ship a `SHA256SUMS.txt`, which is the file to check against.

Formal signing and notarization are described as coming in a later version.

## The core package has no Electron in it, which is why the pipeline is testable

The architecture is a workspace with two packages. `packages/core` is pure TypeScript with zero Electron dependencies and runs a six-stage pipeline: discover, third-party clues, clean, select, render, audit. `packages/app` is the shell, Electron 43 with React 18 and zustand built by electron-vite. The stated intent for the split is that the core could later be reused as a CLI or a web service.

That boundary shows up in the scripts. `npm test` runs the core and app test suites, and `npm run verify` chains version consistency, a lockfile registry check, an icon check, a third-party licence check, the tests, the build and an integration suite. Node 22.12.0 or newer is required. Running from source:

```bash
git clone https://github.com/fanbuz/codesucker.git
cd codesucker
npm install
npm run dev        # 启动桌面应用
npm test           # core 流水线冒烟测试
npm run verify     # 版本一致性 + 测试 + 完整构建
```

There is also a network note for users in China, since the Electron binary download can fail:

```bash
ELECTRON_MIRROR=https://npmmirror.com/mirrors/electron/ node node_modules/electron/install.js
```

The README cites the copyright registration rules and the China Copyright Protection Center's published review practice as its basis, and repeats in two places that it is not legal advice and that the registration authority's current requirements win.

## Conclusion

Adopt CodeSucker if you are filing Chinese software copyright registration and your project has clean licence boundaries, because the parts that normally fail a filing, pagination and comment stripping, are the parts this tool treats as engineering problems. Do not adopt it expecting legal assurance, since it states plainly that it reaches no conclusion on ownership or licence compliance and that its own output is a preparation draft rather than a submission. Verify first two things: that your `@author` and `Copyright` lines match the copyright holder you will file under, because the audit checks them, and that your macOS download really came from the project Release with a matching SHA256, since it ships unsigned.

## FAQ

### Is CodeSucker's generated document ready to submit for copyright registration?

It is a preparation draft. The README says to read the step five report, clear any return-for-correction risks before submitting, and follow the registration authority's current requirements, since the tool states it is not legal advice.

### Does CodeSucker upload my source code anywhere?

No. Scanning, cleaning, layout and export run entirely on your machine with deterministic local rules and no language model. The only network request is a startup check of public GitHub Release version metadata.

### How does CodeSucker decide which files to exclude as third-party code?

It does not exclude them. It parses dependency manifests and combines third-party directories, licence headers and generated-code markers into evidence with a confidence level and a suggestion. Files are never removed without confirmation, and unprocessed clues do not block export.

### Why is the CodeSucker macOS build blocked by Gatekeeper?

The package is not yet signed with an Apple Developer ID and not notarized. Verify the download came from the project Release, check the SHA256 against `SHA256SUMS.txt`, then allow it in System Settings under Privacy and Security or remove the quarantine mark.

### Which file encodings and sizes does CodeSucker scan?

UTF-8, UTF-8 with BOM, GBK, GB18030, UTF-16LE and UTF-16BE, with LF, CRLF and CR treated the same. Files must be at or below 2 MiB to enter content scanning, and larger, empty, binary or undecodable files stay listed in the report with a reason.

### Will CodeSucker tell me if my code was written by AI?

No. It states explicitly that it does not judge whether source was generated by AI, and that it gives no automatic legal conclusion about provenance, ownership or licence compliance. It also says the dependency manifest only proves a dependency is used, and that an SPDX identifier may be the project's own declaration.

## Sources

- [fanbuz/codesucker on GitHub](https://github.com/fanbuz/codesucker)
- [Issues](https://github.com/fanbuz/codesucker/issues)
- [License: Apache-2.0](https://github.com/fanbuz/codesucker/blob/main/LICENSE)
- [README](https://github.com/fanbuz/codesucker/blob/main/README.md)
- [Releases](https://github.com/fanbuz/codesucker/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/fanbuz-codesucker
