# kanasimi/work_crawler: a CeJS crawler that turns novel and comic sites into epub files

> work_crawler is a JavaScript downloader for online novels and webcomics, with a GUI, a CLI and an API. It ships a Dockerfile and a BSD-3-Clause licence, but the README leaves installation, rollback and site coverage largely unexplained.

**kanasimi/work_crawler** — Download comics novels 小说漫画下载工具 小説漫画のダウンローダ 小說漫畫下載:腾讯漫画 大角虫漫画 有妖气 咪咕 SF漫画 哦漫画 看漫画 漫画柜 汗汗酷漫 動漫伊甸園 快看漫画 微博动漫 733动漫网 大古漫画网 漫画DB 無限動漫 動漫狂 卡推漫画 动漫之家 动漫屋 古风漫画网 36漫画网 亲亲漫画网 乙女漫画 webtoons 咚漫 ニコニコ静画 ComicWalker ヤングエースUP モアイ pixivコミック サイコミ;アルファポリス カクヨム ハーメルン 小説家になろう 起点中文网 八一中文网 顶点小说 落霞小说网 努努书坊 笔趣阁→epub.

- Repository: https://github.com/kanasimi/work_crawler
- Stars: 4,244 · Forks: 359
- Language: JavaScript
- License: not declared
- Published: 2026-09-23 · Updated: 2026-09-23 · Language: en
- Canonical page: https://hysenlabs.com/projects/kanasimi-work-crawler

## What work_crawler downloads, and who it is for

The problem is serialised fiction and comics published inside a browser. A novel on 小説家になろう or 起点中文网 is a chain of chapter pages; a comic on ニコニコ静画 or 動漫之家 is a chain of page images. Reading offline means either a screen-scraping script per site or a paid reader. work_crawler is a single JavaScript program that knows the page structure of a long list of sites and writes the result to disk, with novels exported to epub.

The audience is narrow and specific. The README's own description is "Tools to download novels (→ epub) and comics", and the site list spans Chinese, Japanese and Korean services: 腾讯漫画, 快看漫画, webtoons, 咚漫, ComicWalker, ヤングエースUP, pixivコミック, サイコミ, アルファポリス, カクヨム, ハーメルン, 小説家になろう, 起点中文网, 笔趣阁. If your reading list is English-language comics from a Western publisher, this tool has little to offer you. If it is Japanese web novels, it is aimed directly at you.

Three interfaces are listed as supported: GUI, CLI and API. That matters more than it sounds. A GUI downloader is a desktop app you click; an API is something you can call from your own script. work_crawler claims both, which is unusual for this category.

## How the crawler is put together: CeJS, per-site modules and an epub writer

The repository layout tells most of the story. Top-level directories are named by language and medium: comic.cmn-Hans-CN, comic.cmn-Hant-TW, comic.en-US, comic.ja-JP, novel.cmn-Hans-CN, novel.ja-JP, book.cmn-Hant-TW. Each directory holds the site definitions for that language and medium, so adding a site means adding a file in the matching directory rather than editing one monolithic scraper.

The engine underneath is CeJS, listed as the first keyword in package.json and named in the package title, "CeJS online novels and comics downloader". The entry points are work_crawler_loader.js and work_crawler.default_configuration.js, with work_crawler.updater.js used to fetch the application itself. The package manifest sets main to gui_electron/gui_electron.js, so the default program is the Electron desktop window, not a terminal script.

Data flow is conventional for a crawler: resolve the work's page, enumerate chapters or pages, fetch each one, assemble. The novel path ends in epub, which is why epub appears as a keyword and why the description says "novels (→ epub)". The comic path writes image files. The README shows screenshots of a search-and-download flow and of a settings panel with many download options, and mentions an optional dark theme.

The limitation here is that site definitions age. A scraper is only as good as its knowledge of the target's HTML, and the README gives no compatibility matrix, no list of which sites currently work, and no statement about how breakage is handled. The per-language directories are the place to look, but you have to read them yourself.

## Installing work_crawler and downloading a first novel

The README itself does not contain install instructions. It links to four documents, one per language, and those are where setup is described: document/README.en-US.md, document/README.cmn-Hant-TW.md, document/README.cmn-Hans-CN.md and document/README.ja-JP.md. Start there rather than guessing.

Two paths are visible in the repository files. The first is npm plus Electron, driven by the scripts block in package.json. From a checkout, install dependencies and launch the desktop GUI:

```bash
npm install
npm start
```

npm start runs node_modules/.bin/electron ., which opens the Electron window defined by gui_electron/gui_electron.js. The postinstall script runs electron-builder install-app-deps, so expect a native-dependency step during install. The pack and dist scripts build a packaged application into the build directory, with productName work_crawler and appId org.kanasimi.work_crawler.

The second path is the Dockerfile, which builds from node:12, sets WORKDIR /app, copies work_crawler.updater.js and runs it with node, exposes port 80, and starts the GUI with sh start_gui_electron.sh inside work_crawler-master:

```dockerfile
FROM node:12
WORKDIR /app
COPY work_crawler.updater.js /app
RUN ["node", "work_crawler.updater.js"]
EXPOSE 80
CMD ["sh", "-c", "cd work_crawler-master && sh start_gui_electron.sh"]
```

The Dockerfile comments give the pull and run commands, including docker pull kanasimi/work_crawler and docker run -it --rm --name kanasimi/work_crawler. Note that the base image is node:12, which is old; if you build this yourself, that pin is the first thing you will want to examine.

On Windows and macOS there is also start_gui_electron.bat and start_gui_electron.sh at the repository root. The README's OS table marks Windows, macOS and UNIX/Linux as supported; Android appears only as a commented-out row, so treat it as unsupported. Once the GUI is open, the README's screenshot shows searching across sites and downloading a work with one click. For a first run, pick one title you already own or can legally read, and confirm the output file opens in your epub reader before queuing anything longer.

## Where work_crawler breaks or is the wrong tool

The most obvious failure mode is a site redesign. Nothing in the README describes error handling for a changed page layout, and the project's own structure, one directory of site definitions per language and medium, implies that each broken site needs a code change. If you depend on one specific site that is not listed in the description, verify it exists in the matching comic.* or novel.* directory before installing anything.

The second issue is provenance. The README does not state which sites permit downloading, and the repository carries no legal notice beyond the BSD-3-Clause licence in package.json, which governs the code and says nothing about the content you fetch. Downloading a licensed comic from 腾讯漫画 or a paid webtoon series is a different act from archiving a freely published web novel, and the tool does not distinguish between them.

The third is release cadence. The newest release listed is v2.14.0 from 2022-08-13, with v2.13.0 in 2021 and v2.12.0 in 2021 before it. The last push to the repository was on 2026-08-03, so development has continued between releases, but anyone pinning to a tagged release is pinning to a 2022 build. The README documents no rollback procedure and no version compatibility statement, so downgrading after a bad update is not something the project tells you how to do.

Finally, it is the wrong tool for bulk archival at scale. There is no documented rate-limit configuration, no proxy setting, and no queue persistence described in the README. For a few dozen chapters this is fine. For a ten-thousand-chapter novel, you are relying on undocumented behaviour.

## How work_crawler differs from gallery-dl and similar downloaders

gallery-dl is the closest widely used alternative, and the difference is architectural rather than cosmetic. gallery-dl is a Python command-line program with a large set of site extractors and a configuration file; it is built around the terminal and around files landing in a directory tree. work_crawler is a JavaScript application built on CeJS with an Electron GUI as its default entry point, and it treats novels as a first-class output format, writing epub rather than chapter text files.

That epub pipeline is the real distinction. Turning 小説家になろう or 起点中文网 chapters into a single epub with metadata and reading order is a different job from saving pages, and gallery-dl does not target it. Conversely, gallery-dl's extractor set is broader for image sites and its configuration is documented in the open, which is exactly what work_crawler's README lacks.

A second comparison is writing your own script. For a single site, a few dozen lines of Node or Python against the chapter list will be faster to debug than learning work_crawler's configuration, and you will know exactly what breaks when the site changes. work_crawler pays off when you read across many sites and want the same output format from all of them.

## Licence, maintenance and the cost of keeping site definitions current

package.json declares "license": "BSD-3-Clause", so the code is permissively licensed and you can modify and redistribute it under that licence's terms. The repository's own metadata does not repeat the licence, which is why the manifest is the place to check. This is a statement about the source code only; it grants no rights over the comics and novels you download, and the README does not address that question at all. If you plan to redistribute anything you fetch, that is a separate legal question and not one the project answers.

Upgrade cost is the part to budget for. The release history shows a gap: v2.12.0 in April 2021, v2.13.0 in May 2021, v2.14.0 in August 2022, and nothing tagged since, while commits continued up to 2026-08-03. In practice that means the useful code may be ahead of the newest release, and a bug fix you need may only exist on master. The Dockerfile pins node:12, an end-of-life Node line, so a container build is a maintenance item on its own.

The recurring cost is site drift. Every supported site is a definition file in comic.* or novel.*, and each one can break independently. There is no test suite mentioned in the README and no compatibility table, so the only way to know whether your site still works is to try it. Budget for that, and expect to read JavaScript when something fails.

## Conclusion

Adopt work_crawler if you read Japanese, Chinese or Korean web fiction and webcomics on the supported sites and want the result as an epub rather than a folder of images. Do not adopt it if you need a documented, versioned install path or a supported API surface; the README points at per-language documents and the repository's own files, and the newest release listed is v2.14.0 from 2022-08-13, while the last push was on 2026-08-03. Before committing, open document/README.en-US.md, run the Dockerfile build or the npm start script, and confirm on one title that the site you actually use is still covered.

## FAQ

### What is work_crawler used for?

It is a tool for downloading online novels and comics in bulk, with novels exported to epub. The README describes it as "Tools to download novels (→ epub) and comics" and lists Chinese, Japanese and Korean sites among its targets.

### How does the work_crawler crawler work?

It resolves a work's page, walks its chapters or pages, fetches each one, and assembles the result, writing novels as epub and comics as image files. The engine is CeJS, and site definitions live in per-language directories such as comic.ja-JP and novel.cmn-Hans-CN.

### Where are the work_crawler installation instructions?

The main README does not contain them. It links to four language-specific documents, including document/README.en-US.md for English, and those are where setup is described.

### Can work_crawler run in Docker?

Yes. The repository includes a Dockerfile that builds from node:12, runs work_crawler.updater.js during the build, exposes port 80, and starts the GUI with start_gui_electron.sh. The Dockerfile comments include docker pull kanasimi/work_crawler and a docker run example.

### Which operating systems does work_crawler support?

The README's OS table marks Windows, macOS and UNIX/Linux as supported. Android appears only as a commented-out row, so it is not listed as supported.

### What licence does work_crawler use?

package.json declares "license": "BSD-3-Clause" for the code. The repository metadata does not state a licence, and the licence covers the software, not the comics or novels you download with it.

## Sources

- [Issues](https://github.com/kanasimi/work_crawler/issues)
- [kanasimi/work_crawler on GitHub](https://github.com/kanasimi/work_crawler)
- [README](https://github.com/kanasimi/work_crawler/blob/master/README.md)
- [Releases](https://github.com/kanasimi/work_crawler/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/kanasimi-work-crawler
