# Lightnovel Crawler ships the command line tool and the server as one distribution, and its documentation links to a branch that is not the default

> lncrawl/lightnovel-crawler crawls 361 sources across 11 languages with 446 crawlers, turns them into e-books in eighteen formats, and doubles as a multi-user server with accounts, translation and scheduled re-downloads. The CLI install pulls a web framework, a database layer and a translation engine, the container ships a browser, and the default compose stack publishes its database on every interface.

**lncrawl/lightnovel-crawler** — Generate and download e-books from online sources.

- Repository: https://github.com/lncrawl/lightnovel-crawler
- Website: https://lncrawl.bitanon.dev/
- Stars: 2,638 · Forks: 504
- Language: Python
- License: GPL-3.0
- Published: 2026-09-28 · Updated: 2026-09-28 · Language: en
- Canonical page: https://hysenlabs.com/projects/lncrawl-lightnovel-crawler

## The documentation links to a branch that is not the default

The default branch of this repository is the development one, and most of the page's own links point somewhere else. The source list, the changelog, the licence file and the contributing guide are all linked through a master path, so a reader following the documentation arrives at a branch that is not where work happens. The install section makes the same split explicit. One command installs from the repository's stable ref, and a second installs a tarball of the development branch, labelled as the newest fixes with no stability promise. So there are two remotes in one repository's instructions, one for the released line and one for the moving one, and the documentation for the moving line is the part that lives on master. The releases themselves are tagged with a three part version and the newest is 4.14.0, which is where a reader looking for a stable reference should start.

## 446 crawlers serve 361 sources, and both numbers are generated

The headline count is a single sentence: 361 sources across 11 languages, served by 446 crawlers. Two details sit behind it. The two counts differ, which tells you the unit of the second number is not the unit of the first: more crawlers than sources means some sites are covered by more than one crawler, presumably for different layouts or devices. And the sentence sits between two markers in the page source that identify it as generated, and the full source list in the repository is described as regenerated by continuous integration on the same basis. That is the right way to keep a list this long honest, and it has a consequence for a reader: the number can change in a commit that touches nothing but the page. It also means the source list is not something you can review in a pull request and expect to stay put.

## The command line tool arrives with the whole server

The page says to pick one of three installs and that all three ship the same engine, the same sources and the same web app. The dependency list is where the cost of that sentence becomes visible. Alongside the crawling libraries, a parser and an HTML toolkit, an ebook writer and an HTTP client with compression and HTTP two, the manifest carries a web framework with its standard extras, an ORM, a database migration tool, a PostgreSQL driver, password hashing with argon two, a JSON web token library with cryptography, an IMAP library, an image library, a gRPC client for an automation service, a translator package and a scraper package with all extras enabled. None of that is needed to write one EPUB file. It is there because the CLI and the server are one distribution, so the smallest install is the largest one.

## The container ships Firefox and a desktop library stack

The image build installs a graphical toolkit, a sound library, several X11 libraries and two font packages, then downloads a browser from Mozilla during the build and symlinks it into the path. The comment above the browser step says what it is for, spoofing, which in this context means presenting as an ordinary browser to the sites being fetched. Two details follow from how that download is written. The download asks for the latest secure build rather than a version, so the browser inside the image floats with each rebuild. And the archive choice is decided by processor architecture, with a mapping for amd64 and for arm64 and an empty result for anything else, in which case the browser step is skipped and the image still builds. A crawler image carrying a browser and a desktop stack is a large image, and it is also the reason the build takes what it takes.

## The compose stack publishes its database on every interface

The stack runs five services. Two of them are deliberately kept off the network: the translation service is not published because it has no authentication and is reached only through an admin-only path on the server, and the proxy pool keeps its ports on loopback because there is no TLS on it. The database does the opposite. Its port is published on all interfaces rather than on loopback, its credentials are written in the compose file as a user and a password, and the same pair is repeated inside the server's connection string, so anyone who can reach the host on that port has the credentials from a file that is meant to be committed. The password is also the one thing a reader is most likely to copy unchanged. The rest of the stack is a normal local arrangement, with a named volume for data, a fixed timezone, and a host gateway entry for reaching the host machine.

## Three installs that promise the same three things

The install section offers a standalone build per platform, a pip install, and an install straight from the repository. The standalone route needs no Python at all and downloads a binary for Windows, Linux or macOS from the project's own download host, and running it opens a desktop application with no login. The pip route requires Python 3.9 or newer and gives the command line tool and the server together, with fallbacks spelled out for when pip fails and for when the command is not on your path. Two commands cover the whole download path:

```bash
pip install -U lightnovel-crawler

lncrawl crawl "https://example.com/novel/page" -f epub --all
```

The repository route is the one that trades stability for freshness. What all three share is the engine, the sources and the web app, which means a reader comparing them is choosing a distribution channel rather than a feature set, and older standalone builds are kept on the releases page.

## The version lives in a file the build reads, not in the manifest

The developer tasks are driven by a makefile whose phony target list covers the usual ground, from version bumps to wheel builds, standalone executable builds, installer builds, Docker targets and dependency add and remove helpers. The version itself is read from a plain text file inside the package, with a different command for Windows shells than for everyone else, and it is echoed by a target whose only job is to print it. Two release targets begin by checking that the working tree has no uncommitted changes, which refuses to tag a dirty checkout. Dependency syncing uses one shared set of flags across every target so an install and a sync cannot drift apart. It is a careful makefile for a project whose users mostly never see it, and the file based version is the one decision worth knowing about when you script against a release.

## A challenge page served as success is treated as a failure

One row in the feature table is about refusal rather than about features, and it is the most revealing line on the page. Blocks are diagnosed rather than retried, and a challenge page that arrives with a success status is a failure, not a chapter. In practice that means the crawler will not treat an interstitial as content and will not keep hammering a site that has answered with a block, which is a design decision with consequences in both directions: fewer accounts get flagged, and more downloads come back empty with a reason attached. The same table also carries the project's own terms, personal backups of content you have legitimate access to, with redistribution and resale ruled out. The rest of the table is ordinary product surface, from a browser reader with text to speech to per novel glossaries for translation.

## Conclusion

Use this tool to keep your own reading library backed up, and read it as two products in one package rather than one. The command line path is genuinely good for a single novel: give it a URL, a format and the all flag, and it writes a file. The server path is a different undertaking, with accounts, quotas, translation and a database behind it, and it arrives with the same install command. Before you run the server, replace the database credentials in the compose file and stop publishing that port on every interface, because the shipped defaults are written for a local stack. And keep the project's own framing in view, which is personal backups of material you can already reach, with blocks diagnosed rather than worked around.

## FAQ

### how to use light novel crawler

Install it with pip, then point the crawl command at a novel page with a format and the all flag, as in lncrawl crawl "https://example.com/novel/page" -f epub --all. It discovers the chapter list, fetches every chapter and writes the file.

### is lightnovel crawler safe

The project states it is for personal backups of content you already have legitimate access to, and that redistribution or resale is out of scope. The code is GPL. Its default compose stack ships database credentials in the file and publishes that port on every interface, so treat those defaults as local only.

### what is lightnovel crawler

A tool that turns a web novel into an e-book with one command, and also a self-hosted server with accounts, a browser reader, translations and scheduled re-downloads. Its source count line says 361 sources across 11 languages served by 446 crawlers.

### How many sites does Lightnovel Crawler support?

The readme's generated count says 361 sources across 11 languages served by 446 crawlers, and the full list in the repository's sources file is regenerated by continuous integration, so the number moves without a feature change.

### Which output formats does Lightnovel Crawler produce?

Eighteen of them. EPUB, TXT and JSON are produced directly and the rest go through Calibre, and one download can emit several formats at once, either per volume or for the whole novel.

### What does Lightnovel Crawler do when a site blocks it?

It diagnoses the block instead of retrying it, and a challenge page served with a success status counts as a failure rather than as a chapter, which is why some downloads come back empty with a reason attached.

## Sources

- [License: GPL-3.0](https://github.com/lncrawl/lightnovel-crawler/blob/dev/LICENSE)
- [lncrawl/lightnovel-crawler on GitHub](https://github.com/lncrawl/lightnovel-crawler)
- [Project website](https://lncrawl.bitanon.dev/)
- [README](https://github.com/lncrawl/lightnovel-crawler/blob/dev/README.md)
- [Releases](https://github.com/lncrawl/lightnovel-crawler/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/lncrawl-lightnovel-crawler
