# zealotCE/America-Against-America: a community OCR transcription of Wang Huning's 1991 book

> The repository is a text-recovery project, not software: it turns a poor PDF scan of Wang Huning's 美国反对美国 into Markdown, EPUB and print layouts. Useful to readers and archivists, not to anyone looking for a library or an API.

**zealotCE/America-Against-America** — 《美国反对美国》是王沪宁先生在上世纪80年代末赴美观察写作的。我们知道在那个年代中国对西方特别是美国的追捧有多高，所以突然看到一个学者在80年代就有如此清楚的认识，十分钦佩。由于网上只有效果很差的PDF扫描版，所以我想利用OCR技术和肉眼（人体OCR）来转成现代化的文本格式。目前已经全部完成。

- Repository: https://github.com/zealotCE/America-Against-America
- Stars: 3,069 · Forks: 366
- Language: Unknown
- License: not declared
- Published: 2026-09-24 · Updated: 2026-09-24 · Language: en
- Canonical page: https://hysenlabs.com/projects/zealotce-america-against-america

## What the repository actually is: a text-recovery release, not a codebase

The README opens with the author, 王沪宁, and a publication year of 1991, then explains the motivation in plain terms: the only copies circulating online were poor-quality PDF scans, so the maintainer set out to produce a modern text format using OCR plus manual proofreading. The phrase used in the README is 人肉OCR, human OCR, because the scan quality defeated the software. Transcription and correction are both reported at 100%.

That framing matters when you evaluate the project. There is no package to install, no command line, no configuration file, no service to run. The top-level directory holds a README, a Markdown file, an EPUB, a DOCX print-proof version, a PDF, a table-of-contents Markdown file, a print-layout folder named 打印适配-A5, and an archive folder named 归档. The deliverable is a book in several formats. If you arrived expecting a tool that performs OCR, you are in the wrong repository: this one is the output of that process, not the process itself.

## Who needs this: readers blocked by scan quality and by download friction

The audience is narrow and specific. First, readers who want the text in a form they can search, quote and annotate rather than a photographed page. Second, people inside mainland China, whom the README addresses directly with mirror options because the repository itself may be inconvenient to reach: a Baidu Cloud link with extraction code x79x, a Shimo document that can be read in the browser, and a GitBook instance at fakegoat.com/aaa/. Third, anyone assembling a plain-text corpus who would rather start from corrected Markdown than run OCR themselves.

The README also notes that individual typos may remain and invites feedback, or a clone-and-push, which tells you the maintainer treats correction as an ongoing, community-facing task rather than a finished artifact. The repository is not archived and its last push was on 2026-05-14, so the project has been touched within the last six months, though no releases are listed and the README does not describe a versioning scheme.

## How the text was produced and how the repository is laid out

The stated pipeline has two stages: automated OCR over the scanned PDF, then human review of the output. The README credits 肉眼, the naked eye, for the parts software could not handle, which is consistent with a book from 1991 whose scan is described as low quality. No OCR engine, model or script is named anywhere in the README, and no OCR configuration is present in the top-level entries, so the tooling behind the transcription is not reproducible from this repository. What you get is the corrected result.

The layout reflects a publishing workflow rather than a software one. 美国反对美国.md is the primary text. 美国反对美国.epub is the e-reader build. 美国反对美国-打印校对版0.1.0.docx carries a version number, 0.1.0, in its filename, which is the only version marker visible at the top level. 打印适配-A5 is a folder for A5 print adaptation, and 归档 holds archived material. 美国反对美国-目录.md is the table of contents as its own file, which is convenient if you want to generate navigation or split the book by chapter. 美国反对美国-wolf.pdf appears to be a PDF build, distinct from the original scan the project set out to replace.

## Getting a readable copy: three routes, one of which is a git clone

Because the project ships files rather than a program, the first real use is simply obtaining the text. The most direct route is cloning the repository and opening the Markdown or EPUB. The default branch is master.

```bash
git clone https://github.com/zealotCE/America-Against-America.git
cd America-Against-America
ls
```

After the clone you should see the top-level entries listed in the repository, including README.md, 美国反对美国.md, 美国反对美国.epub, 美国反对美国-目录.md and the two folders 归档 and 打印适配-A5. If you only want to read, open 美国反对美国.md in any text editor or Markdown viewer; the table of contents lives in its own file, so you can jump by chapter without scrolling the whole book.

For readers who cannot conveniently download from GitHub, the README gives three alternatives: the Baidu Cloud link with extraction code x79x, the Shimo document at shimo.im/docs/jt83JQX9W3ccyPHD, and the GitBook at fakegoat.com/aaa/. The README presents the Shimo and GitBook options as directly viewable, which means no account and no file transfer for the GitBook route. If you want a print copy rather than a screen copy, the 打印适配-A5 folder and the DOCX print-proof file are the relevant artifacts, and the README does not document a build step for either.

## The honest limitations: no licence, no OCR pipeline, and a correction pass that is not a critical edition

The most consequential gap is licensing. The repository does not state a licence, and the README says nothing about rights. The underlying book is a 1991 publication by a named author, so redistribution is a legal question the project does not address. If you plan to republish, mirror or sell anything derived from these files, that is the first thing to resolve, and it is not resolvable from the repository contents alone.

Second, the transcription pipeline is not reproducible. No OCR tool, script or parameter set appears in the README or the top-level entries, so you cannot re-run the process against a better scan or extend it to another book. Third, the correction claim is 100% complete, but the README itself qualifies it: individual typos may remain, and the intended remedy is a report or a pull request. That is a reasonable posture for a reading copy and a weak one for citation. There are no page numbers or scan-image references described, so checking a quotation against the original means going to the PDF yourself. Fourth, the README records a naming decision: after a reader pointed out in issue 6 that against fits the meaning better than the original Oppose, the maintainer adopted America against America as the English title while leaving the repository name unchanged. Anyone searching for the older name will land on the same project under a different label.

## Alternatives, and where the difference actually lies

The obvious alternative is the scanned PDF itself. A scan preserves the physical page, including pagination and any typographic detail, and it cannot introduce transcription errors. Its cost is that it is not searchable, not quotable without retyping, and, per the README, of poor visual quality. This repository trades fidelity to the page for a text you can grep, copy and convert.

A second alternative is doing the OCR yourself from a scan with a general-purpose engine. That gives you control over the pipeline and the ability to regenerate the text when better models appear. What it does not give you is the manual review pass, which is the part the README emphasizes and the part that costs the most human time. The difference between this project and a raw OCR run is precisely that review, and the difference between this project and the original scan is searchability. A third option, if your interest is analytical rather than textual, is secondary literature about the book; this repository offers the primary text and nothing else, with no introduction, annotation or translation.

## Maintenance and upgrade cost, and what the missing licence means for reuse

Maintenance here is editorial, not technical. There is no dependency graph to update and no runtime to patch, so the ongoing cost is the correction backlog: reading reports, reviewing proposed fixes, and pushing corrections to the Markdown, EPUB and print files in step. The README invites exactly that workflow. The last push was on 2026-05-14, and the repository is not archived, so the project has not been abandoned, but there are no releases, which means consumers have no tagged version to pin. If you build anything on top of the text, record the commit hash you cloned, because the files can change in place.

On licensing, the repository states none, and the README does not discuss terms. That is a factual gap, not a judgement about legality, and it should shape what you do with the files. Reading a clone locally and citing it in a paper are different activities from bundling the EPUB into a product or republishing the DOCX. The README's own distribution choices, a Baidu Cloud share, a Shimo document and a GitBook, suggest the maintainer's intent is broad access, but intent is not a licence. Treat the absence of a licence as the open question it is.

## Conclusion

Adopt it if you want a readable, redistributable text of 美国反对美国 rather than a scan, and if you can accept that the README describes the correction pass as complete but not error-free. Do not adopt it if you need a scholarly critical edition with page images, a translation, or any kind of software component, because the repository is a book release and nothing else. Before citing it, diff the Markdown against the 美国反对美国-wolf.pdf scan for the passages you quote, and check the licence, which the repository does not state.

## FAQ

### What is the zealotCE/America-Against-America repository?

It is a community transcription of Wang Huning's 1991 book 美国反对美国, produced with OCR plus manual proofreading and released as Markdown, EPUB, DOCX and PDF files. The README reports transcription and correction both at 100%, while noting that individual typos may remain.

### How do I get the America Against America text if I cannot download from GitHub?

The README lists three alternatives for readers who find GitHub inconvenient: a Baidu Cloud link with extraction code x79x, a Shimo document at shimo.im/docs/jt83JQX9W3ccyPHD, and a GitBook at fakegoat.com/aaa/. The README describes the Shimo and GitBook options as directly viewable.

### Why is the English title America Against America rather than America Oppose America?

The README explains that a reader raised the wording in issue 6 and that against fits the meaning better than the original Oppose, so the maintainer uses America against America as the English title. The repository name was left unchanged because the project had already been uploaded.

### Is there an English translation of America Against America in this repository?

The repository contains the Chinese text in Markdown, EPUB, DOCX and PDF formats, along with a table of contents file and print-layout folders. The README does not mention an English translation.

### Does the repository explain how the OCR was performed?

No. The README says the text was produced with OCR technology plus 人肉OCR, manual proofreading, but it names no OCR engine, script or parameters, and no such tooling appears among the top-level entries. The process is therefore not reproducible from the repository.

## Sources

- [Issues](https://github.com/zealotCE/America-Against-America/issues)
- [README](https://github.com/zealotCE/America-Against-America/blob/master/README.md)
- [zealotCE/America-Against-America on GitHub](https://github.com/zealotCE/America-Against-America)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/zealotce-america-against-america
