AI-reads-books-page-by-page: A Python Script for Page-Level PDF Knowledge Extraction
AI reads books: Page-by-Page PDF Knowledge Extractor & Summarizer. script performs an intelligent page-by-page analysis of PDF books, methodically extracting knowledge points and generating progressive summaries at specified intervals
At a glance
- What is it?
- AI-reads-books-page-by-page is a single-file Python script that processes a PDF book one page at a time, calls the OpenAI API to extract knowledge points from each page, accumulates them in a JSON knowledge base, and generates summaries at configurable intervals. It targets developers and researchers who want structured notes extracted from a book without reading it.
- Who is it for?
- AI-reads-books-page-by-page is the right tool for a developer or researcher who wants to extract structured knowledge points from a technical PDF using the OpenAI API, and who is comfortable reading and editing a single Python file to configure it. The critical setup step is placing the PDF in the script's directory and setting `PDF_NAME` before running.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 96 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Problem This Script Solves
Reading a long technical book and retaining its key points is time-consuming. LLM context windows have grown, but processing an entire book in a single API call is expensive and loses the sequential structure of the text. AI-reads-books-page-by-page takes a different approach: it sends each page individually, keeps a running list of extracted knowledge points, and generates periodic summaries so the cumulative understanding grows incrementally.
The project is a single file: `read_books.py`. The repository also includes two sample PDFs (`meditations.pdf` and `infinite_math.pdf`) that serve as test inputs.
The target user is a developer comfortable with Python and the OpenAI API, who wants to process a PDF book into a structured knowledge base and a set of Markdown summaries. It is not a polished application with a UI. Customization requires editing constants directly in `read_books.py`.
The README describes the project as part of a larger collection of AI-building tools associated with the echohive42 Patreon channel. The full source with walkthroughs is available to Patreon supporters; this public repository contains the script itself.
How the Script Works: Page Processing and Knowledge Accumulation
The script uses PyMuPDF (installed as `pymupdf`) to extract text from each page of the PDF. It sends each page's text to the OpenAI API using the `instructor` library, which wraps the OpenAI client to return structured Pydantic objects.
The response for each page is a `PageContent` Pydantic model with two fields: `has_content` (a boolean indicating whether the page contains substantive content) and `knowledge` (a list of extracted knowledge points). Pages identified as tables of contents, index pages, or similarly low-content pages are skipped by the `has_content` flag.
Extracted knowledge points are appended to a list and saved to a JSON file at `book_analysis/knowledge_bases/`. This file is updated after every page, so a run that is interrupted midway can be resumed: `load_existing_knowledge()` reads the saved JSON on startup.
At the interval configured in `ANALYSIS_INTERVAL` (a number of pages), the script calls the OpenAI API a second time using the `ANALYSIS_MODEL` to generate a cumulative summary of all knowledge points gathered so far. Summaries are saved as Markdown files at `book_analysis/summaries/`. Setting `ANALYSIS_INTERVAL = None` skips interval summaries and only produces a final summary.
The `TEST_PAGES` constant limits processing to a small number of pages, useful for verifying the setup before committing to a full run.
Setting Up and Running the Script
Install the dependencies first:
pip install -r requirements.txtThe requirements are `pydantic`, `openai`, `pymupdf`, and `termcolor`. The script uses the `instructor` library (not in the requirements file as published; the README describes it in the code walkthrough) through the OpenAI client integration.
Place the PDF in the same directory as `read_books.py`. Then open the file and update the `PDF_NAME` constant with the filename. The README gives these configuration constants:
- `PDF_NAME`: The filename of the PDF to process. - `ANALYSIS_INTERVAL`: Number of pages between interval summaries. Set to `None` to skip them. - `TEST_PAGES`: Number of pages to process in test mode. Set to `None` to process the full book. - `MODEL`: The OpenAI model used for per-page extraction. - `ANALYSIS_MODEL`: The OpenAI model used for generating summaries.
Run the script:
python read_books.pyThe script creates three output directories under `book_analysis/`: `knowledge_bases/` for JSON knowledge files, `summaries/` for Markdown summaries, and `pdfs/` for a copy of the input PDF. The README notes that `setup_directories()` clears previously generated files at the start of each run, except the knowledge base which supports resumption.
Terminal output uses `termcolor` to color-code progress messages, making it easier to follow page-by-page processing.
Directory Layout and Output Structure
The repository is minimal: `read_books.py`, `requirements.txt`, `LICENCE`, and two sample PDFs. There is no `setup.py`, no package structure, and no command-line argument parsing. Everything is configured by editing constants at the top of `read_books.py`.
The `BASE_DIR` constant sets the root directory for outputs (defaults to `book_analysis/`). Three subdirectories are created under it: `PDF_DIR` for the input PDF copy, `KNOWLEDGE_DIR` for the JSON knowledge base, and `SUMMARIES_DIR` for Markdown summaries.
Summary files are named with a convention based on whether they are interval summaries or final summaries. The `save_summary()` function takes an `is_final` boolean to set this. Interval summaries are numbered sequentially; the final summary gets a distinct filename.
The knowledge base is a flat JSON list of strings. Each entry is one knowledge point extracted from a page. The `save_knowledge_base()` function prints the count of items when saving, which serves as the progress indicator for how much has been extracted.
Limitations and Cases Where It Falls Short
The script depends on the OpenAI API. There is no local model option as published. Every page and every summary generation makes an API call, so processing a 300-page book at the default settings makes at least 300 API calls for page extraction plus additional calls for interval summaries. The cost scales directly with book length and the models chosen in `MODEL` and `ANALYSIS_MODEL`.
PDF text extraction quality varies. Pages that are scanned images rather than text-based PDFs will produce empty or garbled text from PyMuPDF, and the `has_content` filter may not correctly identify them. The `smart content filtering` mentioned in the README's feature list refers to the `has_content` flag in the Pydantic model, which relies on the AI to classify each page. If the API classifies a content-heavy page as empty, that page's knowledge is lost.
The script does not handle multi-column layouts, mathematical notation, or code snippets gracefully. PyMuPDF extracts text linearly; complex page layouts produce text that is out of reading order.
There is no progress bar or estimated completion time. The script prints page-by-page status to the terminal. For a 500-page book, that is 500 printed messages.
An alternative for structured information extraction from documents is LlamaIndex's document processing pipeline, which supports chunking strategies, metadata extraction, and multiple vector store backends. LlamaIndex is a general-purpose document processing framework; AI-reads-books-page-by-page is a single script specialized for linear page-by-page PDF analysis with a knowledge-accumulation pattern.
Resume Capability and the Knowledge Base JSON
The resume feature is the most operationally useful aspect of the script. If a run is interrupted mid-book (network error, API rate limit, keyboard interrupt), the knowledge base JSON contains all points extracted up to that page. On restart, `load_existing_knowledge()` reads that file and the script can continue from where it left off.
This works because the script saves the knowledge base after every page, not at the end of a run. The knowledge base path (`OUTPUT_PATH`) points to a fixed filename, so re-running the script on the same PDF picks up the existing file.
One caveat: `setup_directories()` is documented as clearing previously generated files. The README says it "clears any previously generated files" but the knowledge base resumption is presented as a feature. The exact behavior in the published script version should be verified before relying on resumption for a long run.
Editorial conclusion
AI-reads-books-page-by-page is the right tool for a developer or researcher who wants to extract structured knowledge points from a technical PDF using the OpenAI API, and who is comfortable reading and editing a single Python file to configure it. The critical setup step is placing the PDF in the script's directory and setting `PDF_NAME` before running. Users who want a no-code book summarizer, support for formats other than PDF, or a local model instead of OpenAI will need to modify the script or look for a different tool. The last push to this repository was on 2026-06-27.
Frequently asked questions
Is there an AI that can read a book and extract knowledge?
AI-reads-books-page-by-page is a Python script that processes a PDF book one page at a time using the OpenAI API, extracting knowledge points from each page and saving them to a JSON knowledge base. It also generates Markdown summaries at configurable intervals.
Does AI-reads-books-page-by-page support resuming a partial run?
Yes. The script saves the knowledge base JSON after every page. On restart, it calls `load_existing_knowledge()` to read the saved file and continues from the accumulated knowledge points. Set `TEST_PAGES` to a small number to verify the setup before a full run.
What AI models does AI-reads-books-page-by-page use?
The script uses two configurable constants: `MODEL` for per-page knowledge extraction and `ANALYSIS_MODEL` for generating summaries. Both are set in `read_books.py` and passed to the OpenAI API. The README documents these constants but does not specify default model names.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/echohive42-ai-reads-books-page-by-page)