book-to-skill: Converting Technical Books and Documents into Agent Skills
GitHub describes it as Turn any technical book PDF into a Claude Code skill , ready to study, reference, and use while you work.. The repository metadata lists Python as its primary language. The metadata lists the MIT license. This article stays within the project description and details documented in the GitHub repository README.
At a glance
- What is it?
- book-to-skill is a Python tool that converts any technical book, document folder, or collection of sources into a structured agent skill your coding agent loads on demand. It replaces the Discovery Loop Tax of re-reading PDFs on every query with a one-time extraction that delivers 24x to 51x fewer tokens per answer, measured on real books.
- Who is it for?
- book-to-skill is a well-matched tool for engineers who buy technical books, use them heavily for the first week, and then never find the right chapter again when the question comes up months later. It is the wrong approach for documents that change frequently, because the extracted skill reflects the content at conversion time and must be manually updated when the source changes.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
DEEP OPEN-SOURCE ANALYSIS
The Problem book-to-skill Solves
A PDF-reading agent does not just read; it navigates. On every query it re-fetches the table of contents, follows links, and reprocesses the surrounding pages to find the relevant passage. The README calls this the Discovery Loop Tax: the agent pays the full navigation cost on every turn, regardless of how many times the same chapter has already been processed.
book-to-skill pays that structuring cost once, at conversion. It runs the book through a deterministic Python extractor (document to clean text and metadata) followed by a spec-driven generator that organizes the result into a SKILL.md index, per-chapter files, a glossary, a patterns file, and a cheatsheet. On subsequent queries, the agent loads only the chapter file that covers the topic, which the README measures at 1,000 tokens per chapter versus the full-book Discovery Loop cost.
The README reports 24x to 51x fewer tokens compared to dumping the book into context, with the methodology and per-book tables at docs/performance.md.
What the Conversion Produces
Running book-to-skill on a source file creates a full skill directory under ~/.agents/skills/<slug>/. The output structure:
- SKILL.md: core mental models and a chapter index, around 4,000 tokens - chapters/ch01-*.md and subsequent: one file per chapter, around 1,000 tokens each, loaded on demand - glossary.md: every key term in alphabetical order with chapter references, around 1,500 tokens - patterns.md: all techniques, algorithms, and design patterns, around 2,000 tokens - cheatsheet.md: decision tables and quick-reference rules, around 1,000 tokens
Chapter files are loaded on demand: they do not count against the skill token budget until the user asks about that topic. This keeps the initial load small while preserving full detail when needed.
The skill is deposited in a user-level cross-agent directory that GitHub Copilot CLI, Amp, and Codex discover natively. When run under Claude Code, the converter also attempts a verified symlink at ~/.claude/skills/<slug>/. Hermes Agent uses a different path structure, landing skills in its own category-partitioned skills directory.
Installing and Using book-to-skill
book-to-skill requires Python 3.9 or later. The README describes the install section as a single command, with full details at docs/install.md. The pyproject.toml defines optional dependency groups for different source formats. To install with PDF support, the relevant extras are:
pdf = ["pdf-inspector>=1.15,<2", "pypdf", "pdfminer.six"]A technical extra adds docling for more demanding document parsing. The all extra combines every format group.
The command-line syntax from the README:
/book-to-skill <path|folder|glob> [skill-name]The simplest invocation from the README's How it works section:
/book-to-skill ./my-book.pdfThe slug is derived from the filename by default, or can be specified as a second argument. The README notes that the tool also accepts folders and glob patterns, making it possible to convert an entire docs/ directory in one run. After conversion the README says: ask `/your-book-slug replication` and the agent reads the right chapter and answers from the real content. After a conversion, the converter can also publish the skill to GitHub (private by default) so any host installs it with npx skills add.
Beyond Books: Folders, Wikis, and Research Clusters
The name says book, but the input is any structured prose. The README documents several non-book use cases:
Internal documentation: architecture decision records, runbooks, and onboarding guides. A whole docs/ folder can be folded into one skill. The agent queries it while writing code rather than opening the wiki separately.
Brand and design systems: voice guidelines, tone documents, and component principles. A brand book becomes a queryable skill instead of a 60-page PDF that gets skimmed once and ignored.
Research clusters: a stack of papers combined with personal notes into a unified skill, updated as new material arrives. The README describes an update and fold-in mode for adding new sources to an existing skill.
Specs and standards: RFCs, API contracts, and compliance documents. Teams that reference these regularly but never memorize them can fold them into a skill that the agent consults during implementation.
The common thread is documents that get re-opened often enough that the Discovery Loop Tax accumulates. If a team member opens the same PDF more than twice a week to answer a recurring question, it is a conversion candidate.
Limitations and Cases Where book-to-skill Is the Wrong Tool
The extraction reflects the document at conversion time. A living specification that updates weekly requires re-running the converter to keep the skill current. The skill does not auto-update when the source changes. Teams using book-to-skill for a fast-moving internal wiki should treat skill freshness as an explicit maintenance task.
The tool's effectiveness depends on the source document's structure. A PDF that is a scanned image rather than text requires OCR; the README mentions a scanned PDF use case but the default pdf extra works on text-based PDFs. The docling extra under the technical group handles more complex parsing scenarios.
The on-demand chapter loading relies on the agent host correctly invoking the skill and following the SKILL.md index to select the right chapter file. Hosts that load the entire skill directory at once rather than loading chapters on demand will not benefit from the token-saving design, because they incur the full cost upfront.
book-to-skill creates a private GitHub repository for the skill when the user requests it. The README notes that these repositories are private by default. Teams that want to share a skill across teammates must manage access to that private repository explicitly.
Maintenance and Licence
The project is MIT-licensed, which allows use, modification, and distribution without copyleft restrictions. The last push was on 2026-09-27 and the most recent release is v1.4.0, published on 2026-08-10. Release v1.2.0, from June 2026, added installable package support and multilingual chapter detection, according to the release title in the repository metadata.
The pyproject.toml shows the project uses Hatchling as its build backend and Ruff for linting, targeting Python 3.9 as the minimum version. The tests directory is set to tests/ in the pytest configuration. The CHANGELOG.md file in the repository tracks changes per release.
The SKILL.md file in the repository root means book-to-skill installs its own skill for agents working in the repository, covering the codebase documentation and development workflow.
Editorial conclusion
book-to-skill is a well-matched tool for engineers who buy technical books, use them heavily for the first week, and then never find the right chapter again when the question comes up months later. It is the wrong approach for documents that change frequently, because the extracted skill reflects the content at conversion time and must be manually updated when the source changes. The core extraction works on any structured prose; the 24x to 51x token savings are measured on real books and documented at docs/performance.md. Verify Python 3.9 compatibility and install the correct optional extras for your document types before running a conversion. The project is MIT-licensed.
Frequently asked questions
How to turn a book into a Claude skill?
Install book-to-skill with pip install book-to-skill[pdf] for PDF support, then run book-to-skill ./your-book.pdf from the command line. The tool creates a skill directory under ~/.agents/skills/<slug>/ and, when run under Claude Code, also creates a symlink at ~/.claude/skills/<slug>/ so Claude Code discovers it automatically.
What file types does book-to-skill support?
The pyproject.toml defines optional extras for PDF (pdf-inspector, pypdf, pdfminer.six), EPUB (ebooklib, beautifulsoup4), DOCX (python-docx), HTML (trafilatura), RTF (striprtf), and technical documents via docling. Install the relevant extra alongside the package, or use book-to-skill[all] to install support for all formats.
How does the on-demand chapter loading reduce token usage?
book-to-skill splits the book into per-chapter files of about 1,000 tokens each. Only the chapter file covering the queried topic is loaded when the agent answers a question. The README reports this reduces token cost by 24x to 51x compared to loading the full book text into context on every query, with methodology documented in docs/performance.md.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/virgiliojr94-book-to-skill)
Community notes