wdoc: sourced RAG answers across PDFs, audio, Anki and more
Summarize and query from a lot of heterogeneous documents. Any LLM provider, any filetype, advanced RAG, advanced summaries, scriptable, etc
At a glance
- What is it?
- wdoc is a Python RAG tool and library that summarizes and queries heterogeneous document collections through any LLM provider. Its pitch is a sourced, indented markdown answer rather than a fluent guess, and its cost is a large dependency surface and a CLI you have to learn.
- Who is it for?
- Adopt wdoc if you already have a pile of mixed-format material (lecture recordings, PDFs, EPUBs, flashcard decks) and you want one query to cite the exact source passage instead of paraphrasing it. Skip it if you need a small dependency footprint, a stable Python API you can pin without reading the changelog, or permissively licensed code you intend to embed in a closed product, since the licence is AGPL-3.0.
- Can I use it commercially?
- Yes, with strict conditions. AGPL-3.0 is a network copyleft licence: if people use a modified version over a network, for example as a hosted service, you must offer them its source code under the same licence.
- Is it still maintained?
- Yes. The repository last received commits 38 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem wdoc was written to solve
The README is unusually direct about its origin: the author is a psychiatry resident who needed a definitive answer pulled from several sources at once, listing audio recordings, video lectures, Anki flashcards, PDFs and EPUBs. That is a specific shape of problem. Most retrieval tools assume a corpus of text files or a web crawl. A medical student's corpus is a folder where the same lecture exists as a recording, a slide deck and a set of flashcards, and none of them is the canonical version.
So the target user is not a team building a customer support bot. It is one person with a heterogeneous pile who wants to ask a question once and get an answer that points back to the passage it came from. The README frames the goal as "perfectly useful" summaries and "perfectly useful" sourced answers, and claims the system handles tens of thousands of documents at once. The word doing the work there is sourced. A summary that reads well but cannot be traced is exactly what the author says he was frustrated by in existing RAG solutions.
The project is also explicitly both a tool and a library, and has been wrapped as an Open-WebUI tool, so the intended audience includes people who want to script it rather than only run a command.
How a wdoc query actually flows
The repository ships three diagrams, one per task, and they name the stages. A query goes through Raphael the Rephraser, then the VectorStore, then Eve the Evaluator, then Anna the Answerer, and finally a recursive combining step before output. The summary task is shorter: loading and chunking, then Sam the Summarizer, then concatenation into a wdocSummary. The search task stops after Eve the Evaluator.
That split matters. Query and search share the front half (rephrase, retrieve, evaluate) and diverge at the point where an answer must be written. The evaluation stage is where the design gets opinionated: the README describes using both an expensive and a cheap LLM so that recall stays high, on the reasoning that fetching a lot of documents per query is affordable when retrieval is embedding-based. In other words, the expensive model is not spent on every candidate document, it is spent where judgement is needed.
Answers are aggregated gradually in semantic batches rather than in one pass, which is what the recursive combining step in the diagram refers to. The output format is indented markdown with source pointers to the exact portion of the source document. The backend stack is LangChain plus LiteLLM, which is the mechanism behind the "any LLM provider" claim: provider support is inherited from LiteLLM rather than hand-written per vendor. ARCHITECTURE.md exists at the repository root for readers who want more than the diagrams give.
Installing wdoc and running a first query
The README gives a short answer for installation: when in doubt, use the full extra. Plain wdoc ships only PDF and URL or web loaders, and everything else (YouTube, audio, Anki, office formats, Logseq) lives in optional extras. The full extra bundles all of them.
uvx wdoc[full]If you install the plain package instead, expect loader errors on anything that is not a PDF or a URL, and reinstall with the extra rather than debugging the parser. Note that setup.py runs a post-install step that calls playwright install, and refreshes yt-dlp to a pre-release only if yt-dlp is already importable, which is how the YouTube extra stays optional while still tracking extractor fixes. That step is documented in setup.py, not in the README, so it is easy to miss.
The README's quick start begins by querying a PDF, and the example link it uses is the situational-awareness.ai PDF. The README does not print a complete query command in the section reproduced here; it defers flag-level detail to the wdoc-skill/ directory, which contains SKILL.md, REFERENCE.md and EXAMPLES.md. REFERENCE.md holds the full argument, environment variable, filetype and Python API tables, and EXAMPLES.md holds copy-pasteable shell and Python recipes. Start there rather than guessing at flags.
What you should see, once a query runs, is an indented markdown answer with source pointers to the exact portion of the source document, not a wall of prose. If you get a provider error instead, the failure is almost certainly in LiteLLM configuration, since that is the layer that resolves model names.
For a browser-based path, the repository has a docker/ directory with a Gradio web interface, described in its own README as an experimental deployment. That is the option for people who do not want to touch the CLI at all.
Where wdoc is the wrong tool
The dependency surface is the first real cost. A full install pulls in the PDF pipeline, the audio and video loaders, the office format parsers, Playwright browsers and yt-dlp, and setup.py performs work after the package install finishes. In an air-gapped environment or a locked-down container, that post-install step is a failure point, and the setup script itself catches the playwright error and prints it rather than aborting, so a broken browser install can pass silently and only surface later when a URL loader is used.
The second constraint is scope. This is built for one person interrogating a corpus, not for a service with latency budgets. The README says the system deliberately fetches a lot of documents per query and aggregates answers in batches, which is a design that trades time and tokens for recall. If your requirement is a fast single-document lookup, the rephraser and evaluator stages are overhead you are paying for nothing.
The third is the API surface. The README states the project is still under active development with tens of planned features, and that the main branch is more stable than the dev branch. Anyone pinning wdoc as a library dependency should read the release notes between versions rather than assuming argument stability. The README also asks users to open an issue before making a pull request, which tells you the internals are still moving.
Finally, the licence is AGPL-3.0. That is a network-copyleft licence, and it is a genuine consideration for anyone embedding wdoc in a hosted product. This is not legal advice; read the licence text and talk to counsel if the distinction matters to you.
How wdoc differs from a plain LangChain retrieval chain
The closest comparison is not another application, it is the raw stack wdoc is built on. LangChain gives you document loaders, a vector store and a chain; assembling a working heterogeneous RAG pipeline from those pieces is the job wdoc has already done. The difference in approach is the evaluator stage. A standard retrieval chain embeds the question, takes the top-k chunks and hands them to one model. wdoc inserts an evaluation step between retrieval and answering, and splits model spend between a cheap and an expensive model so that more candidates can be considered without every candidate costing full price.
The second difference is the summary task. Most RAG tooling treats summarization as a prompt you write yourself. wdoc ships it as a first-class task with its own pipeline and output type, and the README describes the aim as getting the thought process of the author rather than generic takeaways. Whether that lands depends on your documents, but it is a stated design position, not a default.
The third is filetype breadth. Fifteen or more loaders are implemented according to the README, spanning audio, video, Anki and office formats. A hand-rolled LangChain pipeline typically covers text and PDF and stops there. If your corpus is text and PDF only, the breadth is dead weight and a smaller stack will be easier to operate.
Maintenance, releases and what upgrading costs
The last push to the default branch was on 2026-08-24, and release 5.2.2 is dated the same day. Before that, 5.2.0 landed on 2026-07-03 and 5.1.3 on 2026-06-23. That is a steady release rhythm across the summer, and the repository is not archived, so describing it as under active development matches the record rather than a claim from the README.
The upgrade cost is concentrated in two places. Model defaults move: the repository contains bump_default_models.sh and bumpver.toml at the root, which means default model identifiers are refreshed by script as providers deprecate old ones. If you rely on a default rather than pinning a model explicitly, an upgrade can change which model answers your queries. The second is yt-dlp, which setup.py deliberately upgrades to a pre-release when the YouTube extra is present, so that dependency is intentionally not pinned to stable. On a shared machine that is a moving part.
For the licence, the practical question is distribution. AGPL-3.0 applies to the code in this repository. If you run wdoc locally for your own research, the licence question is mostly academic. If you expose a modified wdoc as a network service, the obligations are different, and the README offers no guidance on that because it is not the project's job to give it.
Editorial conclusion
Adopt wdoc if you already have a pile of mixed-format material (lecture recordings, PDFs, EPUBs, flashcard decks) and you want one query to cite the exact source passage instead of paraphrasing it. Skip it if you need a small dependency footprint, a stable Python API you can pin without reading the changelog, or permissively licensed code you intend to embed in a closed product, since the licence is AGPL-3.0. Before committing, install with uvx wdoc[full], run one PDF query end to end, and check that your chosen provider is reachable through LiteLLM, because the README lists no per-provider verification steps. The argument tables in wdoc-skill/REFERENCE.md are the place to confirm flag names before you script anything.
Frequently asked questions
What file types can wdoc read?
The README says fifteen or more are already implemented, covering PDFs, audio, video and YouTube, Anki flashcards, EPUBs, office formats and Logseq, among others. Only PDF and URL or web loaders ship with the plain package; the rest require optional extras, and the full extra bundles all of them.
Does wdoc work with local LLM providers or only cloud APIs?
The README states that wdoc supports virtually any LLM provider, including local ones, and mentions extra layers of security for sensitive material. Provider resolution goes through LiteLLM, so the practical answer is that any provider LiteLLM can address is available to wdoc.
Why does wdoc need the [full] extra instead of a plain install?
The plain wdoc package ships only PDF and URL or web loaders. YouTube, audio, Anki, office formats and Logseq loaders live in optional extras, and the README recommends uvx wdoc[full] when in doubt so that no loader dependency is missing at runtime.
Is there a web interface for wdoc, or is it command line only?
There is an experimental Docker web UI built on Gradio, documented in the docker/ directory of the repository. The README describes it as a way to process documents without CLI interaction, and also notes that wdoc has been turned into an Open-WebUI tool.
What licence does wdoc use?
The repository is licensed under AGPL-3.0, as listed in the LICENSE file at the root and in the project metadata. That is a network-copyleft licence, which matters if you plan to expose a modified wdoc as a service rather than run it locally.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/thiswillbeyourgithub-wdoc)