# thepipe: pulling markdown out of documents with vision-language models

> A Python package that converts PDFs, Word documents, slides, notebooks, and media into clean markdown chunks, with chunking strategies and an OpenAI compatible client as first class arguments.

**emcf/thepipe** — Get clean data from tricky documents, powered by vision-language models ⚡

- Repository: https://github.com/emcf/thepipe
- Website: https://thepi.pe
- Stars: 1,526 · Forks: 98
- Language: Python
- License: MIT
- Published: 2026-10-06 · Updated: 2026-10-06 · Language: en
- Canonical page: https://hysenlabs.com/projects/emcf-thepipe

## What sits underneath the scrape call

thepipe is a wrapper, and it is more honest than most wrappers about which parts do the work. The public entry point is `scrape_file`, imported from `thepipe.scraper`, and the arguments that matter are a filepath, an optional `openai_client`, an optional `model`, and an optional `chunking_method`. Everything else is dependency wiring. The package requires Python 3.9 or newer and publishes under the name `thepipe-api` while importing as `thepipe`, which is the first small trap when you are reading installation instructions.

The pinned dependency list in `requirements.txt` explains the architecture better than the README does:

```toml
pymupdf4llm==0.0.21
PyMuPDF==1.25.5
magika>=0.5.0
markdownify==0.12.1
openai>=1.51.0
```

Two of those exact pins are the text extraction layer, one is Google's content type detector doing the AI-native file type sniffing the README advertises, one turns HTML into the markdown the package promises, and `openai` is there because every model call goes through the OpenAI client interface rather than a vendor SDK. Add `python-docx`, `python-pptx`, and `openpyxl` to that and the picture is a format-by-format set of readers with a single markdown output, and `moviepy` for the video path. There is no release history on the project at all, so the version in `setup.py`, 1.7.3, is the only version number a reader can point at.

## Four optional extras instead of one heavy install

The default install is deliberately CPU friendly so it can run in CI. Heavier libraries are opt in, and `setup.py` shows exactly what each extra pulls:

```python
EXTRAS = {
    "audio": ["openai-whisper>=20231117"],
    "semantic": ["sentence-transformers>=2.2.2"],
    "llama-index": ["llama-index>=0.10.50,<0.11"],
    "gpu": [
        "torch>=2.5,<2.6",
```

That `torch>=2.5,<2.6` ceiling is worth noticing. The GPU extra is pinned to a narrow version window, so when PyTorch moves on, this extra stops resolving cleanly rather than silently upgrading under you. It is the kind of decision that keeps a package stable for its users and turns into an upgrade chore for its maintainer.

For a CPU only machine that still wants embeddings, the README gives a two step install that installs the CPU wheels from the PyTorch index before adding the extra:

```bash
pip install torch==2.5.1+cpu torchvision==0.20.1+cpu torchaudio==2.5.1+cpu \
  --index-url https://download.pytorch.org/whl/cpu
pip install thepipe-api[semantic]
```

Media rich work also needs a system package, `ffmpeg`, installed outside pip. That is an easy thing to miss because the failure looks like a bad path rather than a missing binary. There is also a console entry point registered in `setup.py`, `thepipe=thepipe.__init__:main`, so the package ships a command line surface as well as a library one.

## Scraping a file takes four lines and optionally a client

The minimal call needs no API key and no model, which is the fastest way to see what the non vision path produces:

```python
from thepipe.scraper import scrape_file

# scrape text and page images from a PDF
chunks = scrape_file(filepath="paper.pdf")
```

Passing a client switches on the vision path:

```python
chunks = scrape_file(
  filepath="paper.pdf",
  openai_client=client,
  model="gpt-4o"
)
```

The default model is `gpt-4o`, and because the client is just the OpenAI interface, pointing at another provider is a matter of base url. The README names `https://openrouter.ai/api/v1` for OpenRouter and `http://localhost:3000/v1` for a local server, and notes that a non OpenAI provider still needs its key passed into the client. That single decision is what keeps the package from locking you into one vendor, and it is also why the model name is a parameter on the scrape call instead of a global setting. The project itself is MIT licensed, sits at 1,526 stars and 98 forks, and lists document AI, scraping, scanned PDF, and vision language model as topics, which is an accurate summary of a package whose centre of gravity is turning messy documents into text a model can read.

## Seven chunking methods decide what the model ever sees

Chunking is where document pipelines usually get their quality, and thepipe treats it as a first class choice rather than an afterthought. The README lists seven methods. `chunk_by_document` returns one chunk for the whole file. `chunk_by_page` returns one chunk per page, which means per PDF page or per PowerPoint slide. `chunk_by_length` splits on size, and `chunk_by_section` splits on markdown headings. `chunk_by_keyword` breaks at keywords you nominate. The remaining two are marked experimental: `chunk_semantic` splits on spikes in semantic change using a configurable threshold and needs `sentence-transformers`, and `chunk_agentic` delegates the split to an LLM that hunts for semantically meaningful sections and needs the OpenAI client.

The useful detail is that chunking can happen after the scrape, not only during it:

```python
chunks = scrape_file(
  filepath="paper.pdf",
  chunking_method=chunk_by_document
)
```

That means you can scrape once and re split repeatedly while you work out what a retrieval step actually wants, instead of paying model cost again for each experiment. The trade off between the seven is straightforward. Page and section splits respect the document's own structure and are free, length splits are blunt but predictable, keyword splits need you to know the terms, and the two experimental ones cost either local compute or model calls in exchange for boundaries that line up with meaning. For a first run, `chunk_by_page` is the honest baseline because a page boundary is never wrong, only sometimes coarse.

## Handing chunks to an OpenAI chat call

The last integration piece is a helper in `thepipe.core` that converts chunks into chat message content, so you do not have to write the joining logic yourself:

```python
messages += chunks_to_messages(chunks)

response = client.chat.completions.create(
    model="gpt-4o",
    messages=messages,
)
```

`chunks_to_messages` also takes an optional `text_only` parameter, which is the switch you want when you are paying for vision tokens you do not need, for example when re-running a scrape against a document that already produced usable text. The overall flow the README documents is scrape, chunk, convert to messages, call the model, and every step is a plain function you can substitute. There is no agent framework in the middle and no orchestration object with a lifecycle, which makes the library easy to embed in something like a Flask endpoint or a notebook and easy to leave behind.

The repository itself is small. The tree is `.github/`, `.gitignore`, `LICENSE`, `README.md`, `requirements.txt`, `setup.py`, `tests/`, and `thepipe/`. There is a CI workflow and a codecov badge in the README header, so there is a test suite, though the repository does not publish coverage numbers and the test files are not documented individually. Reading `thepipe/chunker.py` and `thepipe/scraper.py` is the fastest way to understand the real behaviour, and the presence of a `tests/` directory means edge cases have at least been thought about even if the docs do not enumerate them.

## Where the README stops and the questions begin

Two things are conspicuously unstated. First, accuracy. Nothing in the documentation says how the vision path compares against a good text layer on the same file, and there is no guidance on what to do when a document has both a text layer and page images that disagree. If you have a corpus with ground truth, that comparison is worth running on a small sample before you trust the output at volume.

Second, cost control. The chunker choices bound how many tokens reach the model, and the extraction step itself calls a model once per page for image and layout work, which is the part that actually moves the bill. The README does not give a way to cache extraction results, so if the same corpus is scraped more than once, doing it once and keeping the chunks on disk is the obvious move.

What the docs do settle is the setup, and they settle it well. Installation is one command, the extras are named for what they enable, the CPU only path is spelled out with exact wheel versions, the chunking menu is complete with its dependencies labelled, and the vendor question is answered by the base url on a standard client. What is not settled is where extraction quality plateaus, which for most people is the only question that matters. The last commit on the repository is dated 2026-09-02, so the package is being worked on, and with 16 open issues it is active rather than settled.

## Conclusion

thepipe is worth a look if your input is documents whose layout matters, because its answer to a scanned page is to look at the page rather than guess at a text layer. The shape of the library is small and clear: `scrape_file` in, chunks out, a chunking method chosen at call time, and an OpenAI client passed in when you want vision or the agentic chunker. Its weak spot is versioning, since `setup.py` still reports 1.7.3 with `PyMuPDF` pinned at 1.25.5 and `pymupdf4llm` at 0.0.21, both of which are exact pins you will eventually have to relax. Start with `pip install thepipe-api` and one `scrape_file` call on a file that has no usable text layer, then read `thepipe/chunker.py` to see how the split boundaries are actually chosen.

## FAQ

### What kinds of files can thepipe read?

The README lists PDFs, Word documents, PowerPoint decks, Python notebooks, videos, and audio, plus whatever magika detects as a document. Video and audio paths need ffmpeg installed on the system, and the audio extra adds openai-whisper for local transcription.

### Do I need an OpenAI API key to use thepipe?

Not for the basic path. A plain scrape_file call on a digital PDF needs no client at all. The vision path and the agentic chunker take an OpenAI client, and because it is the standard interface you can point its base url at OpenRouter or a local server such as http://localhost:3000/v1 and pass that provider's key instead.

### Which chunking method should I use first?

chunk_by_page, because a page or slide boundary is always a real boundary even when it is coarse. Move to chunk_by_section if your documents have reliable markdown headings, and to the experimental chunk_semantic or chunk_agentic only when you need split points that follow meaning rather than layout.

## Sources

- [emcf/thepipe on GitHub](https://github.com/emcf/thepipe)
- [Issues](https://github.com/emcf/thepipe/issues)
- [License: MIT](https://github.com/emcf/thepipe/blob/main/LICENSE)
- [Project website](https://thepi.pe)
- [README](https://github.com/emcf/thepipe/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/emcf-thepipe
