# spaCy Layout: turning PDFs and Word files into spaCy Doc objects with Docling

> spacy-layout wraps Docling's document conversion inside a spaCy pipeline, so a PDF becomes a Doc with layout spans, tables as DataFrames and a Markdown rendering. It is a thin integration layer, and the thinness is both the appeal and the constraint.

**explosion/spacy-layout** — 📚 Process PDFs, Word documents and more with spaCy

- Repository: https://github.com/explosion/spacy-layout
- Stars: 912 · Forks: 64
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/explosion-spacy-layout

## The gap spacy-layout fills between a PDF and a spaCy pipeline

spaCy has no native concept of a page, a heading or a table cell. A Doc is a sequence of tokens with annotations on top, and everything downstream (matchers, NER, text classification) assumes that sequence. A PDF is the opposite: a fixed-layout canvas where reading order has to be inferred. Anyone who has tried to feed a PDF into a spaCy pipeline ends up writing their own converter, and that converter is where the project quietly stalls.

spacy-layout targets that seam. It is a plugin that integrates with Docling, the document conversion library from the same organisation, and it produces a spaCy Doc rather than a bespoke dictionary. The README frames the output as "clean, structured data" and names two audiences: teams applying NLP techniques such as named entity recognition or text classification to documents, and teams building chunking for RAG pipelines. The second audience is the one the topics list leans into, with rag and generative-ai both tagged on the repository.

The scope is deliberately narrow. The package does not parse PDFs itself. It is a preprocessor plus a set of extension attributes, and the requirements file confirms the split: spacy>=3.7.5, docling>=2.5.2, pandas and srsly, with pandas and srsly version ranges deferred to Docling and spaCy respectively. If you want a document converter with no NLP framework attached, this is the wrong layer to adopt.

## How a PDF becomes a Doc: Docling conversion, layout spans and extension attributes

The mechanism is a two-stage pipeline. Docling converts the source file into a structured intermediate representation, and spaCyLayout maps that representation onto a spaCy Doc. The README describes the result as layout spans that "map into the original raw text and expose various attributes, including the content type and layout features." Those spans live in a SpanGroup under doc.spans["layout"].

The span labels are the document structure: "text", "title", "section_header" and, for tables, "table". Each span carries a layout extension attribute with features including a bounding box, and a heading attribute that the README explicitly qualifies: the closest heading to the span, with the note that accuracy "depends on document structure." That caveat is worth taking at face value. Heading association is a heuristic over inferred structure, not a guaranteed hierarchy.

Four document-level attributes carry the rest. Doc._.layout holds document layout features, Doc._.pages returns a list of page tuples pairing a PageLayout with the spans on that page, Doc._.tables collects every table span, and Doc._.markdown returns a Markdown rendering of the whole document. The Markdown attribute is the cheapest way to get text out for an LLM context window, and it is produced without you writing a serializer.

Tables are where the design gets interesting. Each table span exposes a data attribute holding a pandas.DataFrame of the extracted contents. By default the span text is the literal placeholder TABLE, which means the table's figures are absent from doc.text. You can change that by passing a display_table callback that receives the DataFrame and returns a string. The README gives the example of rendering the column names into the text so that a trained NER or text classifier can see them. This is a real decision point: leave the default and your token-level models are blind to table contents, or write a callback and decide how much of the table to flatten into a linear token stream.

## Installing spacy-layout and processing a first document

The README states that the package requires Python 3.10 or above, and installation is a single pip command.

```bash
pip install spacy-layout
```

After that, you initialize spaCyLayout with an nlp object, which the README says is used for tokenization, and call it on a path. A blank English pipeline is enough for the layout work itself.

```python
import spacy
from spacy_layout import spaCyLayout

nlp = spacy.blank("en")
layout = spaCyLayout(nlp)

doc = layout("./starcraft.pdf")
print(doc._.markdown)
```

What you should see is the Markdown rendering of the document, plus, if you print doc._.layout and doc._.tables, the document layout features and the list of table spans. To inspect structure rather than text, iterate the layout span group and read the label and offsets.

```python
for span in doc.spans["layout"]:
    print(span.text, span.start_char, span.end_char)
    print(span.label_, span._.layout)
```

Each iteration prints the span text, its character offsets into doc.text, its label such as "section_header" or "table", and its layout features including the bounding box. The step that makes the package worth its dependency weight is applying a real pipeline to the already-created Doc. spaCy lets you call nlp on a Doc, so linguistic annotations and named entities land on the same object the layout spans describe.

```python
nlp = spacy.load("en_core_web_trf")
layout = spaCyLayout(nlp)
doc = nlp(layout("./starcraft.pdf"))
```

The README notes that en_core_web_trf installs separately with python -m spacy download en_core_web_trf. For batches, spaCyLayout.pipe takes an iterable of paths or bytes and yields Doc objects, which is the form you want when the conversion cost matters.

## Serialization, the extension attribute trap, and what it costs at scale

Document conversion is the expensive part of this workflow, and the README is direct about the remedy: serialize the resulting Doc objects in spaCy's binary format with DocBin so you do not re-run the conversion. The example passes store_user_data=True, which is what preserves the custom extension data.

There is a catch, and the README flags it as a known rough edge. The extension attributes such as Doc._.layout are registered when spaCyLayout is initialized. Loading a DocBin back without initializing the class first leaves those attributes unpopulated. The README's diff shows the workaround: construct a spaCyLayout instance before reading the file.

```python
layout = spaCyLayout(nlp)
doc_bin = DocBin(store_user_data=True).from_disk("./file.spacy")
docs = list(doc_bin.get_docs(nlp.vocab))
```

The README states that a more elegant approach is planned for an upcoming version, so treat this as a current constraint rather than a permanent design. The practical consequence is that any code path which loads serialized documents needs a spaCyLayout instance in scope, even if it never converts a file. That is an awkward dependency for a serving process whose only job is to read pre-processed docs.

The second cost is the dependency chain. Docling is a large library, and pandas comes with it. Pulling spacy-layout into a service means pulling in a document conversion stack whether or not that service ever converts anything. If your architecture separates conversion from inference, the conversion workers need the full stack and the inference workers need spaCy plus an initialized spaCyLayout. Splitting those two roles is the cleanest way to keep the serving path small.

## Where spacy-layout is the wrong choice

The package inherits its failure modes from Docling, and the README does not document rollback, fallback or error handling for documents that convert badly. There is no described mechanism for detecting a failed extraction, no confidence score on the layout spans, and no documented behaviour for a scanned PDF with no text layer. The related searches include "spacy layout ocr", but the README describes no OCR step, so if your corpus is image-only scans, you need to confirm Docling's OCR behaviour before assuming this package handles it.

The heading attribute is the other soft spot. The README itself says accuracy "depends on document structure." For a well-structured report with consistent section headers, that is fine. For a two-column academic paper, a slide deck or a form with floating labels, the inferred structure is the part most likely to be wrong, and there is no validation hook described for catching it. If your downstream task depends on knowing which section a sentence belongs to, you are trusting a heuristic.

Finally, consider the audience mismatch. If you want to extract tables from PDFs and write them to CSV, spaCy is overhead: you are constructing a Doc, a vocabulary and a token sequence to reach a DataFrame you could have obtained from the converter directly. The package earns its place when the Doc is the point, meaning you intend to run spaCy components over the text. If the Doc is just a transport container, drop a layer.

## How it compares with running Docling directly or using a text extractor

The most honest alternative is Docling on its own. Docling performs the conversion and produces the structured representation; spacy-layout is the adapter that turns it into a Doc. Choosing between them is a question of what you need downstream. If you are building a RAG ingestion job that chunks text and writes embeddings, Docling's own output plus the Markdown rendering may be sufficient, and you avoid the spaCy dependency entirely. If you need tokenization, linguistic annotations, rule-based matching or a trained NER model over the document text, the Doc object is the interface those tools expect, and writing the adapter yourself is the work this package saves.

The second alternative is a plain text extractor, the kind that returns a string per page. That approach is simpler and has fewer dependencies, but it discards exactly the information this package exists to preserve: page boundaries, section headers, bounding boxes and table structure. Once a document is a flat string, recovering which sentence was a heading requires heuristics you would then have to write and maintain. spacy-layout's value proposition is that the structure survives into the object your NLP code already consumes.

Worth noting for scope: the README positions the workflow around PDFs and Word documents, and the topics list adds docx explicitly. The README does not enumerate every format Docling supports, so if your corpus is something else, check Docling's own documentation rather than assuming coverage from this page.

## Maintenance, licence and upgrade exposure

The repository is not archived. The last push was on 2026-03-27, which is roughly six months before today, so calling it actively developed would be a stretch; the release history shows v0.0.12 in March 2025, with v0.0.11 and v0.0.10 in December 2024. The version numbers are still in the 0.0.x range, which tells you the API is not yet frozen. The README's own note about the serialization workaround being replaced "in an upcoming version" is a concrete signal that the extension attribute registration behaviour will change.

That combination matters for upgrade planning. A 0.0.x package with a documented API change on the roadmap means you should pin the version in your requirements and test before bumping, particularly if you serialize Doc objects to disk, because a change in how extension attributes are registered could affect how previously written DocBin files load. The requirements.txt is also loose in places: spacy>=3.7.5 and docling>=2.5.2 are lower bounds, while pandas and srsly have their ranges set by those two dependencies. A Docling upgrade can therefore move your pandas version without spacy-layout changing at all.

The licence is MIT, which is permissive and imposes no copyleft obligation on your own code. That applies to spacy-layout itself; Docling, spaCy and pandas carry their own licences, and the README does not restate them. If you are shipping a product, check each dependency's licence separately rather than assuming the MIT label on this repository covers the stack.

## Conclusion

Adopt spacy-layout if you already run spaCy and want document layout, tables and Markdown in the same Doc object your existing components consume. Do not adopt it if you need a standalone converter with no spaCy dependency, or if you cannot run Python 3.10 or above. Before committing, check that Docling handles your specific PDFs, since the README attributes extraction quality to Docling rather than to this package, and confirm that your serialization path re-initializes spaCyLayout before loading a DocBin, because the extension attributes are only registered at initialization.

## FAQ

### What is spaCy Layout used for?

It processes PDFs, Word documents and other formats through Docling and returns spaCy Doc objects with layout spans, tables as pandas DataFrames and a Markdown rendering. The README names linguistic analysis, named entity recognition, text classification and chunking for RAG pipelines as the intended uses.

### What is spaCy Layout in Python?

It is a Python package that acts as a preprocessor for a spaCy pipeline, initialized with an nlp object and called on a document path or bytes. It requires Python 3.10 or above and installs with pip install spacy-layout.

### Which is better, NLTK or spaCy?

The README does not compare the two and says nothing about NLTK. It only describes spaCy Layout's role inside a spaCy pipeline, where the output is a Doc object that spaCy components can annotate.

### What is spaCy used for?

The README points to spaCy for linguistic analysis, named entity recognition and rule-based matching, and shows loading the transformer-based English pipeline en_core_web_trf to annotate a Doc with POS tags, dependencies and entities. spacy-layout feeds document text into that pipeline.

### Is spaCy still relevant?

The README does not address this. The repository facts show spacy-layout itself is not archived, with a last push on 2026-03-27 and v0.0.12 released in March 2025.

## Sources

- [explosion/spacy-layout on GitHub](https://github.com/explosion/spacy-layout)
- [Issues](https://github.com/explosion/spacy-layout/issues)
- [License: MIT](https://github.com/explosion/spacy-layout/blob/main/LICENSE)
- [README](https://github.com/explosion/spacy-layout/blob/main/README.md)
- [Releases](https://github.com/explosion/spacy-layout/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/explosion-spacy-layout
