# PdfPig: reading PDF text in C# without a commercial licence

> PdfPig is an Apache-2.0 C# library for extracting text and content from PDFs, with layout analysis tools that reorder words into readable text. It is aimed at .NET developers who need PDF extraction inside their own code.

**UglyToad/PdfPig** — Read and extract text and other content from PDFs in C# (port of PDFBox)

- Repository: https://github.com/UglyToad/PdfPig
- Website: https://github.com/UglyToad/PdfPig/wiki
- Stars: 2,571 · Forks: 334
- Language: C#
- License: Apache-2.0
- Published: 2026-09-28 · Updated: 2026-09-28 · Language: en
- Canonical page: https://hysenlabs.com/projects/uglytoad-pdfpig

## What PdfPig extracts, and who needs that

PdfPig reads text and other content from PDF files and also supports basic PDF file creation. The repository describes it as a port of PDFBox, so the mental model is the same as the Java library: parse the document structure, expose pages, letters, words and images, and let the caller assemble meaning. The audience is .NET developers who already have PDFs arriving from somewhere (uploads, a scanner pipeline, an archive) and need the text inside them as data. The topics listed for the repository include alto-xml, hocr, page-xml, layout-analysis and document-analysis, which signals that the project cares about the intermediate representation, not just a flat string. That matters because a flat string is the easy part and the wrong part: two columns, a header, a footer and a table read as one stream of words will defeat most downstream parsing. PdfPig gives you the pieces to do better, and the README is honest that you often have to.

## How the extraction pipeline works: letters, words, blocks, reading order

The mechanism visible in the README is a chain. PdfDocument.Open gives you a document; document.GetPages() enumerates Page objects; each page exposes Letters. From letters you run a word extractor, in the examples NearestNeighbourWordExtractor.Instance, which groups letters into Word objects. Words go into a page segmenter, DocstrumBoundingBoxes.Instance, which produces text blocks. Blocks go into a reading order detector, UnsupervisedReadingOrderDetector.Instance, which returns ordered blocks. The advanced example wires exactly those four stages and then draws each block's bounding box and reading order number onto a copy of the PDF with PdfDocumentBuilder, which is a neat way to see whether the segmentation matches what a human sees. The README is explicit that you should not use page.Text directly unless you know what you are doing, because the Text property preserves the internal content order, which is rarely the order you want. The higher-level convenience is ContentOrderTextExtractor.GetText(page). So there are three levels of effort: raw letters, a single call for plain text, and a hand-assembled pipeline when the single call is not good enough.

## Installing PdfPig from NuGet and extracting your first page

The package is published on NuGet as PdfPig. The README gives the package manager console command, and links the NuGet page and the releases tab as the other places to get it. Run this in the Visual Studio package manager console for the project that needs it.

```bash
Install-Package PdfPig
```

After installing, the smallest useful program opens a file, iterates pages and pulls text. Note the using statements in the README example: the text extractor and word extractor live in the DocumentLayoutAnalysis namespaces, not the root namespace, which is a common first stumble. The README's own snippet looks like this, with the two using directives shown as comments.

```cs
// using UglyToad.PdfPig.DocumentLayoutAnalysis.TextExtractor;
// using UglyToad.PdfPig.DocumentLayoutAnalysis.WordExtractor;
using (PdfDocument document = PdfDocument.Open(@"C:\Documents\document.pdf"))
{
    foreach (Page page in document.GetPages())
    {
        string text = ContentOrderTextExtractor.GetText(page);
        IEnumerable<Word> words = page.GetWords(NearestNeighbourWordExtractor.Instance);
    }
}
```

What you should see is one string per page, in the order the layout analysis decides, plus the word objects with their positions. If the string comes out scrambled, that is the signal to move to the advanced pipeline described above rather than to keep tuning the simple call.

## Creating a PDF, and where creation stops

PdfPig also writes PDFs through PdfDocumentBuilder. The README's example registers a Standard 14 font with AddStandard14Font(Standard14Font.Helvetica) before use, adds an A4 page, places text at a point with a size, and calls builder.Build() to get a byte array. Fonts must be registered with the builder before use so pages can share font resources, and only Standard 14 fonts and TrueType (.ttf) files are supported. The limitation paragraph is worth reading twice. Document creation supports very limited changes to existing PDFs and does not support editing forms, copying or changing annotations, metadata or document structure data, or adding or removing text with existing fonts. In practice that means PdfPig is a reader with a small writer attached, not a PDF editor. If your task is filling an AcroForm or stamping an existing document's metadata, the README says the library will not do it, and no amount of layout analysis changes that.

## Encrypted files, passwords and the ParsingOptions object

Encrypted documents can be opened. The README shows two forms: a single owner or user password supplied through ParsingOptions with the Password property, and a list of candidate passwords through a Passwords list, which PdfPig will try. The single-password form appears inline in the README without a fenced block, and the list form is given as a code sample. This is a small feature but it decides architecture: if your documents arrive password protected, you need to thread ParsingOptions through whatever opens them, and you should decide up front what happens when every candidate fails. The README does not document the exception behaviour for a wrong password or a document that cannot be decrypted, so that is something to establish from the source or from a test against your own files before you build error handling around it.

## Where PdfPig is the wrong tool

Three cases stand out. First, form and annotation work: the README lists editing forms and copying or changing annotations and metadata as unsupported, so a workflow built on filling government PDFs or extracting embedded metadata is out of scope. Second, editing existing text: adding or removing text with existing fonts is unsupported, so redaction or template rewriting is not a PdfPig job. Third, anything where you need a guaranteed byte-identical round trip of a document's structure. The debug overlay example is instructive here: it does not modify the source PDF, it builds a new one from the source page and draws on it. That is the shape of the writer. There is also a subtler cost. The README points out that page.Text preserves internal content order, and the recommended path is ContentOrderTextExtractor.GetText or a hand-built pipeline. That means extraction quality is partly your responsibility: the library gives you letters and geometry, and the layout heuristics decide the rest. For clean, single-column, machine-generated PDFs this is rarely a problem. For scanned or heavily designed documents, expect to evaluate output per document family rather than assume one configuration fits all.

## PdfPig versus PDFsharp and other .NET PDF libraries

The search data around this project shows people comparing it with PDFsharp and with iTextSharp. The difference in approach is worth stating plainly. PdfPig's centre of gravity is reading and analysing content: letters, words, text blocks, reading order, and export formats such as ALTO XML, hOCR and PAGE XML, which is what the repository topics list. PDFsharp is a general PDF manipulation library, and its emphasis is drawing and document assembly. iTextSharp is a commercial-licensed lineage, so the licence question, not the API, is often the deciding factor for teams that cannot take a copyleft or paid dependency. If your task is generating invoices with precise drawing commands, a manipulation-focused library is the more natural fit and PdfPig's creation support will feel thin. If your task is pulling text and layout out of documents that already exist, PdfPig's extractor chain is the part you actually want, and the Apache-2.0 licence removes the commercial review step. The honest summary is that these libraries overlap in the middle and diverge at the edges, and the edges are where most real projects live.

## Version pinning, licence and the cost of upgrading

The README carries a warning that is easy to skim past: while the version is below 1.0.0, minor versions will change the public API without warning, and SemVer will not be followed until 1.0.0 is reached. The release history shows the pattern, with 0.1.14, 0.1.15 and 0.1.16 arriving across 2026 rather than a stable major line. For a library you embed in a product, that means pinning an exact version and budgeting for a read of the release notes and a compile pass at each bump, because a minor number does not promise compatibility here. The licence is Apache-2.0, which permits commercial use and modification and requires preservation of notices; the repository also carries a NOTICES.txt file, which is the kind of thing your legal review will ask about. This is not legal advice, and the licence text is the authority, but for most teams the practical effect is that PdfPig does not introduce the commercial licensing question that iTextSharp does. The maintenance picture is straightforward: the last push to the default branch was on 2026-09-28, and the most recent release is v0.1.16 from 2026-08-22.

## Conclusion

Adopt PdfPig if you are writing .NET code that must read text, words or page geometry out of PDFs and you want an Apache-2.0 dependency rather than a commercial SDK. Do not adopt it as a form filler, a metadata editor, or a general PDF manipulation library, because the README lists those as unsupported. Before committing, verify two things against your own files: that the version you pin is the one you actually want, since the README warns minor versions below 1.0.0 can change the public API without warning, and that ContentOrderTextExtractor.GetText produces the reading order your documents need, since page.Text preserves internal content order and is explicitly not recommended.

## FAQ

### How do I install PdfPig?

Install the PdfPig package from NuGet, either from the NuGet page linked in the README or with the package manager console command Install-Package PdfPig. The README also notes the package is available from the releases tab.

### How do I use PdfPig to extract text from a PDF?

Open the file with PdfDocument.Open inside a using statement, iterate document.GetPages(), and call ContentOrderTextExtractor.GetText(page) for each page. The README warns against reading page.Text directly, because that property preserves the internal content order rather than the order you want.

### What is PdfPig?

PdfPig is a C# library that reads text and other content from PDF files and also supports basic PDF file creation. The repository describes it as a port of PDFBox, and it is published under the Apache-2.0 licence.

### Is PdfPig free for commercial use?

The repository is licensed under Apache-2.0, which permits commercial use and modification subject to the licence terms, including preservation of notices. The project also ships a NOTICES.txt file. This is not legal advice; read the licence text for your own situation.

### Is PdfPig open source?

Yes. The repository is licensed under Apache-2.0 and the source is on GitHub under UglyToad/PdfPig. The README links the wiki for additional examples.

## Sources

- [License: Apache-2.0](https://github.com/UglyToad/PdfPig/blob/master/LICENSE)
- [Project website](https://github.com/UglyToad/PdfPig/wiki)
- [README](https://github.com/UglyToad/PdfPig/blob/master/README.md)
- [Releases](https://github.com/UglyToad/PdfPig/releases)
- [UglyToad/PdfPig on GitHub](https://github.com/UglyToad/PdfPig)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/uglytoad-pdfpig
