smalot/pdfparser: text and metadata extraction from PDFs in plain PHP
PdfParser, a standalone PHP library, provides various tools to extract data from a PDF file.
At a glance
- What is it?
- A standalone PHP library that loads PDF objects, pulls out metadata and ordered page text, and stays compatible with supported PHP versions under a limited-maintenance policy. Useful when you already run PHP and need text, not a rendering engine.
- Who is it for?
- Adopt smalot/pdfparser when your stack is PHP and the job is pulling text or metadata out of ordinary, unencrypted PDFs, especially in batches where a full renderer would be overkill. Do not adopt it if you need form field values, password-protected documents, or layout-accurate output, because the README states secured documents and form data extraction are not supported.
- Can I use it commercially?
- Yes, with conditions. LGPL-3.0 is a weak copyleft licence: you can use it inside commercial and closed-source software, but if you distribute changes to its own files, you must publish those changes under the same licence.
- Is it still maintained?
- Yes. The repository last received commits 8 days ago.
- What is it written in?
- Mainly PHP, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 24, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What smalot/pdfparser does that a PDF viewer cannot
A PDF viewer renders pages for a human. smalot/pdfparser reads the file structure for a program. The README describes it as a standalone PHP package that provides various tools to extract data from PDF files, and the feature list is concrete: load and parse objects and headers, extract metadata such as author and description, extract text from ordered pages, handle compressed PDFs, support the Mac OS Roman charset encoding, and handle hexadecimal and octal encoding inside text sections.
The audience is narrow and specific. You are writing PHP, you have PDFs arriving from somewhere (an upload form, a mail attachment store, a document archive), and you want strings or metadata out of them without shelling out to a Java service or a Python worker. Because it is a Composer package rather than a binary, it drops into an existing PHP application with no extra runtime to deploy. The README states the library requires PHP 7.1 or later since v1.0.0.
What it is not: a renderer, an OCR engine, or a form processor. The README says plainly that secured documents and extracting form data are not supported. If your input is a scan, there is no text layer to extract and this library will not invent one.
How parsing works: objects in, document object out
The data flow is two steps. First you construct a Parser and hand it a file path; the parser reads the PDF and builds the internal objects, including the document header and the object graph. What comes back is a document object, not a string. Second, you ask that object questions: getText() for the page text in order, or the metadata accessors for author, description and similar fields.
That split matters for cost. Parsing is the expensive half, and it happens once per file. If you need both metadata and text, you parse once and call both accessors on the same object. The README also points to a CustomConfig document, which implies the parser's behaviour can be adjusted through configuration rather than by patching the source, though the README itself does not enumerate the available keys; the doc folder is where those details live.
The encoding features are worth reading twice. Compressed streams, Mac OS Roman charset, and hexadecimal or octal escapes inside text sections are all listed as supported. Those are exactly the cases where a naive byte-scanning approach produces garbage, so the library is doing real work rather than just concatenating stream bytes. Note the repository ships samples/InternationalChars.pdf and samples/ImproperFontFallback.pdf, which suggests the maintainers treat character mapping and font substitution as known trouble spots rather than solved problems.
Installing smalot/pdfparser with Composer and parsing your first file
Installation is a single Composer command. Run it from your project root and the package lands in vendor/ with its autoloader registered.
composer require smalot/pdfparserIf Composer is not available in your environment, the README offers a fallback: include alt_autoload.php-dist, which the README says includes all required files automatically. That file sits at the top level of the repository.
The README's quick example is the shortest path to a working parse. It constructs the parser, parses a path, and echoes the text.
<?php
// Parse PDF file and build necessary objects.
$parser = new \Smalot\PdfParser\Parser();
$pdf = $parser->parseFile('/path/to/document.pdf');
$text = $pdf->getText();
echo $text;What you should see is the extracted text printed to stdout. If the output is empty or garbled, the first thing to check is whether the file actually contains a text layer. The repository's samples/ directory is a good place to start, since it ships files from several different generators (Foxit Reader, PDFCreator, PDF-XChange, Word) plus samples/corrupted.pdf and samples/not_really_encrypted.pdf, which let you compare behaviour across producers before you point the parser at production documents. The README also points to doc/Usage.md for further usage information and doc/CustomConfig.md for configuration.
Where smalot/pdfparser breaks down
The README's own limitation sentence is the most important line on the page: secured documents and extracting form data are not supported. If your workflow depends on reading filled form fields, this library is the wrong tool and no amount of configuration will change that, because the feature is listed as absent rather than as partially working.
Text extraction from PDFs is inherently approximate. A PDF stores glyphs positioned on a page, not sentences, so any extractor has to reconstruct reading order from coordinates and font runs. The presence of samples/ImproperFontFallback.pdf and samples/Document-Word-Landscape-printedaspdf.pdf in the repository is a signal that font substitution and non-standard page orientation are the kinds of inputs that stress that reconstruction. Expect column layouts, tables and multi-column academic papers to come out in a different order than a human would read them.
There is also a maintenance caveat that belongs in any adoption decision. The README states the library is under limited maintenance: it is kept compatible with supported PHP versions and community contributions may be accepted, but there is no active feature development and no guarantee that pull requests will be reviewed or merged in a timely manner. The README directs anyone planning a contribution beyond a small, well-scoped fix to read CONTRIBUTING.md first. The last push to the repository was on 2026-09-22, so the code is being touched, but the stated policy is compatibility work rather than new capability. If your roadmap assumes new features will arrive, that assumption is not supported by the project's own documentation.
smalot/pdfparser compared with a Python or CLI extractor
The obvious alternative for many teams is a Python PDF text library invoked as a subprocess, or a command-line extractor wrapped in a shell call. The difference is not accuracy, it is where the parsing lives. A Python tool means a second runtime in your deployment: another interpreter version to pin, another dependency set to audit, and a process boundary to manage when you are parsing thousands of files. smalot/pdfparser keeps everything inside the PHP process, which matters most when your PDFs arrive inside a PHP request or a PHP queue worker.
A second alternative is an OCR pipeline. That is a different problem, not a competing implementation. The README lists no OCR capability at all, so scanned documents with no text layer are simply out of scope. If a meaningful share of your inputs are scans, you need OCR regardless of which text extractor you choose, and adding smalot/pdfparser on top only helps for the born-digital portion of the corpus.
A third option is a full PDF toolkit that renders pages. Those give you layout information and images, at the cost of a heavier dependency and slower per-file processing. For the common case of pulling the text out of an invoice or a report to feed a search index, that machinery is unnecessary, and the README's feature list suggests this library targets exactly that narrower job.
Licence and the cost of staying current
The library is released under LGPL-3.0, and the README links to LICENSE.txt at the repository root. The practical consequence for a PHP application is that you are linking against the library rather than modifying it, which is the scenario the LGPL is designed for; if you fork the source and ship a modified version, different obligations apply. That is a description of the licence, not legal advice, and any organisation with strict dependency policy should have its own review rather than relying on a README.
Upgrade cost is low by design. Installation is a Composer require, releases are tagged (v2.12.5 on 2026-04-21, v2.12.4 on 2026-03-11, v2.12.3 on 2026-01-08), and the stated maintenance commitment is keeping the library compatible with supported PHP versions. That means the main upgrade pressure you will feel is a PHP version bump, not a breaking API change. The flip side is the same sentence read from the other direction: no active feature development means bugs you hit in text extraction are unlikely to be fixed on your schedule. If you need a fix, the README's guidance is to read CONTRIBUTING.md before proposing anything beyond a small, well-scoped change, and even then it makes no promise of a timely review.
The repository ships a Makefile that shows how the maintainers run their own checks: install-dev-tools runs composer update in dev-tools, and run-phpunit, run-phpstan and run-php-cs-fixer invoke the corresponding tools from dev-tools/vendor/bin. If you fork the library, that is the test and static-analysis setup you inherit.
Editorial conclusion
Adopt smalot/pdfparser when your stack is PHP and the job is pulling text or metadata out of ordinary, unencrypted PDFs, especially in batches where a full renderer would be overkill. Do not adopt it if you need form field values, password-protected documents, or layout-accurate output, because the README states secured documents and form data extraction are not supported. Before committing, parse a representative sample of your own files, including a scanned one and one with international characters, and check what getText() actually returns; the samples directory in the repository is a useful starting set for that check.
Frequently asked questions
How do I install smalot/pdfparser?
Install it with Composer by running composer require smalot/pdfparser from your project root. The README notes that if you cannot use Composer, you can include alt_autoload.php-dist instead, which includes all required files automatically.
How do I use smalot/pdfparser to get text from a PDF?
Construct a Smalot\PdfParser\Parser instance, call parseFile() with the path to your document, then call getText() on the returned object. The README's quick example echoes that text directly.
What is smalot/pdfparser used for?
It is a standalone PHP package for extracting data from PDF files. The README lists loading and parsing objects and headers, extracting metadata such as author and description, and extracting text from ordered pages.
Is smalot/pdfparser a replacement for OCR?
No. OCR is not in the README's feature list, which covers parsing objects and headers, metadata extraction, ordered page text, compressed PDFs, Mac OS Roman charset support, and hexadecimal and octal encoding in text sections. Scanned pages with no text layer are outside its scope.
What is the best alternative to smalot/pdfparser?
The README does not compare the library with other tools. It does state what the library does not cover, namely secured documents and form data extraction, so an alternative is only needed if your inputs fall into those categories or are scanned images.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/smalot-pdfparser)