# LiteParse: local PDF parsing with bounding boxes, OCR routing and a shared lit CLI

> LiteParse is LlamaIndex's open source PDF parser for spatial text, bounding boxes and Markdown, running entirely on your machine. It is fast and rule-based, and the README says so about the cases where it will not be enough.

**run-llama/liteparse** — A fast, helpful, and open-source document parser. The representation follows LlamaParse PDFium path extraction; LiteParse calls the shape rectangle bbox rather than PDFium's coords, and uses width / height rather than w / h.

- Repository: https://github.com/run-llama/liteparse
- Website: https://developers.llamaindex.ai/liteparse/
- Stars: 12,739 · Forks: 868
- Language: Rust
- License: Apache-2.0
- Published: 2026-08-08 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/run-llama-liteparse

## What LiteParse is for, and who should reach for it

LiteParse is a standalone PDF parsing tool from the run-llama organisation, the group behind LlamaIndex. The README describes it as focused exclusively on fast and light parsing: spatial text extraction with bounding boxes, no proprietary LLM features, no cloud dependency, everything running locally. It is a Rust project under Apache-2.0, and the repository is a Cargo workspace with separate crates for the PDFium bindings and each language binding.

The audience is narrower than a general document AI product. If your pipeline already has PDFs and you need positioned text so you can reconstruct layout, feed a RAG index, or hand page screenshots to an agent, LiteParse targets exactly that. The README is explicit that the spatial representation follows the LlamaParse PDFium path extraction, with two naming choices worth knowing when you port code: LiteParse calls the shape rectangle `bbox` rather than PDFium's `coords`, and uses `width` and `height` rather than `w` and `h`. That is the kind of detail that breaks a migration silently if you assume field names carry over.

It is not an OCR-first product. Tesseract is bundled and there is an HTTP OCR interface for EasyOCR, PaddleOCR or a custom server, but the primary path is text already present in the PDF. The README frames OCR as selective, merged with native text, not as the default route through a document.

## The Rust core: format conversion, PDFium extraction, selective OCR, grid projection

The README includes a Mermaid flowchart of the architecture, and it is the clearest description of the data flow available. Input formats are PDF, DOCX, XLSX, PPTX and images. Those enter a Rust core with four named stages: format conversion, text extraction, selective OCR, and OCR merge, followed by grid projection for spatial layout reconstruction.

Format conversion handles the non-PDF inputs. The flowchart attributes it to LibreOffice plus the Rust image, resvg and usvg crates, so DOCX, XLSX and PPTX are not parsed natively; they are converted first. That is a deployment dependency, not an implementation detail. The Dockerfile installs `libclang-dev`, `libtesseract-dev`, `libleptonica-dev`, `cmake` and `g++` in the build stage, and the runtime stage carries `libtesseract5`, `liblept5`, `tesseract-ocr-eng` and the pdfium shared library under `/usr/local/lib/pdfium-rs`, with `LD_LIBRARY_PATH` set to match. If you build from source outside Docker, those are the libraries you are signing up for.

Text extraction goes through the PDFium C library. OCR then runs selectively, meaning the tool decides where native text is insufficient rather than OCRing every page. Results are merged back with the native text, and grid projection reconstructs spatial layout from the positioned fragments. Complexity detection sits alongside this as a cheap pre-check: the README says it lets you check whether a document needs OCR or heavier parsing so you can route, reject or estimate cost before a full parse. For a batch pipeline, that pre-check is the difference between paying for a full parse on every file and filtering first.

## Installing LiteParse and running a first parse

Every distribution except WASM ships the same `lit` CLI, so the install channel you pick does not change the commands. The README gives one install line per language: `npm i -g @llamaindex/liteparse` for Node.js and TypeScript, `pip install liteparse` for Python, `cargo install liteparse` for the CLI or `cargo add liteparse` for the library, and `npm i @llamaindex/liteparse-wasm` for the browser build.

For a first run, install the Python package and parse a PDF to Markdown:

```bash
pip install liteparse
lit parse document.pdf --format markdown -o output.md
```

You should get a Markdown file at `output.md`. The README describes this output as structured Markdown with headings, tables, lists, images and links, reconstructed from the spatial layout by heuristics and rules. It also states plainly that complex documents may not render perfectly, which is the trade for speed.

The JSON path is where the spatial data lives. To get bounding boxes and per-item text metadata:

```bash
lit parse document.pdf --format json --extract-text-metadata -o output.json
```

The `--extract-text-metadata` flag adds rich per-item PDF text metadata to the JSON. If you also need vector content, `--extract-vector-graphics` includes page-scoped vector path data. Page selection uses `--target-pages` with a range syntax the README shows as `"1-5,10,15-20"`, and `--no-ocr` skips OCR entirely when you know the document has a text layer.

Image handling has its own rules worth reading before you write output paths. `--image-mode` accepts `placeholder` (the default, emitting references in reading order), `off`, and `embed`. Only `--extract-images` actually extracts embedded image bytes, and `--image-output-dir` requires it. JSON output carries each image's name, path, page bbox, pixel dimensions, rotation, format and duplicate relationship, but never the pixel bytes themselves. Identical image resources reuse the same output file, which matters if you are counting files after a batch run.

## Worker pool mode and the PDFium concurrency limit

One design constraint gets a parenthetical in the README that deserves more attention than it usually gets: PDFium serializes concurrent parses. That means a naive threaded pipeline around LiteParse will not scale the way the Rust implementation might lead you to expect, because the bottleneck is in the C library, not in the Rust code around it.

The answer is worker pool mode, available in the Python and Node.js bindings. The README describes it as parsing in persistent worker processes for true parallelism, with hard per-parse timeouts: rogue documents are killed, identified by name, and never stall the pipeline. That last clause is the real feature. A malformed or pathologically complex PDF can pin a parser indefinitely, and a timeout that names the offending file turns an outage into a log line you can act on.

The trade is operational. Persistent worker processes mean you manage process lifecycle, memory growth and recycling yourself, in a runtime that the Rust CLI user never has to think about. If you are parsing a few hundred documents interactively, the pool is overhead you do not need. If you are parsing continuously and one bad file can block a queue, the pool is the reason to choose the Python or Node binding over shelling out to `lit`.

## Where LiteParse is the wrong tool

The README answers this before a critic can. It names the documents that will do badly: dense tables, multi-column layouts, charts, handwritten text and scanned PDFs. For those it recommends LlamaParse, the company's cloud parser, and says you will get significantly better results there. That is an unusually direct statement of limits, and it should be read as a product boundary rather than marketing.

The mechanism behind the boundary is that Markdown reconstruction is purely heuristics and rule-based. There is no model deciding whether a run of text is a table cell or a caption; the grid projection and the rules decide. On a clean single-column report that is fast and predictable. On a two-column academic paper, reading order is exactly the thing rules get wrong, and you will get interleaved text that looks plausible and is not.

There is a second boundary that is easy to miss: images. The Markdown image modes control presentation only. The README states that placeholder references are still discovered without bytes, so `--image-mode placeholder` does not mean the images are available to you. If downstream code needs the actual image files, `--extract-images` is the only option that enables extraction, and `--image-output-dir` requires it. Teams that assume the default mode captures images will find empty references.

Finally, the cloud upgrade path in the README is not neutral. It links to LlamaParse with tracking parameters and a signup call to action. That is normal for an open source project with a commercial sibling, and it also tells you where the maintainers expect the hard cases to go.

## LiteParse against LlamaParse and Docling

The comparison people actually search for is LiteParse versus LlamaParse, and the difference is architectural rather than a matter of degree. LiteParse runs locally, has no cloud dependency, and does spatial text parsing plus rule-based Markdown. LlamaParse is the cloud parser the README describes as built for production document pipelines, handling dense tables, multi-column layouts, charts, handwriting and scans. One sends nothing off the machine and accepts the accuracy ceiling that comes with heuristics; the other sends the document to a service and targets the cases LiteParse names as out of scope.

The choice is often not either-or. Complexity detection exists precisely so you can route: check cheaply, then decide whether a document goes through the local parse or gets escalated. That is a more useful pattern than picking a winner, because the cost profile differs per document rather than per organisation.

Docling is the other name that comes up in searches. It is a document conversion project with its own model-based pipeline, and the difference in approach is that Docling brings layout and table models into the conversion step rather than relying on rule-based reconstruction. That is heavier to install and slower per page, and it is aimed at the document classes LiteParse declines. If your corpus is clean single-column PDFs, the extra machinery buys you nothing. If your corpus is the hard list from the LiteParse README, the rule-based path is the wrong starting point regardless of how fast it is.

## Licence, maintenance and what an upgrade costs you

LiteParse is Apache-2.0, declared in the workspace `Cargo.toml` and shown in the README badge row. Apache-2.0 is a permissive licence with an explicit patent grant, which is generally the least friction option for commercial embedding, but the repository also ships a `SECURITY.md` and the bindings pull in third-party components with their own terms. PDFium is a notable one: it is bundled through the `pdfium-sys` and `pdfium` crates in this workspace, and its own licensing is separate from LiteParse's. Check the dependency licences for your distribution model rather than assuming the top-level Apache-2.0 covers everything you ship.

The repository is not archived, and the last push was on 2026-08-27. The recent releases listed are `wasm-v2.14.2` and `python-v2.14.2`, both dated 2026-08-27, with `wasm-v2.14.1` earlier the same day. Version numbers are per-language, so the WASM and Python packages move on their own tags rather than one unified release line.

That versioning scheme is the main upgrade cost. Because Node, Python, Rust and WASM are published separately, a fix landing in the Rust core does not automatically appear in your `pip install liteparse` until the Python tag is cut. Pinning versions and reading `CHANGELOG.md` before bumping is the practical approach. There is also a V1 code path preserved on a separate branch, which the README links to; if you adopted LiteParse before the current line, that branch is where the old behaviour lives, and the field naming differences between PDFium and LiteParse mean a migration is not a drop-in replacement.

## Conclusion

Adopt LiteParse when you need spatial text, bounding boxes or Markdown from ordinary PDFs without sending files to a cloud service, and when a rule-based renderer is acceptable. Do not adopt it for dense tables, multi-column layouts, charts, handwriting or scanned documents; the README itself points those at LlamaParse. Before committing, run the same document through `lit parse --format markdown` and `lit parse --format json --extract-text-metadata` and compare the two outputs against what your pipeline actually consumes, because the JSON path and the Markdown heuristic renderer do not carry the same information.

## FAQ

### What are the key differences between LiteParse and LlamaParse?

LiteParse runs entirely locally and does spatial text parsing with bounding boxes plus rule-based Markdown, with no cloud dependency. LlamaParse is the cloud-based parser the README recommends for dense tables, multi-column layouts, charts, handwritten text and scanned PDFs.

### Is LlamaParse free?

The README links to a free LlamaParse signup, but it does not document pricing tiers or usage limits beyond that call to action. LiteParse itself is the open source local alternative and carries no cloud dependency.

### Is LiteParse open source?

Yes. The repository is licensed Apache-2.0 and the README describes it as a standalone OSS PDF parsing tool that runs locally with no proprietary LLM features.

### How to install LiteParse?

Install through your package manager: `pip install liteparse` for Python, `npm i -g @llamaindex/liteparse` for Node.js and TypeScript, `cargo install liteparse` for the CLI, or `npm i @llamaindex/liteparse-wasm` for the browser. All versions except WASM ship the same `lit` CLI.

### How to use LiteParse?

Run `lit parse document.pdf` for a basic parse, or add `--format markdown -o output.md` to render Markdown, `--format json` for bounding boxes and spatial data, and `--target-pages "1-5,10,15-20"` to limit which pages are parsed. The README notes that Markdown reconstruction is heuristics and rule-based, so complex documents may not render perfectly.

### What is LiteParse?

LiteParse is a standalone open source PDF parsing tool from run-llama focused on fast and light parsing. It provides spatial text parsing with bounding boxes and can output Markdown, JSON or text, using PDFium for extraction and Tesseract or an HTTP OCR server for selective OCR.

## Sources

- [Official documentation](https://developers.llamaindex.ai/liteparse/)
- [Official README](https://github.com/run-llama/liteparse#readme)
- [Project repository](https://github.com/run-llama/liteparse)
- [Release notes](https://github.com/run-llama/liteparse/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/run-llama-liteparse
