ParseBench: a document parsing benchmark for AI agent workflows
ParseBench - A Document Parsing Benchmark for AI Agents
At a glance
- What is it?
- ParseBench scores PDF parsers on tables, charts, faithfulness, formatting and visual grounding across roughly 2,000 human-verified enterprise pages, and the leaderboard is led by LlamaParse Agentic Plus.
- Who is it for?
- Adopt ParseBench if you are choosing or regression-testing a PDF parsing pipeline for agent workflows and want per-dimension scores instead of a single similarity number. Skip it if you only need OCR text, if your documents are not enterprise PDFs with tables and charts, or if you cannot supply API keys for the providers you want to compare.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What ParseBench measures that a text-similarity score does not
Most parser evaluations compare extracted text to a reference string and call it done. ParseBench starts from the opposite end: whether the parsed output preserves the structure and meaning an agent needs in order to act. That reframing matters because a parser can reproduce every word of a page and still destroy the table that the agent was supposed to read.
The target audience is narrow and identifiable. Teams building retrieval or agent pipelines over enterprise PDFs, insurance and finance documents in particular, need to know which parser survives contact with merged table cells, footnotes and chart legends. The benchmark covers roughly 2,000 human-verified pages drawn from real enterprise documents in insurance, finance and government. It is not a general OCR quality suite and it is not aimed at academic text extraction.
The five dimensions are the substance of the project. Tables use GTRM, a combination of GriTS and TableRecordMatch. Charts use ChartDataPointMatch. Content Faithfulness and Semantic Formatting score the same underlying text documents against different rule sets, which is why the page counts in those two rows are nearly identical. Visual Grounding reports an Element Pass Rate. Each dimension exists because its failure mode breaks a production agent workflow, not because it was convenient to measure.
How the harness runs a pipeline and scores it
A pipeline in ParseBench is a named document parsing tool or configuration. The README states there are more than 180 of them, listed in docs/pipelines.md and printable with the CLI. The paper baselines alone cover 21 pipelines, spanning LlamaParse variants, OpenAI and Anthropic and Google Gemini models, cloud services such as Azure Document Intelligence, AWS Textract and Google Cloud Document AI, commercial APIs such as Reducto, Extend and LandingAI, and open-weight options including Qwen 3 VL, Dots OCR 1.5 and Docling.
The data flow is straightforward. The CLI loads configuration from a .env file on startup, resolves a pipeline name to its provider SDK, sends the benchmark PDFs through that provider, and scores the returned output against the ground truth for each dimension. Ground truth is stratified: table.jsonl carries 503 pages across 284 documents, chart.jsonl carries 568 pages and 4,864 rules, text_content.jsonl carries 506 pages and 141,322 rules, text_formatting.jsonl carries 476 pages and 5,997 rules, and layout.jsonl carries 500 pages and 16,325 rules. The unique total is 2,078 pages and 1,211 documents, with 169,011 rules overall.
The dependency list tells you what the scoring layer actually is. apted provides tree edit distance, python-Levenshtein, rapidfuzz and fuzzysearch handle string matching, lxml and beautifulsoup4 parse markup, and anls-star supplies the ANLS metric. An optional fast extra pulls in numba for a JIT-accelerated TEDS table metric; the README says scores are identical to the default path and the difference is speed on large tables. That claim is worth verifying on your own hardware rather than taking on faith.
Installing ParseBench and running a first pipeline
Prerequisites are Python 3.12 or newer and a .env file holding the API key for whichever parsing tool you intend to evaluate. The repository ships .env.example with the variable names already laid out, including LLAMA_CLOUD_API_KEY, OPENAI_API_KEY, ANTHROPIC_API_KEY, GOOGLE_GEMINI_API_KEY, AZURE_DOCUMENT_INTELLIGENCE_KEY and AZURE_DOCUMENT_INTELLIGENCE_ENDPOINT, plus AWS credentials and region. Copy it to .env and fill in only the providers you plan to run.
Install from PyPI, choosing the extras for your providers. The runners extra installs every provider SDK and is the heaviest option; single-provider extras such as llamaparse, openai, anthropic and google keep the environment smaller.
pip install "parse-bench[runners]"
pip install "parse-bench[llamaparse]"If you prefer working from a checkout, the README gives the uv path instead, and the optional fast extra adds the numba-backed TEDS metric.
uv sync --extra runners
pip install "parse-bench[runners,fast]"Before committing to a full run, use the test mode. It uses a small dataset with three files per category, which the README describes as good for trying things out. Drop the uv run prefix if you installed from PyPI.
uv run parse-bench run llamaparse_agentic --test
uv run parse-bench run llamaparse_agentic
uv run parse-bench serve llamaparse_agenticThe first command should finish quickly and produce scores for the reduced set. The second runs the full benchmark and will take far longer because it hits every page in the dataset. The third starts a local report server so you can inspect results in a browser. To see what else is available before choosing, list the pipelines.
uv run parse-bench pipelinesThat prints the pipeline names you can pass to run. If you want to add providers, pipelines, products or rule types from another package, the README points to docs/extending.md.
The inclusion criteria and what they exclude
ParseBench states three inclusion criteria for the leaderboard, and they are more restrictive than they first appear. The model or API must be publicly accessible through open weights or a self-serve API any user can sign up for. The benchmark run must finish in roughly single-digit hours. Concurrency can be tuned to a provider's recommended settings, but providers may not require custom framework changes, so the comparison stays fair.
The practical effect is that a slow but accurate parser can be excluded on runtime alone, and a private internal parser cannot appear at all. That is a defensible rule for a public leaderboard, but it means the ranking is a ranking of accessible, reasonably fast parsers, not of the best parsers that exist. If your organisation runs an in-house model, ParseBench can still score it locally through the extending path, but the published table will never reflect it.
The second limitation is cost. The leaderboard carries a cents-per-page column, and the spread is wide: the top entry lists 5.62 cents per page while another LlamaParse configuration lists 0.38 cents per page, and two entries show no price at all. A single Overall number hides that. A pipeline that scores a few points lower at a fraction of the cost may be the right choice, and the raw data in leaderboard.csv is where that comparison has to happen.
Where ParseBench is the wrong tool
ParseBench is a benchmark harness, not a parser. Installing it gives you no parsing capability. If you arrived looking for something to convert your PDFs, this is not it, and the runners extra will install a large set of provider SDKs you may not want.
The dataset is enterprise PDFs, weighted toward insurance, finance and government material with tables and charts. If your documents are scanned handwriting, receipts, or plain prose with no structure to preserve, the five dimensions measure things your pipeline does not depend on, and the scores will not predict your production behaviour. Content Faithfulness and Semantic Formatting dominate the rule counts, so a text-heavy corpus will be judged mostly on dimensions you may not care about.
There is also a cost and access barrier. Every cloud pipeline needs credentials, and a full run against 2,078 pages is a real API bill, not a free experiment. The --test mode exists precisely because the full run is expensive, but three files per category is a smoke test, not a measurement. Treat any conclusion drawn from --test as unverified.
Finally, the project is classified as Development Status 4 - Beta in pyproject.toml, and there have been three releases within September 2026 (v1.0.2, v1.0.3, v1.0.4). Rapid patch releases on a young version line mean you should pin the version you benchmark with, because a metric change between v1.0.3 and v1.0.4 would silently invalidate a comparison you ran last week.
How ParseBench differs from general OCR and VLM evaluation suites
The obvious alternative is a general OCR benchmark that scores character or word error rate on document images. The difference is what gets compared. An OCR benchmark asks whether the characters are right. ParseBench asks whether the table structure, the chart data points, the formatting semantics and the element positions survive, using metrics built for each of those, and it adds a cost column. A parser can win an OCR benchmark and lose badly on GTRM.
The second alternative is to build your own evaluation set from your own documents. That is the approach many teams land on, and it is not wrong: your corpus is the distribution that matters. The trade-off is ground truth labour. ParseBench reports 169,011 rules across its dimensions, all human-verified, which is not something a team assembles in a sprint. The sensible split is to use ParseBench to narrow the field and your own set to make the final call.
The third alternative is to skip benchmarking and pick the parser with the best reputation. The leaderboard is a reason not to: rank 1 and rank 4 are both LlamaParse configurations, separated by roughly ten Overall points and by a factor of about fifteen in cents per page. Reputation does not distinguish between those two, and the configuration does.
Licence, maintenance and the cost of keeping up
ParseBench is Apache-2.0, with the licence declared in pyproject.toml and a LICENSE file at the repository root. That is a permissive licence and imposes no copyleft obligation on your own code. It says nothing about the terms of the parsing providers you benchmark, and the runners extra pulls in SDKs from Anthropic, OpenAI, Google, Azure, AWS and several commercial vendors, each under its own terms. Check those separately; the Apache-2.0 grant covers ParseBench, not the services it calls.
The dataset is hosted on HuggingFace at llamaindex/ParseBench and carries its own terms, which the README does not spell out. If you intend to redistribute scores or derived data, verify the dataset licence rather than assuming it matches the code licence.
On maintenance, the repository is not archived and the last push was on 2026-09-10. The upgrade cost is real but bounded: pin parse-bench in your environment, and re-run your chosen pipelines whenever you bump the version, because metric definitions live in the scoring code and a patch release can change a number without changing your parser. The fast extra is the one upgrade worth taking early if you score large tables, since the README states the scores match the default path.
Editorial conclusion
Adopt ParseBench if you are choosing or regression-testing a PDF parsing pipeline for agent workflows and want per-dimension scores instead of a single similarity number. Skip it if you only need OCR text, if your documents are not enterprise PDFs with tables and charts, or if you cannot supply API keys for the providers you want to compare. Before trusting a score, check which pipeline name you ran, whether --test was dropped, and which dataset version the metric was computed against.
Frequently asked questions
What does it mean to parse a file in the ParseBench context?
In ParseBench, parsing means converting a PDF into structured output that an AI agent can act on, preserving tables, charts, formatting and element positions rather than producing a flat text string. The benchmark scores that output across five capability dimensions.
What is meant by document parsing?
Document parsing is the step that turns a document such as a PDF into a structured representation. ParseBench evaluates how well that step preserves the structure and meaning needed for autonomous decisions, not just how similar the text looks to a reference.
What does "AI parse" mean in ParseBench's evaluation?
ParseBench treats AI parsing as the use of a model or API, including vision-language models and commercial parsing services, to convert PDFs into structured output. Its pipelines cover both open-weight models and self-serve commercial APIs.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/run-llama-parsebench)