# GROBID: Machine Learning Library for Extracting Structure from Scholarly PDF Documents

> GROBID (GEneRation Of BIbliographic Data) is an open-source Java library for parsing and restructuring scholarly PDF articles into structured XML/TEI output, covering everything from bibliographic metadata to full-text section structure and citation context resolution. It is production-tested at ResearchGate, Semantic Scholar, and CERN, and runs as a REST service via Docker with no GPU required in its default CRF configuration.

**grobidOrg/grobid** — A machine learning software for extracting information from scholarly documents

- Repository: https://github.com/grobidOrg/grobid
- Website: https://grobid.readthedocs.io
- Stars: 5,130 · Forks: 569
- Language: Java
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/grobidorg-grobid

## What GROBID Extracts and Who Uses It in Production

GROBID addresses a specific problem: scientific PDF articles contain rich structured information, but that information is encoded in a page-layout format that does not preserve structure. Author names, affiliations, keywords, section headings, figure captions, bibliographic references, and citation context markers are visually distinguishable to a human reader, but buried inside a binary format that treats them all as positioned text runs.

GROBID converts that binary PDF into structured XML in the TEI (Text Encoding Initiative) format. TEI is a standard used in digital humanities for encoding scholarly documents with semantic markup. Each header field, each reference, each section, and each citation callout gets its own XML element with the appropriate attributes. A downstream system can then process the structured output without needing to understand PDF internals.

The feature list in the README covers 68 final labels that GROBID uses to identify structures: title, abstract, author names, affiliations, keywords, dates, DOIs, PMIDs, section headings, paragraphs, figures, tables, data availability statements, footnotes, and more. It handles patent documents in addition to journal articles.

Production deployments documented in the README include ResearchGate, Semantic Scholar, HAL Research Archive, scite.ai, Academia.edu, Internet Archive Scholar, INIST-CNRS, and CERN's Invenio repository software. These are significant-scale scholarly infrastructure organizations, which gives a reasonable indication that the tool handles large volumes of PDFs reliably.

The README traces the project's origins to 2008 and notes it was made open source in 2011. Development has continued as a side project supported by Inria (Institut national de recherche en informatique et en automatique) in France.

## The Extraction Pipeline: CRF Models and Deep Learning Models

GROBID can run two different types of machine learning models for each extraction task: Conditional Random Field (CRF) models and Deep Learning models based on transformer architectures or RNNs with optional layout feature channels.

The CRF models are the default. They are feature-engineered, meaning that hand-crafted features describing the text layout, font, and position feed the statistical model. The CRF models run on any hardware without special dependencies beyond OpenJDK 21. They are what you get when you run GROBID out of the box.

The Deep Learning models require additional setup: Python 3.10 or 3.11 with JEP (Java Embedded Python) to bridge the Java and Python runtimes, and ideally a NVIDIA GPU with CUDA support for acceptable throughput. The README states explicitly that the Deep Learning models are not used by default to 'accommodate out of the box hardware.' You must configure which DL models to enable by editing the GROBID configuration file.

The tradeoff between CRF and DL models is concrete. The README states that 'some GROBID Deep Learning models perform significantly better than default CRF, in particular for bibliographical reference parsing.' For teams whose workload is primarily reference parsing and who have GPU infrastructure, enabling the DL models delivers measurably better F1-scores on that task.

The DL framework GROBID uses is DeLFT (Deep Learning Framework for Text), which is a separate repository developed by the same group. JEP is the bridge that allows calling the DeLFT Python code from the GROBID Java process. The visual and layout information that informs both CRF and DL models comes from pdfalto, another component in the same ecosystem.

## Installing GROBID and Running It as a Service

GROBID runs on Linux (64-bit) and macOS (Intel and ARM) for native builds. The README notes that Windows support is no longer guaranteed, though it worked in previous versions. The two primary deployment paths are building from source and using the Docker images.

Building from source requires OpenJDK 21. The build system is Gradle, and the repository includes a Gradle wrapper. Two Docker images exist: one using the CRF models only (Dockerfile.crf) and one that also supports Deep Learning (Dockerfile.delft). The Dockerfile names correspond to the Gradle build targets.

Docker is the recommended path for most users who do not need to modify the source. The Docker Hub images at grobid/grobid and lfoppiano/grobid are listed in the README's badge section. The documentation at grobid.readthedocs.io describes the exact docker run commands and port mappings.

Once running, GROBID exposes a REST API. The web service documentation covers the available endpoints, the request format (typically a PDF file upload), and the response format (TEI XML or plain text, depending on the endpoint). A free live demo is available at grobidOrg-grobid.hf.space, which is hosted on Hugging Face Spaces.

Batch processing is available as an alternative to the REST API for offline processing of large collections of PDFs. The README mentions batch processing as one of the four deployment modes alongside the web service API, the Java API, and a training and evaluation framework.

## Extraction Accuracy: What the Benchmarks Show

GROBID documents specific accuracy figures in the README, measured against independent test sets.

For bibliographic reference extraction from PDFs, the README states approximately 0.87 F1-score on a set of 1,943 PubMed Central PDFs containing 90,125 references, and approximately 0.90 on a similar 2,000-paper bioRxiv set. Both figures are for the Deep Learning citation model. The evaluation is at the field level: each reference field (author, title, year, journal, volume, pages, DOI, PMID, and others) is scored independently.

For citation context resolution, the accuracy depends on the collection: between 0.76 and 0.91 F1-score. This task requires both identifying the citation callout in the text (the number in brackets or the author-year in parentheses) and correctly linking it to the full bibliographic reference. Getting both right is harder than getting either alone.

For reference parsing in isolation (a single reference string without a surrounding article), the DL model achieves above 0.90 F1-score at the instance level and 0.95 at the field level.

For consolidation and resolution of bibliographic references against external databases (CrossRef or biblio-glutton), DOI and PMID resolution performance is above 0.95 F1-score from PDF extraction.

These figures are for the Deep Learning models. The CRF models produce lower scores on most tasks but remain useful for high-volume deployments on commodity hardware where the per-document inference time with DL models is prohibitive.

## The REST API and Integration Patterns

GROBID's primary integration point for most projects is the REST API. Once the service is running on its default port, client code submits a PDF file as a multipart/form-data request and receives TEI XML in the response. The service exposes separate endpoints for different extraction scopes: header extraction only, full-document extraction, reference extraction only, and several others.

A Python client library for GROBID is a separate community project that simplifies calling the REST API from Python code. The grobid-client-python repository (not part of this repository) handles connection management and response parsing. The README mentions Python client access as a well-known use pattern, and the RELATED SEARCHES data for this project confirms that 'how to use grobid in python' and 'grobid python client' are common search queries.

The Java API is available for projects that embed GROBID directly in a JVM application rather than running it as a separate service. This avoids the HTTP overhead for very high-throughput scenarios.

PDF coordinates are a feature worth noting for interactive applications: GROBID can return the bounding box coordinates in the PDF for each piece of extracted text. This allows building annotation tools or augmented PDF viewers that highlight identified structures in the original page layout, not just in the extracted text.

The training and evaluation framework is part of the same Gradle project. Teams that want to fine-tune models for domain-specific literature (clinical trials, patent documents, or specific subfields of physics, for example) can use GROBID's built-in tooling to retrain and evaluate models on their own annotated datasets.

## Limitations: Windows, GPU Requirements, and Model Configuration Complexity

Windows support has been dropped. The README states: 'We cannot ensure currently support for Windows as we did before.' This affects teams whose build or deployment infrastructure is Windows-only. The Docker option works on Windows via Docker Desktop, but native builds without Docker require Linux or macOS.

Enabling the Deep Learning models requires a specific Python version. The README says Python 3.10 or 3.11 is required with JEP. Other Python versions are not supported. Python 3.12 and later are not listed, which may create friction as operating system Python packages move to newer versions. The JEP bridge between the Java and Python runtimes also adds complexity to the deployment: Java and Python must share the same library path, and incompatible JEP versions have historically caused difficult-to-debug startup failures.

The GPU requirement for Deep Learning models is a GPU with CUDA support, which means NVIDIA hardware. There is no documented support for AMD GPUs or Apple Silicon Metal acceleration. Teams using macOS ARM machines can run the CRF models natively but not the DL models at GPU speed.

Model selection is configuration-level work. The README points to a configuration file where you enable or disable specific DL models per task, but the configuration details are in the documentation at grobid.readthedocs.io rather than in the README itself. Teams adopting GROBID for the first time should budget time to understand which models to enable for their specific extraction tasks and what hardware those models require.

The project does not support documents in formats other than PDF and patents. Web pages, Word documents, and images of printed text are not in scope.

## GROBID Compared to MinerU, Marker, and General-Purpose PDF Tools

Tools like MinerU and Marker take a different approach to PDF extraction than GROBID. They are designed primarily to convert PDFs into clean markdown or structured text for use with large language models and information retrieval systems, handling a broad range of document types including presentations, reports, and books. Their output format is plain text or markdown rather than TEI XML, and their evaluation metrics focus on text extraction quality rather than scholarly metadata accuracy.

GROBID's comparative strength is its depth of scholarly document understanding. Its 68 structural labels cover granular academic elements that general-purpose converters do not model: affiliation parsing (institution, department, city, country), citation callout detection and linking, funder extraction, copyright and license identification, and PDF coordinate output for every extracted element. This level of detail is what makes GROBID suitable for building academic knowledge graphs, citation analysis pipelines, and literature review systems.

The comparative weakness is scope. GROBID is calibrated for peer-reviewed scientific literature. It performs poorly on other document types and has no mechanism for extracting content from scanned documents that are not born-digital PDFs. MinerU and similar tools handle a wider range of input documents.

For a research infrastructure project that needs to index large corpora of scientific articles with accurate metadata, references, and citation links, GROBID is the more appropriate choice. For a product that needs to make arbitrary PDFs searchable in plain text, a general-purpose converter is simpler to operate and does not require GPU-backed model deployment for high accuracy.

## Conclusion

GROBID is the right tool when you need structured, fine-grained extraction from scientific PDFs, particularly for bibliographic reference parsing, author affiliation extraction, or citation context resolution. It is not the right tool for general-purpose PDF conversion to plain text or markdown, for extracting content from non-scientific documents, or for Windows deployments, which the README acknowledges are no longer guaranteed to work. Before deploying, decide whether to use the default CRF models (fast, runs on any hardware) or the Deep Learning models (higher accuracy, requires Python 3.10-3.11 and a GPU). Version 0.9.1 was released on 2026-08-04.

## FAQ

### What is GROBID and what does it do?

GROBID (GEneRation Of BIbliographic Data) is an open-source Java library that parses scientific PDF articles and restructures them into XML/TEI output. It extracts bibliographic headers, references, citation contexts, full-text sections, figures, and tables using machine learning models.

### How do I install GROBID?

The recommended path is Docker using the official grobid/grobid image, which avoids the OpenJDK 21 build requirement. For building from source, the README directs to the Installation documentation at grobid.readthedocs.io, which covers JDK setup and platform-specific requirements for Linux and macOS.

### Is GROBID free and open source?

Yes. GROBID is released under the Apache License 2.0, which permits free use, modification, and distribution including in commercial products. The repository is at github.com/grobidOrg/grobid and has been open source since 2011.

### How does GROBID compare to MinerU or Marker?

GROBID is specialized for scholarly documents and produces structured XML/TEI output with granular academic metadata including citation linking, affiliation parsing, and PDF bounding box coordinates. MinerU and Marker are general-purpose PDF-to-text converters that handle a wider range of document types but do not offer the same depth of scholarly structure extraction.

## Sources

- [grobidOrg/grobid on GitHub](https://github.com/grobidOrg/grobid)
- [License: Apache-2.0](https://github.com/grobidOrg/grobid/blob/master/LICENSE)
- [Project website](https://grobid.readthedocs.io)
- [README](https://github.com/grobidOrg/grobid/blob/master/README.md)
- [Releases](https://github.com/grobidOrg/grobid/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/grobidorg-grobid
