Apache Tika: Unified Text and Metadata Extraction for Java Pipelines
The Apache Tika toolkit detects and extracts metadata and text from over a thousand different file types (such as PPT, XLS, and PDF).
At a glance
- What is it?
- Apache Tika extracts text and structured metadata from over a thousand file formats through a single Java API. Version 4.x targets LLM and RAG pipelines with Markdown output by default and crash-isolated parsing processes.
- Who is it for?
- Engineers building document ingestion pipelines in Java should consider Apache Tika when format breadth and process isolation matter more than zero-dependency footprint. The 4.x line requires Java 17 and a JSON-based configuration format; teams on Java 8 should note that Tika 2.x reached EOL in April 2025.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 4 days ago.
- What is it written in?
- Mainly Java, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Apache Tika Solves for Java Developers
Apache Tika is a Java library that detects file types and extracts both raw text and structured metadata from them. It covers over a thousand formats: PDF, Microsoft Office documents (PPT, XLS, DOCX), OpenDocument files, images, audio, video, archives, and many others, all through a single API. The project targets two overlapping audiences: data engineers who need to index heterogeneous document collections, and developers building LLM and RAG pipelines who need documents converted into a form a language model can consume.
The 4.x line, requiring Java 17 or later, introduced Markdown as the default output format. This change was designed specifically for the RAG use case, where the consuming application expects structured, readable text with preserved headings and lists rather than stripped plain text. Tika 4.x also runs each parser in a crash-isolated forked process, meaning a malformed or hostile document can crash its parser without bringing down the host application. These two changes make the current version meaningfully different from earlier lines, not just an incremental update.
The library is a project of the Apache Software Foundation and is licensed under Apache-2.0. Pre-built binaries are available at tika.apache.org/download.html, and all jars are published to Maven Central.
Crash-Isolated Forked Processes and Recursive Metadata Extraction
Each parse operation in Tika 4.x runs in a separate process. When a parser for a specific format fails or hangs, only that forked process is affected. The parent JVM continues. This is a significant design commitment for server-side deployments where untrusted documents are a realistic threat, and it is one of the more concrete safety guarantees in Tika's architecture.
The `-J` flag on the command-line tool produces recursive metadata JSON: a structured object containing the metadata and content of the document itself plus any embedded documents found inside it. A PDF with embedded attachments, an email with image attachments, or a ZIP with nested archives all expand into their constituent parts in that output. The README describes this as "structured recursive extraction," and the same capability is available via the `/rmeta` endpoint when running the tika-server REST component.
The repository root contains a `.skills/` directory with standalone agent skills for AI agent frameworks. The `file-to-markdown` skill uses tika-app or tika-server as its backend. The `file-to-markdown-docker` skill wraps a containerized Tika with OCR included. The README states these can be copied into any agent's skill directory without depending on this repository.
Installing Apache Tika and Running a First Parse
For the command-line tool, download the tika-app ZIP from tika.apache.org/download.html and unzip it into a directory. The archive has no top-level folder of its own. Run from inside that directory because the jar is a thin launcher that loads parsers from the adjacent `lib/` folder. Running it from a different directory causes a `NoClassDefFoundError`.
A basic parse using the 4.x defaults:
java -jar tika-app-<version>.jar document.pdf # Markdown (the 4.x default)
java -jar tika-app-<version>.jar --text document.pdf # plain text
java -jar tika-app-<version>.jar -J document.pdf # structured JSON: metadata +
# content for the document AND
# anything embedded in itFor Java projects on Maven, add the standard parsers package:
<dependency>
<groupId>org.apache.tika</groupId>
<artifactId>tika-parsers-standard-package</artifactId>
<version>4.x.y</version>
<type>pom</type>
</dependency>For Gradle:
dependencies {
implementation(platform("org.apache.tika:tika-bom:4.x.y"))
// version not required since bom (platform in Gradle terms)
implementation("org.apache.tika:tika-parsers-standard-package@pom")
}Projects managing multiple Tika modules should import the BOM to avoid version convergence errors. Add tika-bom with scope import to the dependencyManagement section, then declare tika-parsers-standard-package without a version.
In Java, parsing a file to a string takes three lines:
import org.apache.tika.Tika;
Tika tika = new Tika();
String text = tika.parseToString(new File("document.pdf"));
System.out.println(text);Tika returns the extracted text as a String. For Markdown output, use the 4.x default without the `--text` flag.
Markdown Output and Vision-Language Model Parsers in 4.x
The 4.x default output format is Markdown, not plain text. This matters for LLM ingestion because Markdown preserves structural signals such as headings, lists, and tables that a language model can use when generating a retrieval-augmented answer. Plain-text output remains available with the `--text` flag on the CLI.
The 4.x line also added vision-language model parsers for files that OCR cannot read. It can call external VLM APIs from Claude, Gemini, and OpenAI to describe or extract content from images and scanned documents where character recognition fails. The README identifies this as a capability for documents "OCR can't read." These parsers require API credentials for the corresponding VLM services; they are not bundled capabilities that work offline.
The repository also supports reproducible builds: building the same source code with the same JDK version produces byte-for-byte identical artifacts regardless of build machine or time. The `project.build.outputTimestamp` is set in `tika-parent/pom.xml`. This matters for supply-chain verification in regulated environments.
Where Apache Tika Reaches Its Limits
Tika extracts text and metadata. It does not edit, re-render, or convert documents. A developer who needs to modify the content of a Word document or produce a pixel-accurate PDF rendering of a slide deck needs a different tool.
Format coverage is broad but not complete. Proprietary or binary formats without open specifications may produce empty or garbled output. The README does not provide a complete list of parsers with known failure modes. The authoritative list is on the project website.
The forked-process model adds overhead. Each parse spawns a child JVM. In high-throughput scenarios where many small documents are parsed per second, that startup cost accumulates. The tika-server REST API is the more practical choice for server-side batch work because it amortizes JVM startup across many requests, rather than restarting a process per document.
Upgrading from 3.x to 4.x requires code and configuration changes: the move to Java 17, updated `TikaInputStream` SPI interfaces, JSON configuration replacing tika-config.xml, and namespaced metadata keys. This migration is not a dependency version bump. Teams should allocate time for it rather than treating it as routine maintenance.
Apache Tika vs. Docling for Document Ingestion
Docling is a Python library developed by IBM Research and released as open source. It focuses on converting documents into structured Markdown and JSON with high layout fidelity, particularly for PDF documents. The core difference from Tika is scope: Docling handles a narrower range of formats (PDF, DOCX, PPTX, HTML, images) with deeper structural fidelity for each one, while Tika covers over a thousand formats with a uniform but thinner extraction layer.
For a team working in Python on a corpus that is primarily PDF and Office files, Docling may produce higher-quality structured output for those specific formats. For a Java team, or for a project that must handle unusual or legacy formats such as DjVu, CAD files, or proprietary image formats, Tika's format coverage advantage is more relevant. The search phrase "apache tika vs docling" reflects a real decision engineers face when choosing a document ingestion layer; both tools have narrow cases where they are the stronger fit.
Tika and Docling are not mutually exclusive. Some pipelines use Tika for format detection and initial extraction across a wide corpus, then route high-value documents through Docling for more precise structural analysis.
Maintenance History, Upgrade Path, and Licensing
The last push to the repository was on 2026-09-25. Tika 2.x and Java 8 support reached end of life in April 2025. The current supported lines and their schedules are documented on the Tika Roadmap page on the Apache Confluence wiki. The project follows the Apache Software Foundation release process: signed source releases and binary packages published to ASF distribution infrastructure and mirrored to Maven Central. The repository has no GitHub releases; all versioned artifacts come through Maven Central or tika.apache.org/download.html.
The license is Apache-2.0. The Apache License permits use in both commercial and open-source products without requiring derivative works to be open-source. It requires preservation of the original license text and attribution notices. Tika bundles parsers for many third-party formats; some carry their own dependencies with separate licenses, such as Apache POI for Office formats and Apache PDFBox for PDF. The NOTICE.txt and HEADER.txt files in the repository address the cumulative attribution requirements. A production deployment should verify that the parser dependencies for the specific formats in use carry licenses compatible with the project's own distribution policy.
Editorial conclusion
Engineers building document ingestion pipelines in Java should consider Apache Tika when format breadth and process isolation matter more than zero-dependency footprint. The 4.x line requires Java 17 and a JSON-based configuration format; teams on Java 8 should note that Tika 2.x reached EOL in April 2025. Before adopting it, verify that your target file formats appear in the supported-parsers list at tika.apache.org and run a test parse against the most complex sample documents in your corpus.
Frequently asked questions
What is Apache Tika used for?
Apache Tika is used to detect file types and extract text and metadata from over a thousand different formats, including PDF, PPT, XLS, DOCX, images, audio, video, and archives. Version 4.x is specifically designed for LLM and RAG pipelines, outputting Markdown by default so language models can consume the extracted content directly.
How to install Apache Tika?
Pre-built binaries are available at tika.apache.org/download.html. For the command-line tool, download the tika-app ZIP, unzip it into a directory, and run the jar from inside that directory alongside its lib/ folder. For Java projects, add tika-parsers-standard-package as a Maven or Gradle dependency from Maven Central.
What is the Apache Tika server?
The tika-server is a REST server component included in the repository under tika-server/. It exposes parsing and detection endpoints over HTTP, which makes it more efficient for batch processing than the standalone jar because it avoids per-request JVM startup costs. The /rmeta endpoint returns recursive metadata JSON for documents and their embedded content.
Is Apache Tika open source?
Yes. Apache Tika is released under the Apache-2.0 license, which permits use in both commercial and open-source products. It is a project of the Apache Software Foundation.
Does Apache Tika perform OCR?
Apache Tika 4.x added vision-language model parsers that call external APIs from Claude, Gemini, and OpenAI to extract content from documents that OCR cannot read, such as scanned images. These parsers require external API credentials and do not work offline. For documents with extractable text layers, Tika reads them directly without OCR.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/apache-tika)