Library / SDK
drewnoakes/metadata-extractor avatar
drewnoakes/metadata-extractor

metadata-extractor: reading media metadata in Java without an imaging library

Extracts Exif, IPTC, XMP, ICC and other metadata from image, video and audio files

2,832 stars506 forksJavaApache-2.0

At a glance

What is it?
A Java library that pulls Exif, IPTC, XMP, ICC profiles and camera makernotes out of image, video and audio files, with decoding for eighteen camera makers and coverage from JPEG to MP4.
Who is it for?
metadata-extractor earns its place when the job is reading metadata rather than pixels. You get one Maven artifact, a single static call to get started, and format coverage that spans still images, raw camera files, audio and video containers.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 74 days ago.
What is it written in?
Mainly Java, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 24, 2026, and from our analysis. They are not legal advice.

Editorial analysis

One static call is the entire entry point

The usage example in the README is a single line, and it is genuinely the whole public entry point most applications need:

java
Metadata metadata = ImageMetadataReader.readMetadata(imagePath);

What you get back is a `Metadata` instance, and the README's next sentence is the entire API story: with it you can iterate or query the tag values that were read from the image. Both verbs are linked to pages on the project wiki rather than to anything inside the repository, which tells you where the real documentation lives before you read a line of source.

The parameter is a path string, not an `InputStream`, so the first thing to check in the wiki is what other overloads exist and what types they accept. That distinction decides whether the library drops into an existing pipeline that already has open streams, or whether it needs a file on disk. The getting started guide and the SampleOutput wiki page are both linked from the README, and between them they cover the two questions a first user actually has: how to walk the results, and what a result actually looks like when a camera writes an unusual value.

Sixteen metadata families, many of them in one file

The features list separates two things that readers tend to blur together. The first group is metadata formats the library understands: Exif, IPTC, XMP, JFIF and JFXX, ICC profiles, Photoshop fields, WebP properties, WAV properties, AVI properties, PNG properties, BMP properties, GIF properties, ICO properties, PCX properties, QuickTime properties and MP4 properties. The second is the container files it will process: JPEG, TIFF, WebP, WAV, AVI, PSD, PNG, BMP, GIF, HEIF covering both HEIC and AVIF, ICO, PCX, QuickTime, MP4, and camera raw formats from Nikon, Canon, Olympus, Sony, Panasonic, Leica and Samsung.

The README's own framing is that several of these formats may be present in a single image, which is the accurate description of how Exif, IPTC, XMP and an ICC profile coexist in one JPEG. That overlap is the reason to reach for this library rather than a JPEG decoder: the file is not walked twice, once for image data and once for tags, because image data is not what the library is for.

The audio and video entries deserve a second look. WAV and AVI properties sit in the same list as JPEG and PNG, so a library used as an audio tagger and a library used as a photo tagger are the same dependency. For a system that ingests mixed uploads, that consolidation is the actual argument for adopting it.

Makernote decoding across eighteen camera makers

Below the container list sits the part of the project with the most visible track record: camera-specific makernote decoding for Agfa, Apple, Canon, Casio, Epson, Fujifilm, Kodak, Kyocera, Leica, Minolta, Nikon, Olympus, Panasonic, Pentax, Reconyx, Sanyo, Sigma/Foveon and Sony. Makernotes are vendor-private blocks embedded inside an otherwise standard Exif structure, so each maker needs its own tag dictionary, and each new camera body tends to arrive with fields nobody has seen before.

The repository tree shows where the maintenance pressure sits. The root holds `Source/`, `Tests/`, `Samples/`, `Resources/`, `wiki` and `wiki-images`, plus `pom.xml`, `build.gradle` and `CONTRIBUTING.md`. A separate sample image repository is linked for research and testing, and the README credits contributors who sent in sample images from their own cameras. That division is worth understanding: the library ships descriptors, and the samples that prove they work live elsewhere, which is why a raw format gain usually arrives with a batch of new fixtures.

The build is worth noting too. Both `pom.xml` and `build.gradle` sit at the root, so Maven and Gradle are both first-class consumers, and the Maven Central badge in the README means you are pulling a published artifact rather than building from source.

The repository lists capabilities, the wiki explains the API

Reading the README top to bottom, the documentation split is stark. Installation gets a Maven dependency block and a pointer to the releases page:

xml
<dependency>
    <groupId>com.drewnoakes</groupId>
    <artifactId>metadata-extractor</artifactId>
    <version>2.19.0</version>
</dependency>

Everything after that is a list of formats, a list of files, a list of camera makers, and then links out. Questions and Feedback goes to a Stack Overflow tag, bugs go to the issue tracker with a request to attach sample images because most issues cannot be investigated without one, and contributing asks you to open an issue before writing a pull request. There is no usage chapter, no descriptor explanation, and no changelog in the README itself.

That is a deliberate division rather than an omission, and it is common in libraries that have been running for long enough to accumulate a wiki. The practical consequence for evaluation is that the repository cannot answer the API questions. If you need to know whether directories can be filtered, whether description objects expose a human-readable name, or how a malformed tag is surfaced, the SampleOutput page and the getting started guide are where those answers live, and the issue tracker is the fallback when the wiki has not caught up.

Release notes describe hostile files, not new features

Three releases are visible, and reading them changes how you picture the project. Version 2.19.0, published 2023-11-28, is a mixed bag: new Exif tags, a ClassCastException fix for Apple metadata of unexpected type, a note that a height or width of 0 in an ICO file actually means 256, PNG iTXt encoding fixes, and GitHub Actions added for CI builds.

Version 2.20.0, published 2026-04-08, adds OM System II makernote parsing built on the Olympus II makernote, support for AVIF and AV1 images, parsing of illegal dates, handling of a known null value, and a fix for wrong date and time when the timezone is null. It also validates lengths before parsing JPEG data. Version 2.21.0, published 2026-07-22, fixes an out-of-memory denial of service vector in `BmpHeaderDescriptor.formatHex` caused by an unbounded digits parameter, extracts IPTC metadata from PNG text chunks, tolerates duplicate PNG chunks that are not allowed to appear more than once, and ports error checking from a .NET reader.

Almost none of that is a headline feature. It is a series of length checks, null guards, encoding fixes and untrusted-input hardening, which tells you the library is most of its effort spent on files that are malformed, truncated or actively hostile. If you are processing uploads from strangers, that is the profile you want. The repository was last pushed on 2026-07-28 and is not archived, with 2.21.0 landing six days earlier.

Where this fits against a native imaging library

The honest comparison is not with other metadata libraries but with what you would otherwise reach for. Java's own ImageIO can decode a JPEG and, in some configurations, surface a small amount of metadata, but it does not attempt makernote decoding and it treats the image as the primary object. Choosing ImageIO means you are building a metadata layer yourself, and vendor-private Exif blocks are exactly the part you would want to skip.

There is a second comparison worth making, which the README makes for you. It lists ports and wrappers: a complete .NET port maintained alongside the Java library, a PHP project that wraps this one, and two Clojure projects, one returning a subset and one returning all the data. The Clojure wrapper that returns everything also handles image manipulation and honours the orientation tag on resize, preserving portrait mode, which the README notes ImageIO does not do. The existence of a maintained .NET port alongside the original is a reasonable signal about how settled the core parsing logic is.

What you give up by choosing this library is any attempt at writing metadata back. The README describes the project as a library for reading metadata, lists only read paths, and documents no write API, so if your requirement includes stripping a GPS tag before upload you will need a different tool for that half of the job.

Editorial conclusion

metadata-extractor earns its place when the job is reading metadata rather than pixels. You get one Maven artifact, a single static call to get started, and format coverage that spans still images, raw camera files, audio and video containers. Its honest limit is documentation depth: the README lists capabilities well and explains none of the API, so the GitHub wiki, not the repository, is where you will actually learn to iterate directories and format descriptions. The release notes from 2.19.0 through 2.21.0 read as a list of hostile files handled more carefully, which tells you something useful about how the library behaves on input you did not control. Start from the getting started wiki page and the SampleOutput page, because those two answer the questions the README only gestures at.

Frequently asked questions

Which file formats can metadata-extractor read?

The README lists image formats (JPEG, TIFF, WebP, PNG, BMP, GIF, HEIF including HEIC and AVIF, ICO and PCX), container formats (QuickTime and MP4), audio formats (WAV), and camera raw files from Nikon, Canon, Olympus, Sony, Panasonic, Leica and Samsung. It also parses PSD and AVI properties.

Does metadata-extractor remove or write metadata?

The README describes the project as a Java library for reading metadata, and it documents a read path only: call ImageMetadataReader.readMetadata and then iterate or query the tag values. No write or strip API appears in the README, so a requirement to remove a GPS field before upload needs a different tool.

How do I add metadata-extractor to a Java project?

As a Maven dependency under groupId com.drewnoakes and artifactId metadata-extractor. The repository ships both pom.xml and build.gradle, so Gradle projects are equally supported, and the Maven Central badge links to the published artifact page.

Which cameras have makernote decoding?

The README names eighteen makers: Agfa, Apple, Canon, Casio, Epson, Fujifilm, Kodak, Kyocera, Leica, Minolta, Nikon, Olympus, Panasonic, Pentax, Reconyx, Sanyo, Sigma/Foveon and Sony. Version 2.20.0 added OM System II parsing built on the Olympus II makernote.

Official sources

  1. drewnoakes/metadata-extractor on GitHub
  2. Issues
  3. License: Apache-2.0
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/drewnoakes-metadata-extractor.svg)](https://hysenlabs.com/projects/drewnoakes-metadata-extractor)