Apache POI: H??F and X??F, a lite jar by default, and a GitHub repository that is a mirror
Mirror of Apache POI gitbox. The Java API for Microsoft Documents.
At a glance
- What is it?
- The Java library for reading and writing Microsoft Office binary and OOXML formats, where the class naming tells you which format a component handles and the default artifact deliberately omits most of the generated schema classes. The project description opens by saying this repository is a mirror.
- Who is it for?
- Adopt POI if you have to read or write Office documents from Java and cannot shell out to another tool, because the coverage is broad and the naming convention makes the format boundary visible in every class name. Do not adopt it for high-volume writing, where the streaming spreadsheet path is the one you want, or for anything needing full fidelity, because the OOXML classes are generated from Microsoft's schemas and what is not in the schema is not represented.
- Can I use it commercially?
- Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly Java, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.
Editorial analysis
This repository is a mirror, and the bug tracker is elsewhere
The first fact in the project description is that this is a mirror of the Apache POI gitbox. That single word changes how you use the repository. It is a reading surface, and the place work happens is the gitbox instance that the ASF runs. The bug tracker section confirms it by listing two systems with Bugzilla first, pointing at the Apache Bugzilla product for POI, and GitHub issues second. Mailing lists are split three ways, for developers, for users including announcements, and general. So the support path for a POI user is a Bugzilla ticket and a user mailing list, and the GitHub repository is where you read the code and, if you must, leave an issue that may not be the queue anyone watches. The branch story is the other thing to fix in your head. The README states that Java 11 or later is required and that the trunk branch is used for 6.0.0 development, and that POI 4 and 5 releases require Java 8 or later. That last clause is the one that matters in production, because most code in the world is on POI 4 or 5 and therefore on a Java 8 baseline, while anyone reading trunk sees a Java 11 codebase. There are no GitHub releases here, so the published artefacts are the ones in the Maven repository, and the README points at a third-party artefact index as a good place to find the jars and the dependency definitions.
H??F and X??F, and six applications ranked by maturity
The component list is the clearest piece of API design communication in the project, because it is organised on two axes and the naming encodes both. On the format axis, components named with an H are for reading or writing OLE2 binary formats and components named with an X are for the OOXML formats. On the application axis, each Office application gets a letter pair: SS for spreadsheets, WP for word processing, SL for slideshows. So HSSF and XSSF are the two Excel implementations, HWPF and XWPF are the two Word ones, HSLF and XSLF are PowerPoint. Reading a class name tells you which format it touches, which is the question you always have when a method behaves differently on a .doc and a .docx. The second axis is maturity, and the README ranks the applications in descending order, with Excel spreadsheets first, then PowerPoint, then Word, then Outlook, then Visio, and Publisher last. It is also candid in the body text: the common high-level API is most developed for Excel workbooks, and work is progressing for Word documents and PowerPoint presentations. That ordering is a practical warning. If your requirement is spreadsheets you are on solid ground; if it is Word documents you should find out early how much of what you need is implemented, because the two implementations of Word are uneven. The Outlook entry is worth a separate note, since the project says Microsoft opened the specifications to that format in October 2007 and that it would welcome contributions, which after nearly two decades tells you how slowly that path moves.
poi-ooxml-lite is the default, and knowing that early saves a dependency change
The jar list contains one decision that will affect you more than any API detail, and it is explained clearly enough that there is no excuse for missing it. Four jars are named for the OOXML support, and they are not alternatives of equal size. poi-ooxml provides the support for the newer XML formats. poi-ooxml-lite is described as generated classes based on the Microsoft schemas used by poi-ooxml, and it only includes the most commonly used classes. poi-ooxml-full is described as generated classes based on the same schemas, and it can be used instead of lite if you need support for less commonly used features. So the default artifact is a subset. The reason is volume, and it is the right one: the Office schemas are enormous, and generating every class from them would put a large amount of dead code in every application that reads a spreadsheet. The consequence for you is a specific failure mode. You write against a class, it compiles against lite, and one day you touch a less common part of the document model and the class is not there, and the fix is to change the artifact. That is not a configuration tweak, it is a dependency change, and it is much cheaper to make before you have a hundred call sites. There is a poi-ooxml-lite-agent directory in the tree, which is how the lite set is maintained when the schemas move, so this is a maintained process rather than a one-off trimming. Decide which artifact you want on day one. The remaining jars are simpler: the main jar carries shared interfaces, poi-scratchpad adds the legacy binary format classes, and poi-excelant is for driving Excel from Ant scripts.
Java 11, Java 8 and JDK 17 in one README, for three different readers
Three Java versions appear in this README and each belongs to a different audience, which is why the numbers seem to contradict each other. Consumers of the current development line need Java 11 or later, and that is the trunk branch used for 6.0.0 development. Consumers of the production releases POI 4 and POI 5 need Java 8 or later. Contributors are asked to install a JDK of 17 or later, together with Apache Ant 1.8 or later, or Gradle. So the number you need depends entirely on which hat you are wearing, and the most common mistake is reading the contributor line and concluding that Java 17 is needed to use the library, which it is not. The build system is described in a similarly unhurried way. The contributing section offers Ant or Gradle, and the tree contains both a build.gradle with settings.gradle and a wrapper, and a build.xml with a patch file, so the migration to Gradle is complete but not exclusive. The build command for the jars is given as
./gradlew jarand the README then repeats the same command without the leading path separator on the next line, which is a copy and paste artefact rather than a second method. Note also that the contributing instructions are unusually specific about where things live, listing the test directories for the main jar, the OOXML module and the scratchpad module, plus a test-data directory of fixtures, and the matching source directories. For a project of this size, that map is the difference between a first patch and an afternoon of hunting.
A fuzzing module and a file leak detector, which are the two directories that matter
Two directories in the top-level listing tell you what this project's threat model is, and both are unusual to find. The first is a fuzzing module. A library whose job is to parse file formats written by other software, from files that arrive from strangers, is exactly the kind of parser that needs continuous fuzzing, and having a dedicated module for it rather than a side project is the difference between a practice and an intention. The second is a file-leak detector, with an exclude file next to it. That is a subtler signal and just as revealing. A document parser opens a lot of files, and a long-running service that parses thousands of documents will exhaust its file handle limit if any path forgets to close. A leak detector with an exclusion list is how you keep that honest, and the exclusion file is the interesting part, because it is a running record of the places where a leak is known and accepted. Read that file before you run POI inside a long-lived service, and treat anything in it as a question you need answered about your own workload. Both of these sit alongside a benchmark module for the core library and another for the OOXML module, so the project also measures what it does. The last structural item is a security policy file, which tells you where to report a problem in a library that parses untrusted input.
OOXML fidelity is bounded by the schemas, and Excel is where the effort went
The limitations here are structural rather than accidental, and they follow from two facts already established. The first is that the OOXML classes are generated from Microsoft's schemas. That is the right way to get broad coverage cheaply, and it also means anything the schema does not describe is not in the library, and anything the schema describes awkwardly is awkward in the API. The second is the maturity ordering, which puts Excel first and Word behind it. Together they define where the library is strong and where it will cost you. For reading and writing spreadsheets, including the streaming path that keeps memory flat when a workbook is larger than memory, this is the mature answer in the Java ecosystem and there is no serious alternative that avoids a native dependency. For Word documents, expect to write more code and to check the implementation status of the specific class you need before designing around it. For Visio, Publisher and Outlook, treat the API as partial and read the source. And there is a broader point about the formats themselves. The project exists to read and write formats that Microsoft defined for Office 97 through 2008, which is a large installed base and a stable target, but it is not a general document manipulation library, and a document model built on generated schema classes will not express the things a word processor does that the schema does not describe.
Where POI is the wrong choice
Three cases, stated plainly. If your problem is converting or extracting at volume, POI is a library, not a pipeline, and you will be writing the concurrency, the memory management and the retry logic yourself; the streaming spreadsheet writer is the part designed for large output, and everything above it is your code. If your problem is fidelity, meaning a document that must round-trip through another implementation without loss, generated schema classes and partial application coverage are the wrong foundation, and the honest answer is that no Java library will do it. If your problem is anything that is not a Microsoft Office format, this library has nothing to offer, and its low-level components exist to support the Office formats rather than as general file format tools. On the positive side, the two questions that decide it are both answerable from the README. If your documents are spreadsheets and you are on Java 8 or 11 or later, the answer is straightforward. If they are Word documents, check the maturity statement first, because the project tells you work is still in progress there, and then check whether the lite artifact covers the classes you need before you commit to the dependency.
Editorial conclusion
Adopt POI if you have to read or write Office documents from Java and cannot shell out to another tool, because the coverage is broad and the naming convention makes the format boundary visible in every class name. Do not adopt it for high-volume writing, where the streaming spreadsheet path is the one you want, or for anything needing full fidelity, because the OOXML classes are generated from Microsoft's schemas and what is not in the schema is not represented. Three things to check first. Decide between poi-ooxml-lite and poi-ooxml-full before you build, since the default omits less common classes and discovering that later means changing your dependency. Read the maturity ordering, because Excel is the most developed area and Word is described as work in progress. And note that this GitHub repository is a mirror of the gitbox source and that Bugzilla is listed first under bug trackers, so a GitHub issue may not reach the people who triage them.
Frequently asked questions
What is the relationship between the GitHub apache/poi repository and the Apache POI source?
The project description states that this repository is a mirror of the Apache POI gitbox. The bug tracker section lists Bugzilla first and GitHub issues second, and mailing lists exist for developers, users and general discussion.
Which Java version does Apache POI require?
Java 11 or later on the trunk branch, which is used for 6.0.0 development, and Java 8 or later for the POI 4 and POI 5 releases. Contributors are asked to install a JDK of 17 or later, plus Apache Ant 1.8 or later, or Gradle.
What is the difference between poi-ooxml-lite and poi-ooxml-full?
Both are generated from the Microsoft schemas. poi-ooxml-lite includes only the most commonly used classes and is the smaller default, while poi-ooxml-full includes the full generated set and is the one to use when you need less commonly used features.
What do the H and X in POI component names mean?
Components named with an H are for reading or writing OLE2 binary formats and components named with an X are for the OOXML formats. The letters then name the application, so HSSF and XSSF are the two Excel implementations, HWPF and XWPF the Word ones, and HSLF and XSLF the PowerPoint ones.
Which Office formats does Apache POI support most completely?
Excel spreadsheets, which the README lists first in descending order of maturity and describes as the most developed high-level API. Word documents and PowerPoint presentations are described as work in progress, and Outlook, Visio and Publisher follow behind them.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/apache-poi)