Library / SDK
jhy/jsoup avatar
jhy/jsoup

jsoup: the Java HTML parser that builds the same DOM a browser does

jsoup: the Java HTML parser, built for HTML editing, cleaning, scraping, and XSS safety.

11,398 stars2,299 forksJavaMIT

At a glance

What is it?
A mature Java library for real-world HTML and XML, covering URL fetching, CSS and XPath selection, DOM manipulation and safelist-based cleaning, with the WHATWG HTML5 spec as its target.
Who is it for?
jsoup is the rare library where the boring answer is the right one: it implements a published specification, it has been doing that for seventeen years under one primary author, and it does not ask you to accept a private idea of what a document is. The 1.23.1 parser rewrite is the most interesting thing in the recent history, since a faster parser that also holds 58 to 65 percent less memory on source-tracked documents changes what people build on top of it.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 15 days ago.
What is it written in?
Mainly Java, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 21, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The claim that matters is spec conformance, not convenience

Most HTML parsers make a promise about being lenient. jsoup makes a different promise, stated in the second line of the README: it implements the WHATWG HTML5 specification and parses HTML to the same DOM modern browsers produce. That is a much stronger commitment, and it is the reason the library has accumulated 11,398 stars, 2,299 forks and a topic list that includes css-selectors, dom, web-scraping, xpath and java-html-parser.

The specification claim explains the design. Browsers do not fail on malformed markup, they recover, and they recover in a defined way. A scraper that wants to select `#mp-itn b a` and get the same elements a browser would needs those same recovery rules, including the implied elements and the table fixups that no ad-hoc parser bothers with. jsoup describes itself as designed to deal with all varieties of HTML found in the wild, from pristine and validating to invalid tag-soup, and to produce a sensible parse tree either way.

The project is MIT licensed, not archived, and the last push was 2026-09-21. Open issues number four, which for a project this size is a statement about how the issue tracker is used rather than about how busy the maintainer is.

A five-line example that shows the whole API shape

The README's example fetches a page, parses it and selects by CSS selector in one call chain:

java
Document doc = Jsoup.connect("https://en.wikipedia.org/").get();
log(doc.title());
Elements newsHeadlines = doc.select("#mp-itn b a");
for (Element headline : newsHeadlines) {
    log("%s\n\t%s",
        headline.attr("title"), headline.absUrl("href"));
}

Four things are worth noticing. `Jsoup.connect(...).get()` combines HTTP and parsing, so the common case needs no separate client. `doc.select` takes a CSS selector, and the README's feature list says xpath is available as well. `Elements` is a list-like collection you can loop over directly. And `absUrl("href")` resolves a relative link against the document's base URI, which is the sort of detail that silently breaks naive scrapers.

`log` is not part of the library. The snippet is a trimmed version of `src/main/java/org/jsoup/examples/Wikipedia.java`, and the README links both an online sample and the full source so you can see the shape with the logging filled in.

The Cleaner is the feature people underuse

Of the five bullets in the README's feature list, four are about getting data out of a document. The fifth is about what you let back in: cleaning user-submitted content against a safelist to prevent XSS attacks.

That is the piece with real security weight, and it is worth being precise about what it does. A safelist approach means you declare the tags and attributes that survive, and everything else is removed or escaped. The alternative, blacklisting dangerous tags, is the approach that fails open every time somebody finds a new parsing quirk between your sanitizer and the browser that will render the output. Since jsoup parses to the same DOM a browser builds, a safelist checked against that DOM is actually reasoning about the same structure the browser will see, which is the property you want and rarely get.

The 1.23.1 release fixed a security issue in the Cleaner that could expose markup when malformed HTML was cleaned. That is the second time in a year the parser work and the sanitizer work have met, and it is a concrete argument for tracking releases rather than pinning an old version and forgetting about it.

1.23.1 rebuilt the parser and changed the memory profile

The release notes for 1.23.1 are the most substantive thing in the recent history, and the performance numbers in them are specific enough to be worth quoting rather than paraphrasing.

The project reports that in its OpenJDK 21 benchmarks, ordinary string parsing is 18 percent faster on average, parsing from an `InputStream` is 11 percent faster, and parsing with source position tracking is 70 percent faster while allocating 64 percent fewer bytes per document. Source-tracked DOMs retain 58 to 65 percent less memory on representative medium-to-large documents. The notes are careful to add that the improvements hold under concurrent parsing without introducing new contention, and that exact gains vary with document shape, JVM and hardware.

That hedge is appropriate and worth repeating: those are the maintainer's own numbers on their own machine, so treat them as an indication of direction rather than a promise for your workload.

The same release brought standard-alignment fixes in `noscript`, CDATA, SVG and MathML parsing, specification-correct HTTP redirects, an immutable `Element#classList()`, and direct outer-HTML output to an `Appendable`. The 1.23.2 patch that followed is narrower: closer alignment with the HTML, XML, URL and form submission specifications, better XML and W3C DOM conversion, streamed request bodies and broader `HttpClient` reuse for HTTP workloads, plus new node insertion methods for `Elements`.

Dependency coordinates that disagree with each other

The getting-started section gives both Maven and Gradle snippets, and they name different versions. Maven gets 1.23.2:

xml
<dependency>
  <!-- jsoup HTML parser library @ https://jsoup.org/ -->
  <groupId>org.jsoup</groupId>
  <artifactId>jsoup</artifactId>
  <version>1.23.2</version>
</dependency>

Gradle gets 1.23.1:

groovy
// jsoup HTML parser library @ https://jsoup.org/
implementation 'org.jsoup:jsoup:1.23.1'

Both are in the README, and both are correct in the sense that both versions have been released. 1.23.2 was published on 2026-08-26 and 1.23.1 on 2026-07-30. The download link in step one of getting started also points at 1.23.1. So the README's dependency snippets, its download link and its own advice to take the latest jar are three different pointers, and the Maven snippet is the one that has moved.

It is a small inconsistency and the fix is easy, which is to copy the version from the releases page rather than from the README. Worth noting precisely because it is the kind of drift that only matters until the moment you are copying a snippet from a blog post written six months ago.

Android support, project structure and the maintenance question

There is a dedicated Android note and it is not filler. When jsoup is used in Android projects, the README says core library desugaring with the NIO specification should be enabled to support Java 8 and later features. That is a real configuration step that a build will otherwise discover for you at compile time.

The repository itself is small in the way mature Java projects tend to be. At the root sit `pom.xml`, `LICENSE`, `SECURITY.md`, `CHANGES.md`, `change-archive.txt` and the READMEs. Everything else is `src/`. The existence of a `change-archive.txt` alongside `CHANGES.md` suggests the project keeps a long history rather than truncating it, and `SECURITY.md` is a file you want to see in a parser that gets pointed at untrusted input.

On the human side, the README names Jonathan Hedley as the creator and primary maintainer and then adds a sentence that is easy to skim past: many contributors have helped improve it over the years, and the contribution graph is on GitHub. That combination, one long-tenured primary author plus accumulated contributions, is a reasonable description of how a seventeen-year-old library stays coherent. The project states its own status plainly as in general, stable release.

One last piece of practical advice from the README itself: for research or technical documentation, cite it as Jonathan Hedley and jsoup contributors, which is the BibTeX block the README includes. An HTML parser that specifies a DOM everyone else implements is a reasonable thing to cite, and the project asks you to.

Editorial conclusion

jsoup is the rare library where the boring answer is the right one: it implements a published specification, it has been doing that for seventeen years under one primary author, and it does not ask you to accept a private idea of what a document is. The 1.23.1 parser rewrite is the most interesting thing in the recent history, since a faster parser that also holds 58 to 65 percent less memory on source-tracked documents changes what people build on top of it. If you are new to it, start with the cookbook introduction rather than the API, and if you clean untrusted HTML, read the 1.23.1 Cleaner advisory before assuming an older pinned version is fine.

Frequently asked questions

What is jsoup used for?

Four things the README lists: scraping and parsing HTML from a URL, file or string, finding and extracting data with DOM traversal or CSS and xpath selectors, manipulating elements, attributes and text, and cleaning user-submitted content against a safelist to prevent XSS. The fifth entry is producing tidy HTML on output.

What is the current version of jsoup?

1.23.2, published 2026-08-26, following 1.23.1 on 2026-07-30. The README's Maven snippet names 1.23.2 while its Gradle snippet and its download link still say 1.23.1, so take the version from the releases page rather than from the README.

What is the best HTML parser?

It depends on what you need from the parse. jsoup's argument for itself is conformance: it implements the WHATWG HTML5 specification and produces the same DOM modern browsers do, which is what you want when your selectors have to agree with what a reader's browser renders. A parser that is more permissive but recovers markup differently will disagree with the browser at exactly the points where your scraper is fragile.

How can I parse HTML in Java?

With jsoup, in the shape `Jsoup.connect("https://en.wikipedia.org/").get()`, which fetches and parses in one call and returns a Document. You can also parse a string or a file. Elements are then selected with CSS selectors through `doc.select(...)`, and the README's feature list notes xpath is available as an alternative.

Official sources

  1. jhy/jsoup on GitHub
  2. License: MIT
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/jhy-jsoup.svg)](https://hysenlabs.com/projects/jhy-jsoup)