# SwiftSoup: a jsoup-style HTML parser written in pure Swift

> SwiftSoup brings jsoup's DOM, CSS selector and jQuery-style API to Swift, covering Linux, iOS, macOS, tvOS and watchOS. It is a good fit when you need to parse or sanitize HTML inside a Swift process, and the wrong tool when you want a browser engine or a full scraping framework.

**scinfu/SwiftSoup** — SwiftSoup: Pure Swift HTML Parser, with best of DOM, CSS, and jquery (Supports Linux, iOS, Mac, tvOS, watchOS)

- Repository: https://github.com/scinfu/SwiftSoup
- Website: https://scinfu.github.io/SwiftSoup/
- Stars: 5,127 · Forks: 404
- Language: Swift
- License: MIT
- Published: 2026-09-23 · Updated: 2026-09-23 · Language: en
- Canonical page: https://hysenlabs.com/projects/scinfu-swiftsoup

## The gap SwiftSoup fills for Swift projects

Java developers have had jsoup for years: a parser that accepts messy HTML, builds a browser-like document tree, and lets you query it with CSS selectors. Swift had no equivalent, so teams writing an iOS or server-side Swift app that needed to read a page had to either bridge to a C library, shell out to another language, or hand-roll string matching. SwiftSoup is the port that closes that gap. It is a pure Swift library, so it carries no C dependency and no separate runtime, and the README lists macOS, iOS, tvOS, watchOS and Linux as supported platforms.

The audience is narrow but real. If you are building a reader mode, an Open Graph or meta-tag extractor, a feed scraper, a link checker, or a comment sanitizer, and your code already runs in Swift, the alternative is usually a subprocess or a second service. SwiftSoup keeps that work in-process. The README states that the library conforms to the WHATWG HTML5 specification, which is the part that matters most: it means the parse tree is shaped the way a browser would shape it, so malformed markup does not silently produce a different tree than the one your selectors were written against.

## How the parse tree and selector engine work

The entry point is a static parse function that returns a Document. From there, everything is DOM traversal: select returns an Elements collection, first narrows to one element, text returns the rendered text of a node, and attr reads a single attribute by name. The README describes the API as combining DOM traversal, CSS selectors and jQuery-like methods, and the examples bear that out: you can query with a compound selector such as p.message, then walk the resulting elements in a loop.

The interesting design decision is parser auto-detection. SwiftSoup.parse inspects the start of the input for an XML declaration. If it finds one, the XML parser is used; otherwise the HTML parser runs. The README frames this as making feeds and OPML documents work without extra configuration, and gives an OPML example where link and img elements are selected and read. That convenience has a sharp edge: an HTML document that happens to begin with a stray XML declaration will be routed to the XML parser, where HTML5 tag normalization does not apply. When that matters, the README documents explicit escape hatches. parseXML forces XML parsing, parseHTML forces HTML5 rules even when the input carries an XML declaration, and the older form of passing Parser.xmlParser() as an argument still exists.

Modification follows the same tree. Elements expose append, and the README example parses an empty div, appends a paragraph, and prints the resulting document, which comes back wrapped in html, head and body elements. That wrapping is a direct consequence of the HTML5 parsing rules, not a quirk, and it is worth knowing before you diff generated output against your input.

## Installing SwiftSoup and parsing your first document

The README documents three package managers. For Swift Package Manager, add the dependency to your Package.swift and name SwiftSoup in the target's dependencies. The README uses from: "2.6.0" as the version requirement, so adjust that to whatever version you have verified. CocoaPods users add a single pod line to the Podfile, and Carthage users add a github line to the Cartfile.

```swift
dependencies: [
    .package(url: "https://github.com/scinfu/SwiftSoup.git", from: "2.6.0"),
],
targets: [
    .target( name: "YourTarget", dependencies: ["SwiftSoup"]),
]
```

After the dependency resolves, import the module and parse a string. The README's first example builds a small document and reads its title, which is the fastest way to confirm the package linked correctly.

```swift
import SwiftSoup

let html = """
<html><head><title>Example</title></head>
<body><p>Hello, SwiftSoup!</p></body></html>
"""

let document: Document = try SwiftSoup.parse(html)
print(try document.title()) // Output: Example
```

The realistic second step is selector-based extraction. The README's message example parses two paragraphs that share a class and prints their text in document order.

```swift
let messages = try document.select("p.message")

for message in messages {
    print(try message.text())
}
```

If your source is a live page rather than a string, the README shows parsing directly from a URL with SwiftSoup.parse(url). It adds one caveat worth repeating: when Foundation cannot determine a page's text encoding, avoid String(contentsOf:) and parse the raw response bytes instead.

## Sanitizing untrusted HTML with a whitelist

The clean function is the part of SwiftSoup that has the least to do with scraping and the most to do with security. It takes a dirty HTML string and a Whitelist, and returns markup with everything outside the whitelist removed. The README's example feeds in a script tag next to bold text and gets back only the bold text, which is the behaviour you want when rendering user-submitted content in a web view.

The whitelist is configurable rather than fixed. You can start from a built-in such as Whitelist.basic(), or build one by chaining addTags, addAttributes and addCSSProperties. The README's second sanitizing example allows the p tag, permits a style attribute on it, and then restricts which CSS properties survive: color passes through, position:absolute is dropped. That property-level filtering is the detail that makes the API useful, because allowing a style attribute wholesale is a well-known way to reintroduce layout-based attacks.

The limitation is that the README does not document the contents of the built-in whitelists. Whitelist.basic() is named but its tag and attribute set is not listed, so if you rely on it you are trusting a definition you cannot see in the documentation. For anything security-sensitive, build an explicit whitelist and test it against your own payloads rather than assuming basic() matches your policy.

## Where SwiftSoup is the wrong choice

SwiftSoup parses markup. It does not execute JavaScript, and the README makes no claim that it does, so a page whose content is rendered client-side will look empty to it. If your target is a single-page application, a parser is the wrong layer and you need a browser engine.

The second boundary is scale and network behaviour. SwiftSoup can parse a URL directly, which is convenient for one page, but the README documents no connection pooling, retry policy, rate limiting, cookie handling or proxy support. Building a crawler on top of that API means writing all of it yourself. For a handful of pages that is fine; for a crawl of any size, a dedicated scraping stack is the better foundation.

The third boundary is XML. Auto-detection only triggers on an XML declaration at the start of the content. A document without that declaration is parsed as HTML5, which normalizes tags and inserts html, head and body wrappers. If you are processing a format where the wrapper elements would corrupt the output, you have to call parseXML explicitly, and you have to remember to do it. The README does not document what happens with malformed XML beyond the parser choice itself.

## SwiftSoup against Kanna and a Python scraping stack

Kanna is the other Swift HTML/XML parser that appears in searches around this project, and the difference is in the API surface rather than the language. Kanna wraps libxml2 and pairs with an XPath-first query style, while SwiftSoup implements its own parser and presents a jsoup-shaped API built around CSS selectors, DOM mutation and whitelist sanitizing. The practical consequences: SwiftSoup has no C dependency to link, and its clean function gives you an out-of-the-box sanitizer that an XPath wrapper does not provide. The trade-off runs the other way for XML-heavy work, where libxml2's XML handling is the more established path.

Against a Python stack such as BeautifulSoup, the split is architectural. BeautifulSoup is a parser plus tree traversal that usually sits inside a larger Python scraping ecosystem, and it is the tool most people reach for when the parsing is one step in a pipeline with requests, queues and storage. SwiftSoup exists so that step can happen inside a Swift binary instead. If your project is already Python, there is no reason to move; if your project ships as an iOS app or a Swift server, moving the parse out of process is the cost you avoid.

## Maintenance, upgrade cost and the MIT licence

The repository is not archived, and the last push was on 2026-09-22. Releases are frequent and versioned in the 2.13.x line, with 2.13.9 published on 2026-08-27. That cadence suggests fixes and small additions rather than a rewrite in progress, but the README does not publish a deprecation policy or a support window, so upgrading is a matter of reading CHANGELOG.md, which is present at the repository root, before moving a version constraint forward.

The version table in the README is the constraint most likely to bite. Swift 5 maps to SwiftSoup 2.0.0 and later, while Swift 4.2 maps to 1.7.4. If you are pinned to an older toolchain, you are also pinned to an old release, and the newer parser features documented in the README, including the auto-detection and the explicit parse modes, may not exist in the version you can use.

The licence is MIT, which is permissive and places few conditions on redistribution. That is a description of the licence identifier, not legal advice; if you are shipping a closed-source product, have your own counsel confirm the notice requirements.

## Conclusion

Adopt SwiftSoup when HTML parsing, CSS-selector extraction or whitelist sanitizing has to happen inside a Swift target on Apple platforms or Linux, and the input is a string, file or URL you control. Do not adopt it if you need a JavaScript-capable browser engine, a headless crawler with proxy rotation, or CSS selector support beyond what the README shows. Before committing, verify that your toolchain matches the Swift version table in the README, that your Package.swift dependency resolves, and that your own fixture HTML produces the selector results you expect.

## FAQ

### What does an HTML parser do, and where does SwiftSoup fit?

An HTML parser turns markup into a structured document tree. SwiftSoup does this in pure Swift, and the README states that it conforms to the WHATWG HTML5 specification, so the tree is shaped the way modern browsers shape it.

### Can BeautifulSoup be used to parse HTML, and how does that compare with SwiftSoup?

BeautifulSoup is a Python HTML parsing library, so it is usable only from Python. SwiftSoup is a pure Swift library for macOS, iOS, tvOS, watchOS and Linux, which means the parse can happen inside a Swift binary instead of a separate Python process.

### How can I parse HTML in Python, and is SwiftSoup relevant?

Parsing HTML in Python uses Python libraries, not SwiftSoup. SwiftSoup is written in Swift, so it is relevant when the parsing needs to run in a Swift target rather than a Python one.

## Sources

- [License: MIT](https://github.com/scinfu/SwiftSoup/blob/master/LICENSE)
- [Project website](https://scinfu.github.io/SwiftSoup/)
- [README](https://github.com/scinfu/SwiftSoup/blob/master/README.md)
- [Releases](https://github.com/scinfu/SwiftSoup/releases)
- [scinfu/SwiftSoup on GitHub](https://github.com/scinfu/SwiftSoup)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/scinfu-swiftsoup
