Library / SDK
ezyang/htmlpurifier avatar
ezyang/htmlpurifier

HTML Purifier for PHP: a whitelist HTML filter for untrusted markup

Standards compliant HTML filter written in PHP

3,356 stars361 forksPHPLGPL-2.1

At a glance

What is it?
HTML Purifier parses untrusted HTML against a whitelist and rebuilds it, keeping CSS and a full tag set. It suits rich user content in PHP; it is the wrong tool for plain-text escaping.
Who is it for?
Adopt HTML Purifier when PHP code stores or renders rich HTML from users, editors or imported feeds and you need a whitelist plus CSS support rather than a tag stripper. Do not adopt it for plain-text fields, where htmlspecialchars() is enough, or for non-PHP stacks.
Can I use it commercially?
Yes, with conditions. LGPL-2.1 is a weak copyleft licence: you can use it inside commercial and closed-source software, but if you distribute changes to its own files, you must publish those changes under the same licence.
Is it still maintained?
Yes. The repository last received commits 15 days ago.
What is it written in?
Mainly PHP, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem HTML Purifier solves, and who it is for

Any PHP application that accepts formatted HTML from users has the same problem: the markup arrives from a source you do not control, and it is rendered later in a browser. Stripping tags with a regex or a simple sanitizer removes the obvious script elements but usually leaves attribute-based vectors and malformed markup behind. HTML Purifier takes the opposite route. The README describes it as an HTML filtering solution that uses whitelists and aggressive parsing so that XSS attacks are thwarted and the resulting HTML is standards compliant.

The intended audience is narrow and stated plainly. The README says the library is oriented towards richly formatted documents from untrusted sources that require CSS and a full tag-set. If your input is a comment box that should only accept plain text, HTML Purifier is more machinery than the job needs. If your input is a WYSIWYG editor that produces paragraphs, lists, links, tables and inline styles, the whitelist model is what keeps the output predictable.

The README also admits the cost: the library can be configured to accept a more restrictive set of tags, but it won't be as efficient as more bare-bones parsers. It will, however, do the job right, which may be more important. That sentence is the whole trade-off in one place, and it should be read before anything else.

How the whitelist parser rebuilds markup instead of cleaning it

The mechanism described in the README is a combination of whitelists and aggressive parsing. Rather than scanning for known-bad patterns and removing them, the library decides what is allowed and reconstructs the document from that decision. Anything outside the allowed set does not survive the round trip, which means a new browser quirk or a novel attribute trick does not need a new blacklist entry to be neutralised.

The repository layout reflects that design. library/ holds the implementation, configdoc/ generates documentation for the configuration surface, and docs/ carries developer-oriented documentation and code examples. There is a separate WYSIWYG file at the top level for editors like TinyMCE and FCKeditor, and benchmarks/ and smoketests/ sit alongside tests/ as distinct directories. The presence of a configdoc generator matters in practice: the configuration surface is large enough that the project treats its own documentation as generated output rather than a hand-written page.

CSS is part of the scope, not an afterthought. The README says richly formatted documents requiring CSS and a full tag-set are the target, so the filter has to reason about style attributes and stylesheet content as well as tags and attributes. That is a larger job than tag stripping and it is where most of the parsing cost comes from.

Installing HTML Purifier with Composer and a first filter call

The README points to INSTALL for a quick installation guide and to docs/ for an in-depth installation guide with code examples. The package itself is on Packagist, and the README gives one command for Composer users. Run it from the root of the project that will use the library.

bash
composer require ezyang/htmlpurifier

After that, the library is available through Composer's autoloader. The README does not print a full usage example in the section shown here; it directs readers to docs/ for code examples, so the exact configuration array should be taken from there rather than guessed. What the README does make clear is the shape of the work: you configure the filter, you pass untrusted HTML in, and you get filtered HTML out. The configuration object is the part worth reading up on, because it is where tag sets, CSS handling and other behaviour are decided.

If you prefer to work inside the project's own environment, the repository ships a Dockerfile and a docker-compose.yaml. The compose file builds a service named htmlpurifier from the local Dockerfile, names the container htmlpurifier, and mounts the repository at /opt/htmlpurifier. The Dockerfile is based on ubuntu:24.04, sets PHP_VERSION to 8.4, installs php8.4 with the dev, xdebug, iconv, bcmath, tidy and xml extensions, and copies the Composer binary in. That container is aimed at developing and testing the library itself, not at being a production image for your application.

bash
docker compose up -d

The container starts with tty enabled and the repository mounted, so changes on the host appear inside it. Expect a development shell rather than a tuned runtime.

Where HTML Purifier is the wrong tool

The clearest limitation is stated by the project itself: a more restrictive configuration will not be as efficient as more bare-bones parsers. Parsing a document, validating it against a whitelist and rebuilding it costs more than scanning for angle brackets. On a high-traffic endpoint that filters every request, that difference is real, and the README's answer is that doing the job right may matter more. That is a defensible position, but it is a position, not a free lunch.

The second boundary is scope. HTML Purifier filters HTML. If the field you are handling is supposed to contain plain text, filtering it through an HTML parser adds cost and a configuration surface for no benefit; escaping on output is the smaller, more predictable operation. The README's own framing, richly formatted documents requiring CSS and a full tag-set, is the test to apply. If your content does not need CSS or a wide tag set, you are paying for capacity you will not use.

The third boundary is the platform. This is a PHP library. Applications in other languages need a different implementation, and the repository does not present itself as anything else. The related search phrases around JavaScript point at a question this project does not answer.

Finally, configuration is where mistakes happen. A filter is only as strict as the whitelist it is given, so a permissive configuration weakens the guarantee. The configdoc directory exists because the configuration surface is broad enough to need generated documentation; treat it as the reference rather than copying configuration arrays from forum posts.

Alternatives and the difference in approach

The natural comparison is with the escaping functions built into PHP. htmlspecialchars() converts characters that have meaning in HTML into entities, so a string that was meant to be text is rendered as text. It does not parse, it does not keep any tags, and it has no configuration. That is the right tool when the content is plain text and the only requirement is that the browser does not interpret it as markup.

HTML Purifier sits at the other end. It is for content that is supposed to contain markup and should keep most of it, minus anything outside the whitelist. The difference is not strictness, it is intent: escaping says the whole string is text, filtering says the string is a document and only part of it is acceptable. Choosing between them is a question about the field, not about which library is better.

There is also the class of bare-bones sanitizers the README alludes to when it says a restrictive configuration won't be as efficient as simpler parsers. Those tools trade completeness for speed. HTML Purifier's answer is that a whitelist plus aggressive parsing produces standards-compliant output, and that the extra work is the point. If your throughput budget cannot absorb a full parse per request, the honest options are to filter less often, for example at write time rather than on every read, or to accept a weaker guarantee from a lighter tool.

Maintenance, licence and the cost of upgrading

The repository is not archived, and the last push was on 2026-09-15, the same day as the v4.19.1 release. The release history shows v4.19.0 on 2025-10-17 and v4.18.0 on 2024-11-01, so the cadence is irregular: a release roughly once a year, with a patch release following the most recent minor. That pattern matters for planning. You are not tracking a fast-moving dependency, but you should not expect frequent small updates either.

The upgrade cost is concentrated in the configuration surface. Because the allowed tag set, attribute handling and CSS behaviour are configured rather than hard-coded, a version bump can change defaults or the set of recognised keys. The repository ships a configdoc/ directory and a configdoc generation step, which is the mechanism for diffing what your configuration means between versions. A NEWS file at the top level carries release notes, and it is the first thing to read before bumping the constraint in composer.json.

The licence is LGPL-2.1, as stated in the repository's LICENSE file. For PHP applications that depend on the library through Composer, this is a library-level licence and the usual obligation is to preserve notices and to allow relinking; if you modify the library itself and distribute it, different terms apply. That is a summary of the identifier, not legal advice, and the LICENSE file plus your own counsel are the sources that matter.

Editorial conclusion

Adopt HTML Purifier when PHP code stores or renders rich HTML from users, editors or imported feeds and you need a whitelist plus CSS support rather than a tag stripper. Do not adopt it for plain-text fields, where htmlspecialchars() is enough, or for non-PHP stacks. Before shipping, read INSTALL and docs/, then check the configdoc output for the exact keys and defaults your version ships.

Frequently asked questions

How do I install HTML Purifier in a PHP project?

The README says the package is available on Composer and gives composer require ezyang/htmlpurifier as the command. It also points to INSTALL for a quick installation guide and to docs/ for an in-depth guide with code examples.

What is HTML Purifier and what is it used for?

It is an HTML filtering solution written in PHP that uses whitelists and aggressive parsing so that XSS attacks are thwarted and the output is standards compliant. The README says it targets richly formatted documents from untrusted sources that need CSS and a full tag set.

Is HTML Purifier available as a Composer package?

Yes. The README states the package is available on Composer and shows composer require ezyang/htmlpurifier for projects that manage dependencies with it.

Can HTML Purifier be used outside a framework, as standalone PHP?

The README describes it as a PHP library installed through Composer, and the repository ships a Dockerfile and docker-compose.yaml for its own development environment. Nothing in the README ties it to a specific framework.

Does HTML Purifier keep CSS and a full set of tags?

The README says it is oriented towards richly formatted documents that require CSS and a full tag-set, and that it can be configured to accept a more restrictive set of tags. It notes that a more restrictive configuration will not be as efficient as more bare-bones parsers.

Is HTML Purifier as fast as a simple sanitizer?

The README states that a more restrictive configuration won't be as efficient as more bare-bones parsers, and that the library will do the job right, which may be more important. It does not publish benchmark figures in the README section shown here.

Official sources

  1. ezyang/htmlpurifier on GitHub
  2. License: LGPL-2.1
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/ezyang-htmlpurifier.svg)](https://hysenlabs.com/projects/ezyang-htmlpurifier)