# js-xss: whitelist HTML sanitization for Node.js

> The xss npm package filters untrusted HTML against a whitelist of tags and attributes. It is a string-level sanitizer for Node.js and the browser, not a browser sandbox, and its defaults need review before you ship them.

**leizongmin/js-xss** — Sanitize untrusted HTML (to prevent XSS) with a configuration specified by a Whitelist

- Repository: https://github.com/leizongmin/js-xss
- Website: http://jsxss.com
- Stars: 5,309 · Forks: 634
- Language: HTML
- License: NOASSERTION
- Published: 2026-09-22 · Updated: 2026-09-22 · Language: en
- Canonical page: https://hysenlabs.com/projects/leizongmin-js-xss

## What js-xss actually filters, and who needs it

The package targets one job: take a string of HTML that came from a user and return a string with disallowed tags and attributes removed or escaped. The README frames it as a module "used to filter input from users to prevent XSS attacks." That places it in the input-handling layer of a Node.js application, typically where a comment, a forum post, a profile bio or a rich-text field is written to the database or rendered back to other users.

The audience is narrow and specific. You need this when your application must accept markup from users and still show it as markup. If you can render user content as plain text, or inside a sandboxed iframe on a separate origin, the sanitizer is not the right control and adds a parsing surface you do not need. The project's own examples lean toward forum-style software: the README lists nodeclub (a Node.js bbs), cnpmjs.org and cocalc.com among the projects using it. That is the shape of the problem it was built for.

The whitelist framing matters. Older filters worked by blacklisting known-bad patterns, which loses to encoding tricks and new browser behaviour. js-xss inverts that: anything not named in the whitelist is dropped. The README points readers at the OWASP XSS Filter Evasion Cheat Sheet as reference material, which tells you the authors expect you to think about evasion rather than trust a default.

## The whitelist mechanism and the onTag / onTagAttr hooks

Configuration is a plain object. A whitelist maps tag names to the attributes allowed on them, in the form { 'tagName': ['attr-1', 'attr-2'] }. Tags and attributes outside that map are filtered out. The README's example is short enough to quote in shape: with a whitelist of { a: ['href', 'title', 'target'] }, the input <a href="#" onclick="hello()"><i>Hello</i></a> becomes <a href="#">&lt;i&gt;Hello&lt;/i&gt;</a>. The onclick attribute is gone and the i tag is escaped into text rather than removed outright.

That escaping detail is worth pausing on. Non-whitelisted tags are not deleted from the output; they are entity-encoded so they display as literal text. For a comment box that is usually what you want, because a user who types angle brackets sees them. For a system that expects clean markup, it means your stored output can contain visible &lt; and &gt; sequences, and downstream consumers that re-parse the string need to know that.

When the whitelist is not expressive enough, two callbacks take over. onTag receives the tag name, the raw HTML of the tag and an options object carrying isWhite, isClosing, position and sourcePosition. Returning a string replaces the tag with that string; returning nothing falls back to the default path, which routes to onTagAttr for allowed tags and onIgnoreTag for the rest. onTagAttr handles individual attributes on a matched tag. The README also documents allowList as an alias with the same behaviour as whiteList, so either key works.

The library is a string rewriter, not a DOM. It scans HTML text, matches tags and attributes with its own tokenizer, and emits a new string. That is why it runs in Node.js without a DOM shim and why it can run in a Web Worker, which one of the example files (example/worker.js) demonstrates. It also means the sanitizer's model of HTML is its own, and any divergence from a real browser parser is a place where the two disagree about what a document means.

## Installing xss and sanitizing your first string

The README gives two install paths. The npm package is the one most Node.js users want:

```bash
npm install xss
```

After that, requiring the module returns a callable function. Passing a script tag through it returns escaped text, which is the behaviour the README's first example shows:

```javascript
var xss = require("xss");
var html = xss('<script>alert("xss");</script>');
console.log(html);
```

If you need to know what was stripped, filterXSSWithResult returns both the sanitized HTML and a list of removed items. The README's example input <script>alert("xss");</script><a href="#" onclick="evil()">click</a> produces html containing the escaped script text and an anchor with only href, plus a removed array holding two entries of type "tag" for the script open and close and one of type "attr" for onclick on the a tag. That array is the practical way to log or alert on suspicious submissions.

```javascript
var xss = require("xss");
var result = xss.filterXSSWithResult('<script>alert("xss");</script><a href="#" onclick="evil()">click</a>');
console.log(result.html);
console.log(result.removed);
```

The package also ships a command line tool, declared in package.json as the bin entry xss. It processes a file with -i and -o, and has an interactive test mode with -t. Running xss -h prints the full flag list, which the README recommends rather than reproducing.

```bash
xss -i origin.html -o target.html
xss -t
```

For browser use the README shows a shim script tag loading dist/xss.js and then calling the global filterXSS. It explicitly warns against using the rawgit.com URL from that snippet in production, so treat the snippet as an illustration of the call shape and host the file yourself.

## Where the default configuration and the string model fall short

The default whitelist is the first thing to examine. The README says to refer to xss.whiteList for it, and it does not print the list in the documentation. That means you cannot judge the default policy from the README alone; you have to read the source or inspect the exported object at runtime. Shipping the default unchanged is a decision, and it is one the documentation does not help you evaluate.

Because the sanitizer works on strings rather than a parsed DOM, its correctness depends on its tokenizer agreeing with browsers about malformed input. Browsers are famously forgiving: they repair unclosed tags, handle unusual nesting and interpret attributes in ways a hand-written scanner may not. The project ships a substantial test directory and CI configuration, and the README links a test coverage badge, but coverage of a tokenizer is not the same as equivalence with a browser parser. If your threat model includes attackers who control the exact bytes of the HTML, this difference is the interesting attack surface.

CSS is handled by a separate dependency, cssfilter, pinned at 0.0.10 in package.json. That is a distinct codebase with its own rules, and the xss README does not document its behaviour. If you allow a style attribute or a style tag, the filtering of what goes inside is happening somewhere the main documentation does not describe.

The package also has no release notes in the repository documentation, and the repository carries a CHANGELOG.md that is not reproduced in the README. Version 1.0.15 is what package.json declares. Pinning matters more than usual for a security filter, because a sanitizer that changes behaviour between patch versions can silently change what your application accepts.

## xss versus DOMPurify and validator's xss helper

The README's own benchmark section compares the module against the xss() function from validator@0.3.7, reporting 22.53 MB/s for this module and 6.9 MB/s for that one. The README labels the numbers as being for reference only and points to the benchmark directory for the test code. Treat them as the authors' measurement of an old comparison, not as a current claim about either library.

The more useful contrast is architectural. DOMPurify parses input through the browser's own HTML parser, which means its notion of the document is the same one the browser will use when it renders the result. That closes the parser-divergence gap described above, at the cost of requiring a DOM: in Node.js that usually means a DOM implementation such as jsdom. js-xss needs no DOM, which is why it runs in a plain Node process and in a Web Worker, and that is the trade it makes.

validator's xss helper is a different shape again: it is one function inside a broader string-validation library, and the README's benchmark treats it as a slower, less configurable option. If you already depend on validator for other checks, its xss helper may be adequate for simple escaping, but the whitelist model and the onTag/onTagAttr callbacks are the reason to pick this package instead.

There is also the option of not sanitizing at all. If user content can be rendered as text, or isolated in an iframe on a different origin, the correct amount of HTML sanitization is none. Choosing a sanitizer is choosing to accept markup, and that choice should be deliberate.

## Maintenance, licence and the cost of staying current

The repository is not archived, and the last push was on 2026-05-06. That is recent enough that the project is not abandoned, but the repository documentation contains no release notes, so there is no documented upgrade path between versions and no stated policy on security releases. The repository does include a SECURITY.md, which is where the project says to look for its disclosure process; read it before you depend on the package in a security-sensitive path.

package.json declares the licence as MIT and the npm metadata badge in the README points to the same. The repository metadata for this review reports the licence as NOASSERTION, which means the automated classifier could not confirm it. Those two disagree, and the discrepancy is worth resolving against the LICENSE file at the root of the repository rather than either summary. This is a factual note, not legal advice; if the licence matters to your organisation, have someone read the file.

The upgrade cost has two parts. First, the package itself: at version 1.0.15 with a small dependency set (commander and cssfilter), the surface is limited, but a change to the default whitelist or to the tokenizer changes your application's behaviour without any code change on your side. Second, the data you already stored. If a new version escapes a tag differently, old rows keep their old sanitization and new rows get the new one, so a re-render pass may be needed. Neither the README nor the repository documentation describes a migration procedure for stored content.

## Conclusion

Adopt xss when you need to store or render user-submitted HTML in a Node.js service and you are prepared to write and test your own whitelist; the default list is a starting point, not a policy. Do not adopt it as your only defense for content you can render inside a sandboxed iframe or as text, and do not expect it to sanitize CSS beyond what cssfilter covers. Before shipping, run your own payload corpus through filterXSSWithResult and inspect the removed array, because that array is the only visibility the library gives you into what it stripped.

## FAQ

### What is XSS in JavaScript, and what does js-xss do about it?

Cross-site scripting is the class of attack where untrusted input reaches a page as executable markup; the README links the Wikipedia article on it as background. The xss module addresses it by filtering user input against a whitelist of allowed tags and attributes, returning a string with everything else removed or escaped.

### How do I install js-xss and run it on a string?

The README gives npm install xss, then require("xss") and call the returned function with your HTML string. The same page shows filterXSSWithResult if you also want the list of removed tags and attributes.

### Can I use js-xss in the browser instead of Node.js?

Yes. The README shows a script tag loading dist/xss.js and then calling the global filterXSS, and it also documents an AMD shim configuration. It warns explicitly not to use the rawgit.com URL from those snippets in production.

### What is the default whitelist in js-xss?

The README does not print it; it says to refer to xss.whiteList. You have to inspect the exported object or read the source to see which tags and attributes are allowed before you rely on the defaults.

### How do I customize which tags and attributes js-xss allows?

Pass an options object as the second argument to xss(), or construct a FilterXSS instance with new xss.FilterXSS(options). The whiteList key maps tag names to arrays of permitted attribute names, and allowList works the same way.

### Does js-xss need a DOM to run?

No. It rewrites the HTML string with its own tokenizer, which is why the README can show it running in a Web Worker. The trade-off is that its model of malformed HTML is its own rather than the browser's parser.

## Sources

- [Issues](https://github.com/leizongmin/js-xss/issues)
- [leizongmin/js-xss on GitHub](https://github.com/leizongmin/js-xss)
- [Project website](http://jsxss.com)
- [README](https://github.com/leizongmin/js-xss/blob/master/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/leizongmin-js-xss
