Library / SDK
mathiasbynens/he avatar
mathiasbynens/he

he: an HTML entity encoder that follows the spec instead of the shortcut

A robust HTML entity encoder/decoder written in JavaScript.

3,594 stars264 forksJavaScriptMIT

At a glance

What is it?
Mathias Bynens's he library implements the WHATWG named character reference list, resolves ambiguous ampersands the way a browser does, and ships a CLI and a man page that the README never mentions.
Who is it for?
he is the library to reach for when you need output that a browser will parse exactly as you intended, particularly for ambiguous ampersands, non-ASCII text, and symbols above the basic multilingual plane. Its edge over a hand-rolled replace chain is that the entity table is generated from the WHATWG data file rather than typed by hand, and its edge over a DOM parser is that encoding one string does not require constructing a document.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 4 days ago.
What is it written in?
Mainly JavaScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 6, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Why ambiguous ampersands are the whole point

The README opens by naming two differentiators, and the first is unusual enough to be the reason to choose this library at all: ambiguous ampersands. A bare ampersand in text is ambiguous, because the HTML tokenizer has to decide whether what follows begins a character reference or is simply text. The library's author wrote a separate explanation of the problem and links it from the README, and the project's stated goal is to resolve it the way a browser would rather than the way a regular expression would.

The second differentiator is astral symbols. Text above the basic multilingual plane, which sits beyond the sixteenth code plane boundary, needs a surrogate pair in JavaScript and has to be emitted as a single entity carrying a high code point. The README's own example string includes one of these and shows the correct hexadecimal escape in the output.

The supporting claim is coverage. The README says the library supports all standardized named character references as per the HTML specification, which is a stronger statement than having a long list. The evidence is in the build. There is a script that fetches the specification file and writes it into a data directory, a second script that fetches a spec document, and a build step that runs both and then formats the generated data. A table fetched from the specification and regenerated is a different maintenance proposition from a table someone typed out.

The practical consequence is that the output of encode is claimed to be valid HTML as long as the input contains only allowed code points. Code points that cannot be represented by a character reference pass through unchanged rather than being mangled, and the strict option converts that behaviour into throwing an exception, which is what you want inside an HTML parser or validator.

Installing an ES module with a CommonJS escape hatch

Installation is one command:

bash
npm install he

The import shape is named exports first:

js
import { decode, encode, escape, unescape } from 'he';

If you prefer the namespace form used throughout the examples, `import * as he from 'he'` gives you the same four functions as properties. Those four names deserve a pause. Encode and decode are the specification-complete pair, while escape and unescape are deliberately looser. That distinction is the library's most useful design decision, because it lets a caller choose between fidelity to the specification and a quick transform, and the README's structure makes the difference visible from the section headings alone.

Now the place where the README and the package manifest describe the package differently in emphasis. The README says the library is published as an ECMAScript module with named exports, and the manifest agrees: the module type is set to module, the export map points the package root at a single ES module file, and there is no CommonJS entry in the exports field. The README also states that on Node versions supporting require of an ES module, requiring the package returns the module namespace object, so `require('he').encode(...)` also works.

Both statements hold, and the second does less work than it appears to. CommonJS support is delegated to a Node runtime feature rather than shipped as a dual build, so the guarantee is tied to the runtime version and not to the package contents. The manifest states the floor plainly, at Node 22 or newer, and a version file for the runtime sits in the tree. For an application still targeting an older Node, the practical question is whether your runtime can require an ES module at all, and that answer is a property of your Node version rather than something the package can promise.

In a browser, the README's documented path is a module script tag pointing at the package's source file inside node_modules, which is a development setup rather than a bundled one. If you are shipping to browsers, you are expected to run this through your own bundler.

Four options that decide what the output looks like

The option surface is small, and each option changes one visible thing in the output, which makes it easy to reason about. The default is hexadecimal escapes with no named references, and the README's example makes the trade-off obvious:

js
he.encode('foo © bar ≠ baz 𝌆 qux');
// → 'foo © bar ≠ baz 𝌆 qux'

Setting useNamedReferences to true swaps in readable names where they exist, producing copy and ne entities for the same input, while the astral symbol still has to be a numeric escape because no named reference exists for it. The README attaches a warning to this option that is easy to miss and worth repeating: if compatibility with older browsers matters, leave it disabled. Named references are the part of the character reference grammar that browsers implemented later and less uniformly.

The decimal option swaps hexadecimal for decimal escapes, and it composes predictably with named references. Names win, and only entities without a name fall back to decimal, so the two options are not competing but layered.

The encodeEverything option extends escaping to printable ASCII that did not need it, which turns a readable string into a wall of hex codes. The README states that this option takes precedence over the option for allowing unsafe symbols, so setting the latter while this one is on has no effect. That precedence rule is documented rather than left to be discovered, which is a small sign of documentation quality.

The strict option is the one with real consequences for callers. Off by default, the encoder tolerates input that would cause a parse error, which is convenient for arbitrary user text. Turned on, it either throws or returns valid HTML, which is what makes the library usable as a component of an HTML parser or validator rather than only as an escape helper.

A CLI and a man page that the README skips

The repository contains things the README does not document, and they are not incidental. At the root there is a bin directory and a man directory. The manifest's published file list contains exactly four entries: the licence, the source directory, the bin directory and the man directory. It also declares a man page path and maps a binary named after the package to a script inside bin.

A man page and an executable named after the package mean this is a command-line tool as well as a library. The README's installation and API sections never mention either. Evaluating the project from the README alone would lead you to conclude it is library code only, and you would miss that a command-line entry point ships in the tarball. That gap matters if you were about to shell out to something else that does the same job.

The build tooling shows how carefully the library is kept. Formatting is handled by a dedicated formatter rather than a general-purpose linter, the test script runs Node's built-in test runner over a glob of test files, and a separate coverage script restricts collection to the source directory. The development dependencies are few and telling: a DOM implementation and a regenerator, which together let the test suite compare behaviour against a real parser rather than against hand-written expectations.

Those two dependencies are the strongest available evidence for the README's central claim about browser fidelity. A library that claims to handle ambiguous ampersands like a browser, and that tests itself against a DOM implementation while regenerating fixtures, is making a checkable claim rather than an ornamental one. It is also the reason the coverage script exists: the library is small enough that measuring it is cheap, and a percentage nobody looks at is worth nothing on its own.

Version 2.0.0 with no release notes in the repository

The manifest says version 2.0.0. The README carries an npm version badge linking to the package page, and it links to an online demo for trying the library without installing anything. The repository itself carries no GitHub releases, so the version line you depend on lives entirely on the npm registry.

That arrangement is ordinary for a small utility library, and it has one practical consequence worth stating plainly. There is no release-note history on the repository page to read before upgrading. For a library at version 2, the interesting question is what broke in the transition from 1, and that answer is not in this repository.

The last push to the default branch was on 2026-10-04 and the repository is not archived, so the code is being worked on. Its maintenance burden is unusually low, and the reason is structural rather than about discipline. The entity data is fetched from the specification rather than maintained by hand, so the part that would normally rot is generated at build time. What the maintainer does maintain is the parsing and encoding logic, the option surface, and the test suite.

The licence is MIT, stated in the manifest, with the licence text in a file at the root. The manifest's own description calls this an encoder and decoder with full Unicode support, which is a slightly stronger phrasing than the README makes. Full Unicode support includes planes the HTML character reference grammar cannot always represent, and the README is careful to say those code points pass through unchanged. That gap between a one-line package description and a careful README is worth remembering when you read any package summary.

What this library is not, and where it loses

There are two reasonable alternatives, and the README addresses neither, which is worth knowing before you commit.

The first is doing it in the DOM. Any browser, and any server-side DOM implementation, will encode and decode correctly by construction, because it is the same code path a browser uses. If you already have a document, a parser, or a templating step that escapes attributes, reaching for a DOM is one fewer dependency. What you give up is that constructing a DOM to encode one string is a heavy operation, and that is exactly the case this library exists for: a pure string transform, with no document, no parser and no surrounding context.

The second is a hand-rolled escape function. Replacing the ampersand, the angle brackets, both quote characters and the backtick gets you most of the way for a narrow case, and it is four lines long. What it cannot do is resolve ambiguous ampersands, handle astral symbols, or tell you whether an input code point is representable at all. Those three are the entire reason this library exists, so the choice is between a short function you trust for a narrow case and a generated table you trust for a broad one.

The remaining boundary is scope. This library encodes and decodes. It does not sanitise, and encoding with default options escapes the six characters that matter for text and attribute contexts, which is not the same guarantee as HTML sanitisation of arbitrary markup. If you need sanitisation, you need a sanitiser, and no amount of entity encoding substitutes for one. The README's distinction between the strict and non-strict modes is a reasonable place to start reading if you need to know exactly where that line sits.

For dependency cost: this is a zero-runtime-dependency package. Everything in the manifest's development section is a build-time or test-time tool, which means adopting it does not add anything to a production bundle beyond the library's own source.

Editorial conclusion

he is the library to reach for when you need output that a browser will parse exactly as you intended, particularly for ambiguous ampersands, non-ASCII text, and symbols above the basic multilingual plane. Its edge over a hand-rolled replace chain is that the entity table is generated from the WHATWG data file rather than typed by hand, and its edge over a DOM parser is that encoding one string does not require constructing a document. Two things to check before adopting it: the manifest declares version 2.0.0 and requires Node 22 or newer, so older runtimes are out, and the package is published as an ES module, which means CommonJS consumers depend on the runtime being able to require an ES module rather than on a dual build. For command-line use, the executable and its man page ship in the package even though the README documents only the JavaScript API.

Frequently asked questions

How do I install he and import it in JavaScript?

Install it from npm with `npm install he`, then import the named exports with `import { decode, encode, escape, unescape } from 'he'` or use the namespace form `import * as he from 'he'`. The package is published as an ECMAScript module and the manifest requires Node 22 or newer.

What is the difference between he.encode and he.escape?

encode implements the specification's rules for character references, including ambiguous ampersands and astral code points, and honours options such as useNamedReferences, decimal, encodeEverything and strict. escape is the looser, quick transform offered alongside it, and the README documents both under separate headings.

Does he support symbols above the basic multilingual plane?

Yes. The README's own examples include a character above the sixteenth plane and show it encoded as a single hexadecimal character reference, which the library notes other JavaScript solutions often get wrong.

Does he ship a command line tool?

The repository and the package manifest both say yes. The manifest maps a binary named after the package into the bin directory, declares a man page, and the published file list includes both directories. The README documents only the JavaScript API, so the command line surface is undocumented in prose.

Official sources

  1. Issues
  2. License: MIT
  3. mathiasbynens/he on GitHub
  4. Project website
  5. README