# mammoth.js: converting .docx to semantic HTML instead of pixel-perfect markup

> mammoth.js turns Word, Google Docs and LibreOffice .docx files into clean HTML by reading style names rather than formatting. It suits pipelines that want structure; it fails on documents that only use direct formatting, and it does no sanitisation.

**mwilliamson/mammoth.js** — Convert Word documents (.docx files) to HTML

- Repository: https://github.com/mwilliamson/mammoth.js
- Stars: 6,311 · Forks: 669
- Language: JavaScript
- License: BSD-2-Clause
- Published: 2026-09-22 · Updated: 2026-09-22 · Language: en
- Canonical page: https://hysenlabs.com/projects/mwilliamson-mammoth-js

## The mismatch mammoth.js chooses to resolve in one direction

A .docx file is a zip of XML that describes layout: run properties, fonts, sizes, colours, spacing. HTML describes structure. The README states this plainly: there is a large mismatch between the structure used by .docx and the structure of HTML, and the conversion is unlikely to be perfect for more complicated documents. Mammoth picks a side. It reads the semantic information, such as the style name attached to a paragraph, and discards the rest. A paragraph styled `Heading 1` becomes an `h1`. It does not become a span with a 16pt bold font.

That decision defines the audience. If you are building a content pipeline where the .docx is a delivery format and the HTML is the durable artefact, mammoth is aimed at you. If you are building a preview that must look like the Word document, it is the wrong tool, because it deliberately throws away the information you need. The README is explicit that Mammoth works best if you only use styles to semantically mark up your document. That is a constraint on the people writing the documents, not on the code.

## How the conversion pipeline is wired

The published package exposes a single main entry, `./lib/index.js`, and the browser field remaps two modules: `./lib/unzip.js` becomes `./browser/unzip.js`, and `./lib/docx/files.js` becomes `./browser/docx/files.js`. That is the whole portability story. The core reads a zip archive and walks the document XML; in Node the unzip implementation is the one from the dependency list (`jszip`), and in the browser it is swapped for a different implementation behind the same interface.

Conversion returns a promise. The resolved object carries `value` and `messages`, and the README describes `messages` as holding things such as warnings during conversion. That second field matters more than it looks: it is the only channel through which the library tells you it encountered something it could not represent. A pipeline that reads `value` and drops `messages` is discarding the diagnostic signal.

Style resolution is layered. There is a default map that handles common styles, and a user-supplied `styleMap` that takes precedence over the defaults. The map is a list of rules in a small textual syntax, and each rule has a left side matching a document style and a right side naming the HTML to emit. The README gives `p[style-name='Aside Heading'] => div.aside > h2:fresh` as an example, and the `:fresh` suffix on the right side is what closes an element so a subsequent matching paragraph does not nest inside the previous one.

## Installing mammoth.js and running a first conversion

The README gives one installation command. It requires Node, and the package declares `"node": ">=12.0.0"` in its engines field.

```bash
npm install mammoth
```

The package installs a `mammoth` binary, mapped from `bin/mammoth`. The CLI takes the path to a .docx file and an output file; with no output file, output goes to stdout. The result is an HTML fragment, not a full document, encoded as UTF-8.

```bash
mammoth document.docx output.html
```

If you want images written as separate files rather than inlined, pass `--output-dir`. The README warns that existing files will be overwritten.

```bash
mammoth document.docx --output-dir=output-dir
```

From code, the entry point is `convertToHtml`, which returns a promise resolving to an object with `value` and `messages`.

```javascript
var mammoth = require("mammoth");

mammoth.convertToHtml({path: "path/to/document.docx"})
    .then(function(result){
        var html = result.value; // The generated HTML
        var messages = result.messages; // Any messages, such as warnings during conversion
    })
    .catch(function(error) {
        console.error(error);
    });
```

If you only need the text, `mammoth.extractRawText` ignores all formatting and puts two newlines after each paragraph. The README notes that `styleMap` can be a string as well as an array, with each line treated as a separate mapping and blank lines and `#` comment lines ignored, which makes it practical to load a style map from a file. The web demo is the fastest way to see the output shape: clone the repository, run `make setup`, and open `browser-demo/index.html`.

## The supported element set, and where it stops

The README lists what is currently supported: headings, lists, tables, footnotes and endnotes, images, bold, italics, underlines, strikethrough, superscript and subscript, links, line breaks, text boxes, and comments. Two entries on that list come with caveats stated in the same paragraph. Table formatting such as borders is ignored, though the text inside the table is treated like text elsewhere in the document. Text box contents are treated as a separate paragraph that appears after the paragraph containing the text box, which means the reading order in the output can differ from the visual order on the page.

Those are not bugs to be filed. They follow from the same design choice that makes the output clean. Anything that exists only as visual arrangement has no structural equivalent to map onto, so it is dropped. The practical consequence is that a document whose meaning lives in its layout, a form, a two-column comparison, a caption positioned beside a figure, will lose that meaning in conversion and the library will not necessarily flag it.

The other boundary is security, and it is stated twice in the README in bold: Mammoth performs no sanitisation of the source document, and should therefore be used extremely carefully with untrusted user input. A .docx can carry content that becomes dangerous markup once converted. If uploads from users reach mammoth, the HTML it produces must pass through a sanitiser before it reaches a browser. This is a design boundary, not an oversight, but it moves real work onto the integrator.

## Markdown output is deprecated, and the README says so

The package description mentions Markdown, and the CLI accepts `--output-format=markdown`. The README's Markdown section opens by stating that Markdown support is deprecated, and recommends generating HTML and using a separate library to convert that HTML to Markdown, which it says is likely to produce better results. If your pipeline currently passes `--output-format=markdown`, the documented path forward is a two-stage conversion rather than the built-in flag. That is a concrete migration cost to weigh, and the README does not document a deprecation timeline or a removal version.

## mammoth.js compared with a layout-preserving converter

The natural alternative is a converter that renders the .docx as it looks, producing positioned HTML or an image-like representation of the page. The difference is not quality, it is what survives the trip. A layout-preserving converter keeps fonts, sizes, colours, spacing and page geometry, so a two-column layout still reads as two columns. It also emits far more markup, most of it presentational, and the semantic meaning of a paragraph is not recoverable from it.

Mammoth inverts those properties. It emits an HTML fragment with headings, lists and links, and no styling. Downstream, that fragment is easy to restyle with your own CSS, easy to index, and easy to diff. The cost is that any document relying on direct formatting rather than named styles converts to something close to plain paragraphs. Choosing between the two is really choosing whether the HTML is a rendering or a representation. For a CMS, a search index, or a documentation pipeline, representation is what you want.

## Maintenance, licence and the cost of upgrading

The repository is not archived, and the last push was on 2026-09-19. The published version in package.json is 1.12.3, while the most recent release listed is 1.2.5 from 2016-11-13, so the version numbering in the manifest and the release feed do not line up; anyone pinning by release notes should check which number their tooling actually resolves. The dependency list is short and largely stable, with `jszip`, `@xmldom/xmldom`, `argparse`, `underscore`, `xmlbuilder`, `lop`, `base64-js` and `dingbat-to-unicode`. A short dependency tree is the main reason upgrades here tend to be cheap, though `underscore` and `argparse` are older choices than a new project would make.

On licence: the package declares BSD-2-Clause, and the repository carries a LICENSE file at the top level. That is a permissive licence, but it is worth reading the actual file rather than the manifest field, and if you redistribute the bundled browser build, note that the build step uses browserify-prepend-licenses, which implies third-party licences are prepended to the generated file. This is a description of what the repository contains, not legal advice.

## Conclusion

Adopt mammoth.js if your .docx files use named styles and you want an HTML fragment you can post-process, and if you can guarantee the input is trusted or sanitise the result yourself. Do not adopt it if you need to reproduce a document's visual layout, or if you are feeding it user uploads without a sanitiser in front. Before committing, verify three things: that your authors actually apply named styles rather than toolbar formatting, that the element set you need (tables, footnotes, images, text boxes) is covered by the README's supported list, and that your Node runtime satisfies the engines field, which declares node >=12.0.0.

## FAQ

### Is mammoth.js available on NPM?

Yes. The README's installation section gives the command npm install mammoth, and package.json declares the package name as mammoth with main set to ./lib/index.js.

### How can I convert a DOCX file to HTML with mammoth.js?

From the command line, run mammoth document.docx output.html, or omit the output file to write to stdout. From code, call mammoth.convertToHtml with a path and read result.value for the HTML fragment.

### How do I use mammoth.js in my own code?

Require it with var mammoth = require("mammoth"), then call convertToHtml, which returns a promise resolving to an object with value and messages. A custom styleMap can be passed as the second argument to control which docx styles map to which HTML elements.

### What is mammoth.js?

It is a converter that turns .docx documents from Word, Google Docs and LibreOffice into HTML. It aims to produce simple and clean HTML by using the semantic information in the document, such as style names, and ignoring details like fonts and colours.

### Can mammoth.js convert an HTML file back to DOCX?

No. The README describes one direction only: converting .docx documents to HTML, plus a deprecated Markdown output mode. Nothing in the documented options produces a .docx file.

## Sources

- [Issues](https://github.com/mwilliamson/mammoth.js/issues)
- [License: BSD-2-Clause](https://github.com/mwilliamson/mammoth.js/blob/master/LICENSE)
- [mwilliamson/mammoth.js on GitHub](https://github.com/mwilliamson/mammoth.js)
- [README](https://github.com/mwilliamson/mammoth.js/blob/master/README.md)
- [Releases](https://github.com/mwilliamson/mammoth.js/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/mwilliamson-mammoth-js
