# lxml: the Python XML and HTML toolkit that binds libxml2 and libxslt

> lxml is a C-backed Python library for XML and HTML processing, built on libxml2 and libxslt. It is fast, it is not the standard library, and installing it is where most people first get stuck.

**lxml/lxml** — The lxml XML toolkit for Python

- Repository: https://github.com/lxml/lxml
- Website: https://lxml.de/
- Stars: 3,062 · Forks: 651
- Language: Python
- License: BSD-3-Clause
- Published: 2026-09-24 · Updated: 2026-09-24 · Language: en
- Canonical page: https://hysenlabs.com/projects/lxml-lxml

## What lxml solves, and who actually needs it

The Python standard library ships xml.etree.ElementTree, which parses XML and is pure Python. That is enough for configuration files and small documents. It is not enough when you need XPath expressions, XSLT transforms, namespace-aware processing at volume, or HTML that is malformed enough that a strict parser refuses it. lxml fills that gap by wrapping two mature C libraries, libxml2 and libxslt, behind a Python API that mostly mirrors the ElementTree interface.

The audience is narrow but deep: people scraping HTML, people consuming SOAP or RSS feeds, people transforming documents with stylesheets, and people who parse XML in a loop where per-document overhead matters. The README describes the library as feature-rich and memory friendly, and the repository layout backs that up: a src/ tree with Cython sources, a samples/ directory with simple.xml and simple-ns.xml, and a benchmark/ directory. The presence of benchmark/ and of a valgrind-python.supp suppression file says the maintainers treat memory behaviour as something to measure, not something to assume.

## How lxml is put together: Cython over libxml2 and libxslt

lxml is not a pure-Python reimplementation. The build system compiles Cython sources against libxml2 and libxslt. The pyproject.toml pins the versions used for the published wheels: LIBXML2_VERSION is 2.14.6 and LIBXSLT_VERSION is 1.1.43, with ZLIB_VERSION 1.3.2 and LIBICONV_VERSION 1.18. Those pins matter because the behaviour of XPath, HTML recovery and encoding handling comes from the C libraries, not from Python code you can patch.

The repository root also carries libxslt-1.1.43-backport1.patch, which indicates the project patches the upstream C library rather than tracking it untouched. That is a real maintenance commitment, and it is also a reason to trust the wheels more than a from-source build you assemble yourself.

On the Python side, the import surface is split. The documented modules are lxml.etree, lxml.objectify and lxml.html. The cibuildwheel test command in pyproject.toml is a single import line for all three, which is the project's own smoke test for a built wheel. Data flow is conventional: you parse a byte string or file into a tree, query or mutate it, then serialize it back out. The work happens in C, and Python objects are thin handles onto it.

## Installing lxml and parsing your first document

Installation guidance lives in INSTALL.txt, linked from the README as the installation page. The normal path is a wheel from the Python Package Index, which the README notes is also how many Linux and macOS package distributions ship it.

```bash
pip install lxml
```

If a compatible wheel exists for your platform, pip installs a compiled binary and you never touch a compiler. If it does not, pip falls back to building from source, and the build requires Cython 3.3.0 or later, per requirements.txt and the build-system table in pyproject.toml. The setup.py script checks your interpreter version first and exits with a message if it is older than Python 3.9.

To confirm the whole install rather than one module, the project's own check is the test command it runs against built wheels:

```bash
python -c "import lxml.etree, lxml.objectify, lxml.html"
```

No output means all three extension modules loaded. A traceback naming lxml.etree is the failure you will hit most often, and it usually means the extension did not build against a matching libxml2.

The repository's samples/ directory contains simple.xml and simple-ns.xml, which are the documents to parse once the import succeeds. The Makefile shows how to build the extension in place from a checkout, which is the route to take when no wheel exists for your platform:

```bash
make inplace
```

That target runs setup.py build_ext with the Cython and coverage flags the Makefile resolves for your interpreter, and installs the result in place rather than into site-packages.

## The module-not-found error and the build it hides

The single most common lxml complaint is an import failure, and the cause is rarely that the package is missing from pip. It is that the compiled extension could not load. lxml is a binary distribution, so the failure modes are the ones you get with any C extension: a wheel built for a different ABI, a system libxml2 that is too old or too new, or a source build that silently linked against headers the runtime library does not match.

The project's own build configuration shows how much surface this covers. The cibuildwheel skip list in pyproject.toml excludes cp38 and pp38, drops musllinux_i686 and ppc64le builds for several Python versions, and explicitly skips cp313t, with a comment that free-threaded Python 3.13 is not worth supporting. If you are on one of those combinations, there is no wheel and you are in source-build territory.

The second real limitation is scope. lxml parses documents; it does not validate against a schema you have not loaded, and it does not protect you from hostile input by default. XML entity expansion and external entity resolution are parser features inherited from libxml2, and the project ships a SECURITY.md at the repository root, which is the place to look before feeding untrusted XML into a parser. If your input is untrusted and you have not configured the parser, lxml is the wrong tool until you do.

## lxml against xml.etree and against BeautifulSoup

The obvious alternative is the standard library's xml.etree.ElementTree. The difference is not stylistic. ElementTree is pure Python and ships with every interpreter; lxml is compiled and must be installed. ElementTree has a limited XPath subset; lxml supports XPath through libxml2, which is a different capability class. If your queries are getroot, find and iter, ElementTree is sufficient and you avoid a binary dependency entirely. The moment you need a real XPath expression, an XSLT transform, or HTML recovery, the standard library stops being an option.

The other comparison is HTML parsing. BeautifulSoup is a pure-Python parser that accepts multiple backends, and lxml can be one of those backends. That is the honest relationship: they are not always competitors. If your HTML is broken in ways that require guessing, BeautifulSoup's tolerance is a feature lxml.html does not try to match. If you need speed and XPath over HTML, lxml.html is the better fit, and people frequently run BeautifulSoup with the lxml parser underneath to get both.

A third point of comparison is objectify, which lxml ships and the standard library does not. It maps XML elements onto Python objects with attribute-style access, which is convenient for data-shaped XML and awkward for document-shaped XML. The pyproject.toml test command imports it alongside etree and html, so it is a first-class module rather than an experiment.

## Maintenance, releases and what the licence permits

The last push to the repository was on 2026-09-10, and the most recent release listed is lxml-6.1.3-1 from 2026-09-02. A 7.0.0b1 prerelease appeared on 2026-08-22, so a major version is in flight. The repository is not archived. The setup.py source describes the maintenance model for stable series: after an official release, bug fixes land on a branch named lxml-<branch_version>, and pip can install that unreleased branch state directly from a GitHub archive URL. That requires Cython at an appropriate version, which is stated in the same comment.

The upgrade cost is dominated by the C dependencies rather than by API churn. Every wheel bundles specific libxml2 and libxslt versions, and the pyproject.toml currently pins 2.14.6 and 1.1.43. When you upgrade lxml, you may also be upgrading the parser underneath, and that can change how malformed input is recovered or how encodings are detected. Pin lxml in your lockfile and read CHANGES.txt before moving a major version, especially with 7.0.0 in prerelease.

Licensing is BSD-3-Clause for lxml itself, per LICENSE.txt. The repository also carries LICENSES.txt, which is the file to read for the bundled components. libxml2 and libxslt have their own terms, and the wheels ship them statically. If your organisation reviews third-party licences, LICENSES.txt is the document that matters, not the single SPDX identifier on the project page. This is a description of what the repository contains, not legal advice; your counsel decides what obligations apply.

## Conclusion

Adopt lxml when you need XPath, XSLT or HTML parsing at a scale where pure-Python parsers become the bottleneck, and when you can pin a wheel or a C toolchain in your build. Do not adopt it if you only parse small, trusted XML documents and want zero binary dependencies; xml.etree in the standard library covers that case. Before committing, verify that a wheel exists for your Python version and platform, since the project skips cp38 and several ppc64le and s390x builds, and check whether your code depends on libxslt, which is a separate C library bundled into the build.

## FAQ

### Is lxml a standard Python library?

No. lxml is a third-party package distributed through the Python Package Index and through Linux and macOS package distributions, and it must be installed separately. The standard library ships xml.etree.ElementTree, which covers a smaller subset of XML processing.

### Why is the module lxml not found?

The usual cause is that the compiled extension failed to load rather than that pip never installed the package. lxml is a binary distribution built against libxml2 and libxslt, so a missing wheel for your platform or a mismatched system library produces an import error. The project's own check against built wheels is the import line for lxml.etree, lxml.objectify and lxml.html, which tells you whether all three extension modules load.

### What does lxml stand for and what is its purpose?

The name is a contraction of the words it operates on, XML and HTML, with the leading l. Its purpose is processing XML and HTML in Python, including XPath queries and XSLT transforms, by wrapping the libxml2 and libxslt C libraries.

### How do I install lxml in Python?

Run pip install lxml. If a compatible wheel exists for your Python version and platform, pip installs a prebuilt binary. If not, pip builds from source, which requires Cython 3.3.0 or later and a working C toolchain, and the source build requires Python 3.9 or newer.

## Sources

- [License: BSD-3-Clause](https://github.com/lxml/lxml/blob/master/LICENSE)
- [lxml/lxml on GitHub](https://github.com/lxml/lxml)
- [Project website](https://lxml.de/)
- [README](https://github.com/lxml/lxml/blob/master/README.md)
- [Releases](https://github.com/lxml/lxml/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/lxml-lxml
