Library / SDK
zeux/pugixml avatar
zeux/pugixml

pugixml: A C++ XML Parser With XPath, and When It Is the Wrong Tool

Light-weight, simple and fast XML parser for C++ with XPath support

4,653 stars809 forksC++MIT

At a glance

What is it?
pugixml is a small MIT-licensed C++ library that builds a DOM from an XML file or buffer and queries it with XPath 1.0. Its last push was on 2026-06-16, and the design trade-offs are worth knowing before you adopt it.
Who is it for?
Adopt pugixml when you want a permissively licensed C++ XML DOM with XPath and no external runtime dependency, and when you can compile src/pugixml.cpp into your build. Do not adopt it if you need streaming or pull parsing of documents larger than memory, XPath 2.0 or 3.0, schema validation, or a Python binding that the project itself maintains.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 106 days ago.
What is it written in?
Mainly C++, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What pugixml Solves, and Who It Is Written For

pugixml is a C++ XML processing library. It gives you three things in one dependency: a DOM-like tree with traversal and modification, a parser that builds that tree from an XML file or buffer, and an XPath 1.0 implementation for tree queries. Unicode is handled through interface variants and automatic conversion between encodings during parsing and saving.

The audience is a C++ project that already has XML on its input or output path and does not want to drag in a larger stack. The README describes the library as light-weight and simple, and the repository layout backs that up: the implementation is a single source file, src/pugixml.cpp, with a matching header. There is no runtime, no code generation step, and no separate data directory to ship.

The README also states that pugixml is used by open-source and proprietary projects for performance and an easy-to-use interface. That sentence is a claim about adoption, not a benchmark, and the repository does not publish numbers to support it. Treat it as context, not evidence.

How the DOM and XPath Layers Fit Together

The data flow is linear. You declare a pugi::xml_document, call one of its load methods with a file path or a memory buffer, and receive a pugi::xml_parse_result. That result converts to bool, so an if (!result) check is the documented way to detect a failed parse. Once the document is loaded, the tree is yours to walk with child(), children(), attribute() and similar accessors, or to query with select_nodes() and an XPath expression.

The README example shows both styles on the same file. The traversal version walks doc.child("Profile").child("Tools").children("Tool") and reads each Tool's Timeout attribute as an integer. The XPath version replaces that walk with doc.select_nodes("/Profile/Tools/Tool[@Timeout > 0]") and iterates the returned node set. The two produce the same output, which makes the example useful as a comparison rather than just a snippet.

The parser is described as constructing the DOM tree from an XML file or buffer, so the whole document is resident after a successful load. That is the central architectural fact. Everything convenient about pugixml follows from it, and so does its main limitation.

Installing pugixml and Running a First Query

The repository is built with CMake. The top level contains CMakeLists.txt alongside a Makefile used for the project's own test runs, and the documentation links on pugixml.org point to a quick-start guide and a complete reference manual. The README does not include a distribution-specific install command, so the practical route is to add the source to your build rather than to look for a system package.

The Makefile shows how the project compiles the library itself. Its release configuration appends -O3 and -DNDEBUG to the standard flag set, and its sources list is src/pugixml.cpp plus the test files under tests/.

makefile
config=release
SOURCES=src/pugixml.cpp $(filter-out tests/fuzz_%,$(wildcard tests/*.cpp))

With the library compiled into your build, the README's XPath example is the shortest realistic first use. It loads a file, checks the parse result, and selects nodes whose Timeout attribute is greater than zero.

c++
#include "pugixml.hpp"
#include <iostream>

int main()
{
    pugi::xml_document doc;
    pugi::xml_parse_result result = doc.load_file("xgconsole.xml");
    if (!result)
        return -1;

    for (pugi::xpath_node node: doc.select_nodes("/Profile/Tools/Tool[@Timeout > 0]"))
        std::cout << node.node().attribute("Filename").value() << "\n";
}

Running it against a document with the expected structure prints one filename per matching Tool. If the file is missing or malformed, the program exits with status -1 and prints nothing, because the result check short-circuits before any query runs.

The Whole Document in Memory, and What That Rules Out

The parser builds a DOM, so peak memory tracks document size, not the size of the records you actually care about. A 2 GB XML dump is not a pugixml workload. If your input arrives as an unbounded stream, or if the machine parsing it has less headroom than the file, this library is the wrong shape and no amount of tuning changes that.

The same architecture has a second consequence: there is no incremental callback API in the repository's documentation. You cannot hand pugixml a chunk and have it emit elements as they close. The README and the example describe loading, then querying. Nothing in the project description suggests a pull-parser mode.

XPath support is version 1.0, as stated in the project description. Expressions that depend on later XPath revisions are outside the scope of what the library advertises. The README itself warns that the quick-start guide omits or only briefly mentions many important features, and directs you to the complete reference manual for detail. Read the manual before you design queries around a function you assume exists.

The README does not document rollback or recovery behaviour for a partially parsed document, and it does not describe error reporting beyond the parse result object. If your application needs to explain to a user which line of a malformed file broke, verify that the parse result carries enough detail for your error message before you build the UI around it.

pugixml Against tinyxml2, rapidxml and libxml2

The closest comparison is tinyxml2, which also offers a small C++ DOM and is also aimed at projects that want XML without a large dependency. The practical difference to check is API shape and query support: pugixml ships an XPath 1.0 implementation as part of the library, so selection logic can live in an expression string rather than in hand-written traversal code. If your queries are simple and fixed, tinyxml2's traversal style may be enough and you avoid learning XPath semantics.

rapidxml sits at the other end of the convenience spectrum. It is a header-only parser in the same lightweight tradition, but its in-situ model means the parser mutates the character buffer you hand it. That is a real constraint for callers who receive const data or who want to reuse the buffer after parsing. pugixml's documented interface loads from a file or buffer into a document object, which keeps the source buffer out of the picture.

libxml2 is the heavyweight alternative. It is a C library with a far broader feature surface, and it is the name that usually comes up when a project needs validation or a wider range of XML standards. The trade is build and dependency complexity against feature coverage. pugixml's answer is deliberately narrower: one C++ source file, MIT licensed, with XPath 1.0 and Unicode conversion.

Licence, Release Cadence and Upgrade Cost

pugixml is MIT licensed, and the README states the library is available free of charge under those terms. The repository also carries a LICENSE.md and a readme.txt that records the lineage: the work is based on the pugxml parser, which was released into the Public Domain by Kristen Wegner in 2003. That history matters if your legal review asks where the code came from. It is not legal advice, and the LICENSE.md file is the text your counsel should read.

Release cadence is slow and deliberate. v1.14 landed on 2023-10-01, v1.15 on 2025-01-10, and v1.16 on 2026-06-16. The last push to the repository was on 2026-06-16, so the project is not dormant, but you should not expect frequent drops. The upgrade cost is correspondingly low: because the library is one source file and one header, a version bump is a file replacement plus a rebuild, not a dependency graph change.

The Makefile reveals how the project tests itself, and it is worth reading before you vendor the source. It defines config values including debug, release, coverage, sanitize and analyze, and the sanitize configuration adds -fsanitize=address,undefined with -fno-sanitize-recover=all. The standard build flags are strict: -Wall -Wextra -Werror -pedantic and several more warning categories. If your own build uses a different warning set, expect to resolve warnings that the project's own configuration already treats as errors.

Editorial conclusion

Adopt pugixml when you want a permissively licensed C++ XML DOM with XPath and no external runtime dependency, and when you can compile src/pugixml.cpp into your build. Do not adopt it if you need streaming or pull parsing of documents larger than memory, XPath 2.0 or 3.0, schema validation, or a Python binding that the project itself maintains. Before you commit, verify the two things the README does not spell out: that your build toolchain can compile src/pugixml.cpp with the flags your project already uses, and that the XPath subset you rely on is covered in the complete reference manual rather than only in the quick-start guide. A single test file that loads one of your real documents and runs your intended XPath expression will settle both questions in an afternoon.

Frequently asked questions

How do I install pugixml?

The README does not give a package-manager command. The repository is built with CMake, and the top level contains CMakeLists.txt, so the usual route is to add the source to your own build. Because the implementation is a single file, src/pugixml.cpp, compiling it directly alongside your program also works.

How do I install pugixml on Ubuntu?

No Ubuntu package or apt command is documented in the repository. The README points to pugixml.org for the quick-start guide and the complete reference manual, and the repository ships CMakeLists.txt at the top level. Building from a checkout with CMake is the route the repository itself supports.

How does pugixml compare with rapidxml?

Both are small C++ XML parsers, but pugixml ships an XPath 1.0 implementation and a DOM you load a file or buffer into, while rapidxml is a header-only parser that works in situ. The README does not discuss rapidxml, so the comparison rests on pugixml's documented interface rather than on a published benchmark.

How does pugixml compare with tinyxml2?

The repository does not mention tinyxml2. The difference you can confirm from pugixml's own documentation is that pugixml includes XPath 1.0 for tree queries and automatic Unicode conversion during parsing and saving, on top of the DOM traversal API shown in the README example.

How does pugixml compare with libxml2?

pugixml describes itself as light-weight and simple, with one C++ source file and XPath 1.0 support. libxml2 is not mentioned anywhere in the repository documentation, so any claim about its feature set or performance would go beyond what the project documents.

How does pugixml compare with xerces?

The repository does not mention Xerces at all. What can be compared from pugixml's own description is its scope: a DOM interface, a parser that builds the tree from an XML file or buffer, and XPath 1.0, all in a single MIT-licensed C++ source file.

Official sources

  1. License: MIT
  2. Project website
  3. README
  4. Releases
  5. zeux/pugixml on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/zeux-pugixml.svg)](https://hysenlabs.com/projects/zeux-pugixml)