pmu-tools treats installation as optional, because it is a source tree you run from
Intel PMU profiling tools
At a glance
- What is it?
- The readme's first section is about not installing anything: clone the repository, run the tool from the source directory, and let it find its own dependencies. That is a deliberate choice for a tool whose primary consumer is the machine's owner debugging that machine, and it comes with a real cost: the event lists are fetched from a network repository on first run.
- Who is it for?
- Adopt pmu-tools if you are profiling on Intel silicon and raw perf output is not telling you why your code is slow, because the wrapper gives you the full vendor event list under symbolic names and the top-down tool turns counters into a bottleneck, which is a different kind of answer. Do not adopt it if your processor is not Intel, since every tool here is built around the vendor's performance monitoring unit.
- Can I use it commercially?
- Yes, with conditions. GPL-2.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
- Is it still maintained?
- Yes. The repository last received commits 68 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The first instruction is to not install it
The packaged install is one command, marked experimental in the readme:
pip install pmu-toolsThe installation section is titled a quick non-installation and says so in its first line: the tools do not really need to be installed, and it is enough to clone the repository and run the relevant tool out of the source directory. Then there are two optional steps for people who want the tools on their path, either exporting it or symlinking a tool into a local binary directory, and a sentence that matters more than either: the tools automatically find their Python dependencies. That last sentence describes a self-locating dependency scheme, which is a pattern you rarely see outside of this author's projects and which explains the non-installation. There is also an alternative route, a package index install, and the readme marks it as experimental so far. So the supported path is a source tree, and the packaged path exists for convenience. That is a defensible choice for a tool whose users are people debugging a specific machine and who want to be on the latest commit because a new processor model was added yesterday. It also means there is no upgrade process: you pull and run. The readme is explicit that the majority of the tools need no third-party Python packages at all and run in what it calls an included-batteries-only mode, with plot and spreadsheet output being the exceptions, and only those requiring an install of a small requirements file. A small dependency list of four packages, all of them for visualisation, is consistent with that.
First run needs the network, and that is the design's one dependency
There is a single external dependency and the readme is upfront about it. On first run, the two main tools automatically download the vendor's event lists from a separate repository on a code-hosting site, and that requires working internet access. Later runs can be done offline, and the readme links a wiki page explaining how to download the lists ahead of time, which tells you the offline path is supported but is manual. The reason for the download is the right one. The kernel's perf tool knows a small set of generic events, because the event encodings are processor-specific and change with every microarchitecture. The full list of what a given processor can count, with the symbolic names the vendor documents them under, lives in data that has to come from somewhere. Downloading it at runtime rather than vendoring it means the repository does not have to be updated every time a new part ships, and the recent-changes list shows the author adding models for particular processors as they launch. The cost is that first run is not hermetic, and a profiling run on an air-gapped machine needs a manual step first. The tool also depends on a reasonably current version of the kernel's perf, and the readme says that depending on the processor an up-to-date kernel may be required, with a wiki page on kernel support. That is a hard floor: a tool that reads hardware counters through the kernel cannot work below the kernel version that exposes them.
Seven tools, and the readme tells you which one you want
The most useful section in the readme is a list headed by what you want to do, and it maps seven tasks onto seven tools. To understand processor bottlenecks at a high level, use the top-down tool. To see that output graphically, use its spreadsheet or graph output. To know which events to run but use symbolic names for a processor you are adding support for, use the event translation wrapper. To measure interconnect, caches, memory and power management on server parts, use the uncore tool or the top-down tool. To use events from a C program, use the C library. To query processor topology or disable hardware threading, use the topology tool. To change model specific registers, use the register tool. To change PCI configuration space, use the configuration tool. That list is a specification of the project more accurate than any feature list would be, and it also shows the shape of the whole thing. There is a wrapper that gives the kernel's perf a full event list, a top-down analysis tool that turns counters into a bottleneck, an uncore tool for the parts of the processor that are not the core, a C library for programmatic access, a statistics clone, and a set of small utilities. It is a toolbox for one vendor's hardware, and the readme is honest that some tools are obsolete.
The top-down tool is the substantial piece, and multiplexing is its weakness
The top-down tool is described as implementing a named analysis methodology, and the readme links two descriptions of it, one from a research site and one from the vendor. Top-down analysis is a specific idea: instead of collecting every event and hoping you notice something, you measure a small number of top-level metrics, and the slowest one tells you which category of problem you have, and you descend only into that category. The readme's examples show the interface, and they are worth reading for the flags. A level number controls how much detail you get. A single-thread option exists, with a note that on a specific older processor generation with hardware threading the system should be idle, which is a real caveat about measurement rather than a stylistic one. A consolidated view exists, with an explicit warning that it will multiplex performance counters and therefore may produce measurement errors. A sampling option exists that runs a second pass to find where the bottleneck is located. And a drilldown option exists that reruns the workload with minimal multiplexing until the critical bottleneck is found, printing only that one. That last flag is the honest answer to the multiplexing problem: if you want accurate numbers you have to measure fewer events and run more passes. The recent-changes list shows the tool being kept current with a methodology version and with new processor models, which is what maintaining a hardware model actually involves.
A repository with one file per processor generation, and obsolete tools kept on purpose
The top-level listing is the most unusual thing about this repository and it explains the maintenance burden. Dozens of files are named after processor families and their client or server variants, one per generation, and each appears to hold the ratios that the top-down model needs for that part. Then a handful of JSON files named after newer parts with a suffix suggesting latency data, and a shell script and a Python file for processor information. So when a new processor launches, the author adds a file. That is a maintenance model with no upper bound, and it is the kind of commitment that explains why the readme's recent-changes list is dominated by model updates rather than new features. The same listing shows the smaller tools as individual scripts rather than a package layout, which is consistent with the non-installation approach. Two entries in the listing are worth noting for what they say about testing. There is a makefile whose default target echoes that there is nothing to compile, and whose real targets generate documentation from a help tool and render model diagrams from a graph description through a graphviz renderer. And there are two files that look like test harnesses, one named for a test runner and one for a general tester. Three continuous integration badges, one for current Python, one for older Python, and one specifically for the C library, confirm the project runs both interpreter versions and tests the compiled component separately. The readme also states plainly that some obsolete tools are kept as programming reference and may need updates to build on newer kernels, which is a candid admission of dead weight.
Editorial conclusion
Adopt pmu-tools if you are profiling on Intel silicon and raw perf output is not telling you why your code is slow, because the wrapper gives you the full vendor event list under symbolic names and the top-down tool turns counters into a bottleneck, which is a different kind of answer. Do not adopt it if your processor is not Intel, since every tool here is built around the vendor's performance monitoring unit. Four things to verify. That your processor generation is one the tool has a model for, because the readme's recent changes are all about adding and updating models for particular server and client parts, and an unrecognised processor is a degraded experience rather than an error. That your kernel exposes the events, since the readme says a current kernel may be needed depending on the processor and links a page on kernel support. That you can reach the event repository once, because the first run downloads the vendor's event lists and later runs work offline only if you have them. And how you pin the version, since the version string is a date with a patch component, which is honest for a project that tracks hardware, and the readme documents no release tags to install from. The licence is GPL-2.0 and the last push was on 2026-07-24.
Frequently asked questions
Do I need to install pmu-tools?
No. The readme says it is enough to clone the repository and run the tool from the source directory, and that the tools find their Python dependencies themselves. There is a package index install as an alternative, which the readme marks as experimental. Most tools need no third-party Python packages; plotting and spreadsheet output do.
Which tool should I use in pmu-tools?
The readme maps seven tasks to seven tools: high-level bottleneck analysis to the top-down tool, graphical output to its spreadsheet or graph flags, symbolic event names for a new processor to the event wrapper, uncore and power metrics to the uncore tool, event access from a C program to the C library, topology queries and disabling hardware threading to the topology tool, and register or PCI configuration access to the respective tools.
What does the pmu-tools toplev tool do?
It identifies the micro-architectural bottleneck for a workload by implementing a top-down analysis methodology, where a small set of top-level metrics tells you the category of problem and you descend only into that. Options control detail level, thread scope, consolidated view, automatic sampling, and a drilldown mode that reruns with minimal counter multiplexing to find only the critical bottleneck.
Does pmu-tools need internet access?
On first run only. The two main tools automatically download the processor event lists from a separate repository on a code-hosting site, which requires working internet access. Later runs work offline, and the readme links a wiki page explaining how to download the lists in advance. The tools also need a reasonably current perf, and depending on the processor an up-to-date kernel.
What licence is pmu-tools released under?
GPL-2.0-only, with the licence file referenced from the package metadata. The version is date-based with a patch component, the project requires Python 3.9 or later for the packaged form, and the last push to the master branch was on 2026-07-24.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/andikleen-pmu-tools)