Intel PCM: reading CPU performance counters without a profiler
Intel® Performance Counter Monitor (Intel® PCM)
At a glance
- What is it?
- Intel Performance Counter Monitor is a C++ API plus a set of command line tools that read Intel core and uncore counters directly. It is for engineers who need per-socket memory bandwidth, PCIe throughput or energy numbers rather than a call graph.
- Who is it for?
- Adopt Intel PCM when you need hardware-level numbers that a sampling profiler does not expose: per-channel memory bandwidth, per-socket PCIe throughput, DRAM energy, C-state residency. Skip it if you are profiling application code paths, since it reports what the silicon is doing, not which function is slow, and skip it on AMD or ARM hosts because the counters are Intel-specific.
- Can I use it commercially?
- Yes. BSD-3-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly C++, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Intel PCM measures that a profiler does not
Intel Performance Counter Monitor is an application programming interface and a set of tools built on that API. The README states that it monitors performance and energy metrics of Intel Core, Xeon, Atom and Xeon Phi processors across Linux, Windows, FreeBSD, DragonFlyBSD and ChromeOS. That sentence is the whole scope, and it is narrower than it sounds: this is a hardware observation layer, not an application profiler. It answers questions like how many bytes per second crossed each DRAM channel, what fraction of cycles the core was stalled on memory, how much power the package drew, and whether the PCIe link is saturated.
The audience is the person who owns a machine rather than a function. Capacity planners sizing a NUMA topology, people chasing a bandwidth ceiling on a Xeon, anyone who needs to know whether a workload is memory bound or compute bound before rewriting it. The tool list reflects that: pcm-memory for per-channel and per-DRAM DIMM rank bandwidth, pcm-pcie for per-socket PCIe bandwidth, pcm-iio for per bus or device PCIe bandwidth, pcm-numa for local versus remote memory accesses, pcm-latency for L1 miss and DDR or PMM memory latency, pcm-power for sleep and energy states plus CPU frequency throttling reasons, and pcm-bw-histogram for a memory bandwidth utilization histogram. A second group targets accelerators: pcm-accel covers Intel IAA, DSA and QAT.
What it does not do is attribute time to source lines or call stacks. If your question is which function is slow, this is the wrong instrument, and the README never claims otherwise.
Core counters, uncore counters and the shared-memory daemon
The mechanism is direct counter programming. A binary called pcm-raw is described as programming arbitrary core and uncore events by specifying raw register event ID encoding, which is the honest description of what the whole suite does underneath: it writes event selectors and reads model specific registers. pcm-core and pmu-query exist to query and monitor arbitrary processor core events, so you can go past the curated tool list when you know the event encoding you want.
The uncore side is where the interesting data lives. Memory bandwidth per channel and per DIMM rank, PCIe bandwidth per socket, and the accelerator counters all come from uncore units rather than from the core PMU, which is why the numbers do not line up with what a core-only profiler reports.
Two supporting pieces make the data reachable. First, pcm-sensor-server is a collector that exposes metrics over HTTP in JSON or Prometheus exporter text format, so a scrape-based monitoring stack can consume it. Second, the README mentions a daemon that stores core, memory and QPI counters in shared memory so that non-root users can read them. That detail matters operationally: reading MSRs normally requires elevated privileges, and the daemon is the documented way to decouple collection from consumption.
There are also register access utilities that are not about monitoring at all: pcm-msr for model specific registers, pcm-pcicfg for PCI configuration registers, pcm-mmio for memory mapped registers and pcm-tpmi for TPMI registers, on Linux, Windows, FreeBSD and DragonFlyBSD. Treat those as low-level debug tools, not as measurement tools.
Building Intel PCM from source and taking a first reading
The README's build path is a recursive clone followed by CMake. The submodule flag is not optional: the repository has a .gitmodules entry and the README gives the update command separately for anyone who cloned first, so a plain clone leaves you without dependencies.
git clone --recursive https://github.com/intel/pcm
cd pcmInstall cmake and, on Linux, libasan, then choose a build method. The README distinguishes an incremental build that reuses an existing build directory from a clean build that removes it first. For a first build, the clean path is the one it spells out.
cmake -E rm -rf build && cmake -E make_directory build
cd build
cmake ..
cmake --build .The README's own snippet ends at cmake, so the build step above is the standard CMake invocation rather than a quoted command; check the CMakeLists.txt in the repository root if your generator differs. Binaries land in the build tree, and the Dockerfile copies them from build/bin, which tells you where to look.
Once built, the basic utility is pcm. Running it as root gives a live table of instructions per cycle, core frequency including Intel Turbo Boost Technology, memory and Intel Quick Path Interconnect bandwidth, local and remote memory bandwidth, cache misses, core and CPU package sleep C-state residency, thermal headroom, cache utilization, and CPU and memory energy consumption. If you want memory bandwidth broken out per channel instead, pcm-memory is the tool; if you want PCIe per socket, pcm-pcie.
The fastest route to a running instance is the published container, which the README points at a Docker how-to in doc/DOCKER_README.md. The compose file in the repository pins the image to ghcr.io/intel/pcm:latest and maps port 9738.
services:
pcm:
image: ghcr.io/intel/pcm:latest
ports:
- "9738:9738"
cap_add:
- SYS_ADMIN
- SYS_RAWIO
devices:
- /dev/cpu
- /dev/memThe image entrypoint runs pcm-sensor-server on port 9738 with the -r flag, so after starting it you should be able to scrape metrics from that port. Note what the compose file asks for: SYS_ADMIN and SYS_RAWIO capabilities, /dev/cpu and /dev/mem devices, and /sys mounted read-write. That is a lot of privilege for a monitoring container, and it is a direct consequence of the register access model.
Where Intel PCM is the wrong tool
The container recipe is the clearest limitation, and it is not a packaging oversight. Reading MSRs and uncore counters requires raw device access, so the documented deployment grants SYS_ADMIN, SYS_RAWIO, /dev/mem and /dev/cpu, and mounts /sys read-write. On a hardened cluster or a multi-tenant node, that combination is often simply not allowed, and no amount of configuration will change it.
Platform lock-in is the second constraint. The counters are Intel core and uncore events, so the tools are meaningless on AMD or ARM hosts. The README lists Linux, Windows, FreeBSD, DragonFlyBSD and ChromeOS; macOS appears in the repository topics but not in the README's operating system list, so treat macOS support as unconfirmed rather than assumed.
Third, the output is a stream of hardware counters, not an answer. Nothing in the tool list correlates a bandwidth spike with a process, a container or a request. If you need that correlation you are building it yourself on top of the JSON or Prometheus endpoint.
Fourth, counter availability varies by processor model and generation. The README never enumerates which events exist on which part, and the raw-event tools exist precisely because the curated tools cannot cover everything. Expect to discover gaps on your own silicon.
Intel PCM against perf and other counter front ends
The obvious alternative on Linux is perf, which also reads the core PMU and can sample call stacks through perf record. The difference in approach is real: perf is built around events attributed to processes and stacks, while Intel PCM is built around counters attributed to sockets, channels, DIMM ranks and PCIe devices. If your question is which code path is hot, perf is the right answer. If your question is whether the memory controller is saturated, perf will not tell you per-channel bandwidth and Intel PCM will.
A second comparison is against vendor monitoring stacks that expose similar data through their own agents. The trade-off there is reach versus control: an agent gives you a supported integration and a dashboard, while Intel PCM gives you the raw counter and the raw register tools, including pcm-raw for events nobody has wrapped yet. The cost is that you own the interpretation.
Within the project there is also a Grafana front end in the scripts/grafana directory, with its own README, and a Windows perfmon front end called pcm-service. Those are presentation layers over the same counters, so choosing between them is a dashboard decision, not a measurement decision.
Releases, licence and the cost of keeping up
Releases are dated by year and month rather than semantic version: 202604 shipped on 2026-04-10, 202509 on 2025-09-12 and 202502 on 2025-02-26. The last push to the default branch was on 2026-09-16, so the repository is being updated, but the release cadence looks like roughly twice a year. That is the upgrade cost you should budget for: pin a release, and expect to re-validate counter availability when you move to a newer processor generation, because the raw event encodings are tied to the hardware.
The licence is BSD-3-Clause, stated in the repository and repeated in the Dockerfile's SPDX header. That is permissive and imposes no copyleft obligation on your own code, but this is a description of the licence text, not legal advice, and redistribution still requires carrying the copyright notice and disclaimer. The container image ships a Fedora base and installs pci.ids from hwdata so that PCIe device names resolve; if you build your own image, that data file is part of what makes pcm-iio and pcm-pcie output readable.
Editorial conclusion
Adopt Intel PCM when you need hardware-level numbers that a sampling profiler does not expose: per-channel memory bandwidth, per-socket PCIe throughput, DRAM energy, C-state residency. Skip it if you are profiling application code paths, since it reports what the silicon is doing, not which function is slow, and skip it on AMD or ARM hosts because the counters are Intel-specific. Before trusting a rollout, verify three things on your own hardware: that the counters you want are actually exposed on your processor model, that the container path works with your kernel because the documented compose file mounts /sys read-write and passes /dev/mem, and that the release you pin (202604 is the latest listed) matches the processor generation you are measuring.
Frequently asked questions
How do I install Intel PCM?
Clone the repository recursively so submodules come with it, install cmake and libasan on Linux, then configure and build with cmake from a build directory. The README also documents a prebuilt Docker image at ghcr.io/intel/pcm:latest that runs pcm-sensor-server.
What is Intel PCM?
It is an application programming interface and a set of tools that monitor performance and energy metrics of Intel Core, Xeon, Atom and Xeon Phi processors. The tools include pcm, pcm-memory, pcm-pcie, pcm-numa and pcm-power, plus a sensor server that exposes metrics over HTTP.
Does Intel PCM work on Windows?
Yes. The README lists Windows among the supported operating systems, and the project ships a Windows perfmon front end called pcm-service. The register access utilities pcm-msr, pcm-pcicfg, pcm-mmio and pcm-tpmi are also documented for Windows.
Can Intel PCM read memory bandwidth per channel?
Yes, that is what pcm-memory is for: the README describes it as monitoring memory bandwidth per-channel and per-DRAM DIMM rank. The basic pcm utility reports aggregate memory bandwidth instead.
Does Intel PCM need root privileges?
The register-level tools read hardware counters directly, and the documented Docker deployment grants SYS_ADMIN and SYS_RAWIO and passes /dev/mem and /dev/cpu. The README also mentions a daemon that stores core, memory and QPI counters in shared memory so non-root users can access them.
Which licence does Intel PCM use?
BSD-3-Clause, stated in the repository and repeated in the SPDX header of the Dockerfile. That is permissive, but redistributing it still means carrying the copyright notice and disclaimer.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/intel-pcm)