# ccfos/huatuo: eBPF Kernel Observability with AutoTracing and Continuous Profiling

> Huatuo is a Didi open source project that uses eBPF, kprobe and tracepoint hooks to collect kernel-level metrics and capture runtime context on slow paths. Here is what the repository actually documents, and where it stops.

**ccfos/huatuo** — eBPF-based Linux kernel observability 🚀🚀

- Repository: https://github.com/ccfos/huatuo
- Website: https://huatuo.tech/
- Stars: 1,157 · Forks: 149
- Language: Go
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/ccfos-huatuo

## The gap Huatuo targets: kernel evidence during incidents

Most observability stacks start at the application. Traces come from instrumented code, metrics come from exporters that read /proc or a runtime's own counters, and logs come from whatever the process chose to print. When a node stalls because of memory reclaim, scheduler delay or block I/O contention, that layer has little to say. The process is not slow because of its own logic; it is slow because the kernel is not scheduling it.

Huatuo is built for that layer. The README describes it as a cloud-native operating system observability project open-sourced by Didi and incubated under the CCF, delivering kernel-level observability for general-purpose cloud-native computing, AI computing and bare-metal infrastructure. The intended audience is infrastructure and SRE teams running Linux fleets, including Kubernetes nodes, who need to explain CPU idle drops, CPU sys spikes, I/O surges and Loadavg spikes that application telemetry cannot attribute.

The project states it is deployed at scale in Didi's production environment. That is a claim in the README, not something verifiable from the repository, but it explains the shape of the tool: it is designed for operators who own the host, not for developers who only own a service.

## How the eBPF probes, AutoTracing and the metrics endpoint fit together

The mechanism is kernel dynamic tracing. The README lists kprobe, tracepoint, ftrace and eBPF as the underlying technologies, and the repository layout confirms the split: a bpf/ directory holds the C programs, core/ and internal/ hold the Go runtime, cmd/ holds the binaries, and apis/ holds generated API types. The Makefile compiles the BPF sources with build/clang.sh and passes -I bpf/include, so the probes are built from source rather than shipped as opaque blobs.

The README claims BPF keeps performance overhead below 1%, which is a project claim rather than a measured figure. The design that supports it is selective instrumentation: rather than tracing everything, Huatuo instruments kernel slow paths and triggers on events such as page faults and scheduling delays. That is the AutoTracing idea. Instead of continuously sampling, the agent captures runtime context when a slow path fires, which is what makes the output useful for a jitter that lasts milliseconds.

On the collection side, the agent exposes a Prometheus endpoint. The README's quick run section pulls from localhost:19704/metrics, and the go.mod file shows prometheus/client_golang, the Elasticsearch v8 client, Pyroscope, cadvisor, containerd cgroups v3 and the Cilium eBPF library as dependencies. That matches the ecosystem section: Prometheus, Grafana, Pyroscope and Elasticsearch, with automatic association of Kubernetes container labels and annotations.

There is also a public Go client in client/, with node client files prefixed node_. The README shows it being used to push configuration, including a Runtime.CPULimitCores key, to a node agent over an API. That is a control-plane surface, not just a read-only exporter.

## Installing Huatuo with Docker and reading the first metrics

The fastest path in the README is a privileged container. It needs host PID and cgroup namespaces, host networking, and mounts for /sys and /run, because kernel tracing requires visibility into the host rather than the container's own namespace. Run it on a host whose kernel is 4.18 or later.

```bash
docker run --privileged --pid=host --cgroupns=host --network=host -v /sys:/sys -v /run:/run huatuo/huatuo-bamai:latest
```

The README attaches an explicit warning to that command: do not deploy images with the latest tag to production, because it is a development and testing image. Use a formal release image or binary instead. The most recent release listed in the repository is v2.3.0 from 2026-08-13, so that is the tag to pin.

From another terminal, pull the metrics endpoint. What you should see is a Prometheus text exposition, the same format any scraper expects.

```bash
curl -s localhost:19704/metrics
```

If you want the full stack rather than the agent alone, the README points at a Compose file under build/docker that brings up Elasticsearch, Prometheus, Grafana and Huatuo together.

```bash
docker compose --project-directory ./build/docker up
```

Once that is running, the README says the monitoring dashboard is at http://localhost:3000. The screenshots in docs/img show a quickstart components view and an AutoTracing event view, which is where the captured slow-path context surfaces.

The Dockerfile is worth reading before you build your own image. It installs clang, libbpf-dev, bpftool, musl-tools and capnproto in the build stage, then produces either a static Alpine-based runtime image or a nostatic image based on golang:1.24 that carries libelf1 and libnuma1 and sets LD_LIBRARY_PATH to include Ascend driver paths. That split exists because the static build cannot carry the shared libraries some hardware monitoring paths need. The Dockerfile also rewrites the shipped configuration, blanking the Address field and appending KubeletReadOnlyPort=0 and KubeletAuthorizedPort=0, so the default container does not talk to a kubelet.

## Where Huatuo is the wrong tool, and what the docs do not settle

The first constraint is the kernel. The README states support for kernel 4.18 and later and gives a tested distribution table, but the table is anchored to version 1.0.0 for every row except the last, which lists 2.3.0 against kernel 7.0.x on Ubuntu 26.04. There is no published compatibility matrix for 2.1.0, 2.2.0 or 2.3.0 against the older kernels in that table. If you are running 5.10 on OpenEuler 22.03 with a 2.3.0 binary, the repository does not tell you whether that combination is tested. Verify it yourself on a canary node before trusting it.

Second, this is a host-level agent. It needs --privileged, host PID, host cgroup namespace and host networking. On managed Kubernetes, on locked-down clusters with Pod Security Admission at the restricted level, or anywhere a privileged DaemonSet is not acceptable, Huatuo will not deploy. That is not a bug; kernel tracing cannot work without those privileges. But it means the tool is unavailable to a large class of teams by policy alone.

Third, if your problem is a slow database query or a misconfigured HTTP client, kernel data will not help. Huatuo observes the operating system. It does not replace distributed tracing inside your application, and the README makes no claim that it does.

On documentation, the README is a landing page, not a manual. It links to docs.huatuo.tech and a quick-start page, and the repository ships huatuo-bamai.conf and huatuo-apiserver.conf as the configuration surfaces. The README does not document the full set of config keys, does not describe rollback or upgrade procedures, and does not state what happens when the agent is removed from a running node. The Go client example shows one key, Runtime.CPULimitCores, and nothing about how the rest are named.

## Huatuo against Pyroscope and continuous profilers

The clearest comparison is with Pyroscope, which Huatuo lists as an ecosystem integration and depends on in go.mod. The two collect different things. Pyroscope is a continuous profiler: it samples stacks inside a process, whether through language SDKs or eBPF, and answers which functions are consuming CPU or allocating memory. Huatuo's continuous profiling is described as covering the operating system and applications across CPU, Memory, I/O and Locks, and its distinctive feature is event-driven context capture rather than a continuous stack sample.

The practical difference shows up in the question you are asking. If you want to know which Go function in your service burns the most CPU over an hour, a process-level profiler is the direct answer, and Huatuo's kernel view is a detour. If you want to know why a node's CPU idle dropped at 03:14 and what the scheduler was doing at that moment, a stack profiler has no access to that and Huatuo's slow-path capture does. The README positions AutoTracing specifically against the jitter class of problems: CPU idle drops, CPU sys spikes, I/O surges, Loadavg spikes.

The honest reading is that these are complementary, and Huatuo's own dependency list says so. It integrates with Pyroscope rather than replacing it. A team already running Pyroscope with good coverage should treat Huatuo as an addition for host-level questions, not a migration target.

## Maintenance, release cadence and the Apache-2.0 terms

The repository is not archived, and the last push was on 2026-09-10, ten days before this writing. That is a recent push, and the release history supports it: v2.1.0 landed on 2025-11-03, v2.2.0 on 2026-03-30, and v2.3.0 on 2026-08-13. Roughly two releases a year, with the most recent one about a month before the last commit. The cadence is steady rather than fast, so plan upgrades around release tags instead of expecting continuous drift on main.

Upgrade cost is dominated by two things. The first is the BPF compilation toolchain: the Dockerfile pins golang:1.24 and installs clang, libbpf-dev and bpftool, and the Makefile drives the BPF build through build/clang.sh with BPF_DEBUG defaulting to 0. Building from source means matching that toolchain, and a BPF_DEBUG=1 build changes what the probes emit, so debug artifacts should not be what you ship. The second is the configuration files, huatuo-bamai.conf and huatuo-apiserver.conf. The Dockerfile edits the config during image build, which means an in-place upgrade can silently reset keys you set. Diff your config against the new image's default before restarting.

Licensing is Apache-2.0, per the LICENSE file and the badge in the README. That permits commercial use, modification and redistribution with the usual conditions around notices and patent grant. The repository vendors its dependencies under vendor/, so if you redistribute a binary you are also carrying third-party licences. This is a factual description of the licence identifier, not legal advice; have counsel review anything you ship.

## Conclusion

Adopt Huatuo if you operate Linux fleets on kernel 4.18 or later and already have Prometheus, Grafana, Pyroscope or Elasticsearch in place, and if kernel-level evidence is what your incident reviews are missing. Skip it if you only need application-level tracing, if you cannot run a privileged container with host PID and cgroup namespaces, or if your kernels sit outside the tested distribution table. Before rollout, confirm the config keys in huatuo-bamai.conf against the documentation, pin a release tag rather than latest, and check whether your kernel is one of the versions the project lists as primarily tested.

## FAQ

### What is Huatuo and what does it observe?

Huatuo is a cloud-native operating system observability project open-sourced by Didi and incubated under the CCF. It uses kprobe, tracepoint, ftrace and eBPF to collect kernel-level metrics and capture runtime context from kernel slow paths.

### Which kernel versions does Huatuo support?

The README states support for kernel 4.18 and later, with a table listing primarily tested combinations such as 4.18.x on CentOS 8.x, 5.15.x on Ubuntu 22.04, and 6.8.x on Ubuntu 24.04. The table only maps 2.3.0 to kernel 7.0.x on Ubuntu 26.04; the other rows are anchored to version 1.0.0.

### How do I run Huatuo quickly?

The README gives a privileged Docker command with host PID, host cgroup namespace, host networking and mounts for /sys and /run, and then a curl against localhost:19704/metrics. It explicitly warns against using the latest tag in production.

### What is Huatuo's licence?

The repository is licensed under Apache-2.0, as stated in the LICENSE file and the badge in the README.

### Can Huatuo be used without a privileged container?

The README's quick run command uses --privileged together with --pid=host and --cgroupns=host, and the documentation gives no unprivileged alternative. Kernel tracing requires that level of host access.

## Sources

- [ccfos/huatuo on GitHub](https://github.com/ccfos/huatuo)
- [License: Apache-2.0](https://github.com/ccfos/huatuo/blob/main/LICENSE)
- [Project website](https://huatuo.tech/)
- [README](https://github.com/ccfos/huatuo/blob/main/README.md)
- [Releases](https://github.com/ccfos/huatuo/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/ccfos-huatuo
