# microsoft/retina: eBPF Network Observability for Kubernetes

> Retina is Microsoft's MIT-licensed Kubernetes networking observability tool built on eBPF. It ships two things: Prometheus metrics from a set of Linux plugins, and on-demand packet captures driven by a CLI or a CRD.

**microsoft/retina** — eBPF distributed networking observability tool for Kubernetes

- Repository: https://github.com/microsoft/retina
- Website: https://retina.sh
- Stars: 3,176 · Forks: 304
- Language: Go
- License: MIT
- Published: 2026-09-24 · Updated: 2026-09-24 · Language: en
- Canonical page: https://hysenlabs.com/projects/microsoft-retina

## The gap Retina fills between cluster state and packets

Kubernetes tells you a pod is Running and a Service has endpoints. It does not tell you that DNS replies for one workload are being dropped, or that traffic to a specific backend is being forwarded somewhere unexpected. The usual answers are tcpdump on a node, a service mesh sidecar, or a CNI-specific flow log. Each has a cost: tcpdump is manual and node-local, a mesh changes your data path, and CNI flow logs are tied to one vendor.

Retina targets cluster network administrators, cluster security administrators and DevOps engineers who need to answer those questions on demand, and who also want a continuous stream of network metrics. The README frames it as a centralized hub for application health, network health and security, with telemetry exported to Prometheus, Azure Monitor or other vendors and visualized in Grafana, Azure Log Analytics or elsewhere. The distinguishing choice is that collection happens in eBPF on the node rather than in a sidecar, so a workload does not have to be restarted or instrumented to become observable.

## Two products in one repository: metrics plugins and captures

The repository splits into two top-level capabilities the README names explicitly: Metrics and Captures. They share the agent but differ in what they produce and how you drive them.

Metrics run continuously. The Helm chart takes an enabledPlugin_linux value, and the README's metrics example passes dropreason, packetforward, linuxutil and dns. Those plugin names are the unit of configuration: you choose which eBPF programs load, and the output is Prometheus-format metrics. The capture example in the same README adds packetparser to that list, which is the plugin that inspects packet contents. That is why the known-limitations note about high-core-count nodes names packetparser specifically: the deep-inspection path costs more than the counters.

Captures are on-demand. You ask for traffic matching a selector and get a capture back. Two front ends exist. The CLI is a kubectl plugin installed through Krew, and the CRD path uses the operator, which the capture Helm example enables with operator.enabled=true and operator.enableRetinaEndpoint=true. The CRD route is the one to pick if captures should be created by a controller or a GitOps pipeline rather than by a human at a terminal.

The Go module confirms the shape of the codebase: cobra for the CLI, prometheus/client_golang for metric exposition, zap for logging, and k8s.io/client-go for cluster access. The azure-sdk-for-go and cloud-provider-azure dependencies are present as indirect requirements, which fits the cloud-agnostic claim rather than contradicting it.

## Installing Retina from the OCI Helm chart and taking a first capture

Retina's chart is published to GHCR as an OCI artifact, not to a classic HTTP chart repository. The README's quick install resolves the latest release name through the GitHub API and passes it as both the chart version and the image tags. The enabledPlugin_linux value is escaped with backslashes because the comma-separated list is passed through Helm's --set parser.

```bash
VERSION=$( curl -sL https://api.github.com/repos/microsoft/retina/releases/latest | jq -r .name)
helm upgrade --install retina oci://ghcr.io/microsoft/retina/charts/retina \
    --version $VERSION \
    --set image.tag=$VERSION \
    --set operator.tag=$VERSION \
    --set logLevel=info \
    --set enabledPlugin_linux="\[dropreason\,packetforward\,linuxutil\,dns\]"
```

After the install, the README directs you to separate pages for setting up Prometheus and Grafana. Retina does not bundle a visualization layer, so the metrics are only useful once something scrapes them. If you are trying Retina to see whether it fits, that is two more installs before you see a single graph.

The capture path is faster to evaluate. Install the CLI as a kubectl plugin through Krew:

```bash
kubectl krew install retina
```

Running the bare command prints the plugin's help, which the README reproduces: the subcommands are capture, completion, config, help, shell, trace and version, and the banner line reads "Retina is an eBPF distributed networking observability tool for Kubernetes." If you see that output, the plugin is on your PATH correctly. The shell subcommand is marked EXPERIMENTAL, so treat it as a debugging aid rather than a supported workflow.

To create a capture against a workload, the README gives a single selector-based command:

```bash
kubectl retina capture create --pod-selectors <app=my-app>
```

Replace the placeholder with your own label selector. The README points to the CLI capture documentation for the full set of options, and the repository carries example manifests under samples/capture/ plus a samples/metricsconfiguration.yaml, which is where to look for a working configuration rather than reconstructing one from the flag list.

## Where Retina is the wrong tool

The README's own known-limitations section is the most useful part of the documentation for a purchasing decision, because it names two cases where Retina will disappoint.

The first is high-core-count nodes. Community users have reported performance considerations when Advanced metrics with the packetparser plugin run on nodes with 32 or more CPU cores under high network load. The README's suggested mitigation is to start with Basic metrics mode on large node types. Read that as a design trade-off rather than a bug: packet parsing per packet per core does not scale the way a counter does, and the project is telling you to size the plugin set to the node. If your cluster is built from large instances and you wanted packetparser everywhere, you should plan to run it on a subset of nodes or accept the cost.

The second is Windows Server 2019. Retina no longer supports those nodes and the README directs Windows workloads to Windows Server 2022. If part of your fleet is still on 2019, Retina will not cover it, and a partial rollout across a mixed Windows estate is a worse outcome than not adopting it yet.

There is a third constraint that is structural rather than stated as a limitation. Because collection is eBPF-based and the plugin list is Linux-specific (enabledPlugin_linux), the metrics capability is a Linux story. The capture Helm example sets os.windows=true, so Windows nodes are in scope for the capture path, but do not assume plugin parity across operating systems.

Finally, consider whether you need Retina at all. If your only question is "which pod is talking to which service", a CNI that already emits flow logs may answer it with no new agent. Retina earns its place when you need packet contents or drop reasons at the node level.

## How Retina differs from Cilium Hubble and from a service mesh

The obvious comparison is Cilium Hubble, which also uses eBPF for Kubernetes network visibility. The difference is where the eBPF programs live. Hubble's observability is part of the Cilium CNI: adopting it means adopting Cilium as your data plane, and the visibility you get is visibility into Cilium's own datapath. Retina is described as cloud-agnostic and is installed as its own Helm release, so it can sit alongside a CNI you already run rather than replacing it. That is the trade: you take on a second node-level agent, and in exchange you are not committing your networking stack to get observability.

The other comparison is a service mesh. A mesh gives you L7 telemetry, mutual TLS identity and traffic policy in one package, but it does so by putting a proxy in the request path, which means per-pod sidecars or a per-node proxy and a change to how traffic flows. Retina collects at the kernel and does not proxy. You get node-level facts, including packets that never reach an application, but you do not get request-level tracing or identity-based policy. Teams that need both end up running both, and the honest question is whether the mesh's telemetry already covers the incidents you actually have.

## Maintenance, release cadence and what the MIT licence leaves you to do

The repository is not archived and the last push was on 2026-09-23. Releases have been frequent: v1.2.6 on 2026-08-21, v1.2.7 on 2026-08-28, and v1.2.8 on 2026-09-10, roughly weekly patch cadence across that window. The Go module targets go 1.26.0 and depends on k8s.io/client-go v0.35.4, so keeping up with Retina means keeping up with a recent Kubernetes client library, which in practice means a recent Go toolchain if you build from source.

Upgrade cost is mostly the Helm values. The chart takes image.tag, operator.tag and image.pullPolicy separately, so a version bump is three coordinated values rather than one. The retract directives in go.mod (v0.10.0 published accidentally, v0.10.1 containing retractions only) are a reminder that the module is consumed by other Go code and that version selection matters if you import it.

On supply chain: the README states that Retina images published to GHCR are cryptographically signed and shows how to verify provenance with sigstore/cosign, using the retina-operator image as the example. If your policy requires verified images, that step is documented rather than left to you.

The licence is MIT. That is permissive: it allows commercial use and modification with the copyright notice and permission notice retained, and it does not carry the patent grant or the copyleft obligations of other licences. It also means there is no warranty. Contributors sign a Microsoft CLA, and the project follows the Microsoft Open Source Code of Conduct, which matters if you plan to send patches upstream. None of this is legal advice; check the LICENSE file and your own policy.

## Conclusion

Adopt Retina if you run Kubernetes on Linux nodes and want packet-level answers without installing a sidecar per pod, and if you already have Prometheus and Grafana to receive the metrics. Skip it if your fleet is Windows Server 2019, since the README states that node version is no longer supported, or if your nodes have 32 or more CPU cores and you intend to run Advanced metrics with the packetparser plugin under high load, where community users have reported performance considerations. Before rolling it out cluster-wide, verify three things: that your kernel and OS combination is covered by the supported platforms list, that the enabledPlugin_linux value you pass matches the plugins you actually need, and that capture artifacts land somewhere your team can retrieve them.

## FAQ

### What is microsoft/retina?

It is an open-source, eBPF-based network observability platform for Kubernetes, described in the README as a centralized hub for monitoring application health, network health and security. It has two capabilities: continuously exported Prometheus metrics and on-demand packet captures.

### How do I install microsoft/retina?

The README installs it with Helm from the OCI chart at oci://ghcr.io/microsoft/retina/charts/retina, passing a version and an enabledPlugin_linux list. The capture CLI is installed separately as a kubectl plugin with kubectl krew install retina.

### Does microsoft/retina support Windows nodes?

The README states that Windows Server 2019 is no longer supported and directs Windows workloads to Windows Server 2022. The capture Helm example sets os.windows=true, while the metrics plugin list is configured through enabledPlugin_linux.

### What are the known performance limitations of microsoft/retina?

The README notes that community users reported performance considerations when Advanced metrics with the packetparser plugin run on nodes with 32 or more CPU cores under high network load, and suggests starting with Basic metrics mode on large node types.

### What do I need before microsoft/retina metrics are useful?

The README's install guide points to separate steps for setting up Prometheus and Grafana after the Helm install. Retina exports Prometheus-format metrics to storage such as Prometheus or Azure Monitor, and visualization happens in Grafana, Azure Log Analytics or other vendors.

## Sources

- [License: MIT](https://github.com/microsoft/retina/blob/main/LICENSE)
- [microsoft/retina on GitHub](https://github.com/microsoft/retina)
- [Project website](https://retina.sh)
- [README](https://github.com/microsoft/retina/blob/main/README.md)
- [Releases](https://github.com/microsoft/retina/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/microsoft-retina
