# Akvorado's data path is inlet, Kafka, ClickHouse, and four things it enriches on the way

> The flow collector from a French ISP is a good architecture lesson because its stages are named after what they are for: an inlet that speaks the flow protocols, a Kafka hop that lets each stage scale and be upgraded independently, an outlet that enriches with interface names and geolocation, and a console that queries. The readme's one-line warning to read the changelog before upgrading is the honest signal.

**akvorado/akvorado** — Flow collector, enricher and visualizer

- Repository: https://github.com/akvorado/akvorado
- Website: https://demo.akvorado.net
- Stars: 2,368 · Forks: 185
- Language: Go
- License: AGPL-3.0
- Published: 2026-09-30 · Updated: 2026-09-30 · Language: en
- Canonical page: https://hysenlabs.com/projects/akvorado-akvorado

## Four stages named for what they do, with a broker in the middle

The top-level directories are the architecture, and they are named for function rather than for layer. There is an inlet, an outlet, an orchestrator, a console, a common package, a command directory, a configuration directory, and a demo exporter. The inlet receives flows, the outlet enriches and exports them, and the broker between them is what makes the split worthwhile. The readme says the program receives flows, currently NetFlow and IPFIX and sFlow, enriches them with interface names using SNMP and with geolocation information from a commercial IP database, and exports them to ClickHouse, then exposes a web interface to browse what it collected. A broker between collection and enrichment is the decision that makes this a platform rather than a script. Collection and enrichment have different scaling characteristics: flow volume is driven by how many devices export and how chatty they are, while enrichment load is driven by how many distinct interfaces and addresses you care about, and the second is usually far smaller and burstier. Decoupling them means a slow geolocation lookup or an unreachable SNMP target cannot back up your flow ingestion. It also means you can upgrade the console without touching the collector, and the readme's own caution about reading the changelog before upgrading is a reminder that a staged system has more places for a breaking change to land.

## Enrichment is where the operational decisions actually are

Two enrichment sources are named and both carry consequences. Interface names come from SNMP, which means the collector polls your network devices and has to be given credentials that can read interface tables, and that a device which stops answering produces flows with missing names rather than an error. Geolocation defaults to a commercial database, which means your address space is being resolved against somebody's data and, depending on the configuration, IP ranges are being sent to them. That second point deserves a sentence of caution for anyone deploying this on a network they do not own or are not authorised to monitor, and it is one of the reasons the topic list and the readme both treat geolocation as a configurable default rather than a hardcoded one. The dependency list confirms the breadth of what the outlet does beyond those two: a routing information library and a BGP library for address-to-ASN mapping, a gNMI client and a NETCONF library for device configuration, a system information package for polling device metadata, a maximum-geometry database reader, and two SNMP implementations, one of them a server rather than a client. Then there is a SQL parser for the clickhouse driver and several database drivers. In other words the outlet can learn about your network from a routing table, from BGP, from gNMI, and from SNMP, and you should decide how much of that you want switched on before you deploy it.

## The console is a single-page app with its own build pipeline

The console is not a set of server-rendered pages, and the build system says so. The makefile lists generated JavaScript artefacts including a frontend node_modules directory and a built frontend bundle, and the build rules mention a JavaScript package manager invoked through a wrapper binary in the repository. There is also a syntax highlighter in the dependencies, which is what you use when a user types a query into a filter box and you highlight it. That is the tell that this is a query console: you write filter expressions rather than clicking through pre-built reports, and the console ships both a timeseries graph and a Sankey diagram as its two headline visualisations. The Sankey view is worth noting because it is the right shape for flow data, where you want to see volume moving between addresses, protocols and interfaces rather than a line over time. The generated-file list in the build is also unusually long and informative, and reading it is a map of the codebase: a generated protobuf file for raw flows, a generated schema definition, a compiled object for packet socket reuse in the UDP input for two byte orders, generated comma-separated data for autonomous system numbers, protocols, and transport ports, a generated filter parser, and a series of generated enumerations for every configuration option that is a fixed set of values. The console is therefore a typed, validated configuration surface with code generation behind it, not a config file you hand-edit.

## A profile-guided optimisation file in the root, and what it implies

One entry in the top-level listing deserves explanation because it says something about the performance work. There is a default profile file committed at the root, which is Go's profile-guided optimisation artefact. A PGO file is produced by running a representative workload, recording which functions the compiler should favour, and committing the profile so every subsequent build benefits. For a data plane that processes a firehose of flow records, that is a real win and a real commitment: someone had to run the software under load to produce it, and it has to be refreshed when the hot paths change or it becomes quietly counterproductive. It is a small file that signals a mature performance culture. Other top-level entries say the same kind of thing. There is a Nix flake with a lock file, so the build environment is reproducible. There is a linter configuration for a Go linter, an .env and a .env.demo, a GitLab CI file alongside the GitHub Actions workflow referenced by the badge, a licence header checker configuration, and a generated code directory. A project that keeps a demo environment file, a Nix lock and a header checker is one where reproducibility and licence hygiene are somebody's job.

## What the dependency list says about the scale of the thing

The go module file is worth reading as a requirements document. It declares a recent Go language version and pins a toolchain version, and then lists a long set of dependencies that map almost one to one onto the feature list. A native ClickHouse client and a SQL parser, a column-oriented driver library, and a Postgres driver with a pool, plus a MySQL driver, tell you that ClickHouse is the primary store and the others exist for remote data sources. A Kafka client and a topic administration library, plus a fake broker implementation used in tests, tell you the broker is first class. An expression evaluation library, a bit set library, a binary radix trie for prefixes, and a bitset tell you the filter and routing-lookup paths are performance sensitive. A backoff library and a resilience library tell you there is retry logic in front of unreliable dependencies. A command line framework, a structured logger, a Prometheus client and a gRPC middleware for metrics tell you the operational surface. A clock abstraction in the dependencies is a small and good sign, because it means time-dependent logic is testable. And an eBPF library alongside a Go packet library, with the compiled reuse-port objects in the build, tell you the inlet has a fast path for flow UDP that bypasses the ordinary socket path. That is a collector written by people who have profiled one.

## Licence, help, and the demo that runs on fake data

The readme's help section tells you how to report a problem, and the first line of the template is a version stamp:

```bash
# akvorado version | head -2
```

Two things to understand before you deploy. First, the licence is AGPL-3.0 and the project states it was developed by a French internet provider, which is the reason a network operator's tool is available at all. The AGPL is a real constraint in a specific way: if you modify Akvorado and let others interact with it over a network, you owe them the source of your modifications. For an internal operations console that nobody outside your organisation reaches, this is usually not a problem. For anything customer-facing, it is a conversation with whoever owns your licensing. Second, the demo. The readme offers a live site and is unusually clear about what it is: the result of running a container composition on a fresh checkout, running the latest stable version, using fake data, with the flow ports deliberately not accessible so you cannot send your own flows, and a request to be gentle with the resource. That clarity is worth something, because a demo that quietly accepted your production flow export would be a different kind of service. It also means the demo is the fastest way to see the console without installing anything, and that the documentation is browsable there, with a distinction the readme makes explicit: the served documentation is the version you are running, while the documentation in the repository is for the next version. If you are evaluating a release, read the version matching what you would install.

## Conclusion

Adopt Akvorado if you run a network and want flow data you can query yourself, with the readme's condition attached: read the changelog before every upgrade, because the project says so in a callout and the release numbering shows why, with calendar-style versions that reset their minor series between majors. Do not adopt it if your requirement is a hosted or turnkey analytics product, since the demo instance runs fake data and deliberately does not accept your flows. Four things to verify. That you have somewhere to put the data, because ClickHouse is not optional and this is the component that constrains your deployment. That your exporters can reach the inlet, since flow protocols are UDP and the readme's help text shows a pipeline describing an interface loop. That your enrichment sources are acceptable, since interface names come from SNMP polling and geolocation defaults to a third-party commercial database, which is a decision about sending IP space to somebody. And that the AGPL-3.0 licence works for you, because a network operations console modified and served to others is exactly the case that licence was written for. Version 2026.8.1 was released on 2026-08-29 and the last push was on 2026-09-28.

## FAQ

### What flow protocols does Akvorado collect?

NetFlow, IPFIX and sFlow. The collected flows are enriched with interface names using SNMP and with geolocation information, which by default comes from the IPinfo.io databases, and then exported to ClickHouse and browsed through a web interface.

### What are the main components of Akvorado?

The repository is organised as an inlet that receives flows, an outlet that enriches and exports them, an orchestrator, a console for browsing, a common package and a command directory. A Kafka broker sits between collection and enrichment, which lets the two stages scale and be deployed separately.

### Do I need a backend to run Akvorado?

Yes. Flows are exported to ClickHouse, and the build generates a compiled object for UDP socket reuse in the flow inlet, so the collector is written to run at high flow rates. The console is a query interface over that store, with timeseries and Sankey visualisations.

### What is the Akvorado demo site?

It is a live instance running the latest stable version on fake data, described as the direct result of running a container composition on a fresh checkout. Its flow ports are not accessible, so you cannot send your own flows, and it also serves the documentation matching the version it runs.

### What licence is Akvorado released under?

AGPL-3.0. The readme states the project is developed by Free, a French internet provider. The readme also carries a caution to always read the changelog before upgrading.

## Sources

- [akvorado/akvorado on GitHub](https://github.com/akvorado/akvorado)
- [License: AGPL-3.0](https://github.com/akvorado/akvorado/blob/main/LICENSE)
- [Project website](https://demo.akvorado.net)
- [README](https://github.com/akvorado/akvorado/blob/main/README.md)
- [Releases](https://github.com/akvorado/akvorado/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/akvorado-akvorado
