# Fluxcd Flagger: Progressive Delivery on Kubernetes with Canary and Blue/Green Deployments

> Flagger is a CNCF graduated Kubernetes operator that automates canary releases, A/B testing and blue/green mirroring by shifting traffic and checking metrics. It is for teams already running a service mesh or ingress controller who want automated rollback, and it is not a substitute for those components.

**fluxcd/flagger** — Progressive delivery Kubernetes operator (Canary, A/B Testing and Blue/Green deployments)

- Repository: https://github.com/fluxcd/flagger
- Website: https://docs.flagger.app/main
- Stars: 5,415 · Forks: 805
- Language: Go
- License: Apache-2.0
- Published: 2026-09-22 · Updated: 2026-09-22 · Language: en
- Canonical page: https://hysenlabs.com/projects/fluxcd-flagger

## What Flagger automates that kubectl rollout does not

A Kubernetes Deployment rolls out all at once unless you tune maxSurge and maxUnavailable, and it has no opinion about whether the new pods are actually serving good responses. Flagger takes a Deployment, optionally a HorizontalPodAutoscaler, and creates a set of objects around it: Kubernetes deployments, ClusterIP services, and service mesh or ingress routes. Those objects expose the application and drive the canary analysis and promotion. The target audience is platform and release engineers who own a cluster with a mesh or ingress controller already in place and who want the promotion decision to come from measurements instead of from a person watching a dashboard. The README frames the goal plainly: it "reduces the risk of introducing a new software version in production by gradually shifting traffic to the new version while measuring metrics and running conformance tests." That combination, traffic shifting plus metric gates plus rollback, is the whole product.

## The Canary resource and the analysis loop

Everything is driven by a custom resource of kind Canary with apiVersion flagger.app/v1beta1. The spec has three parts that matter. The provider field selects the traffic layer and can be kubernetes, istio, linkerd, kuma, knative, nginx, contour, gloo, traefik or skipper, with gatewayapi:v1 and gatewayapi:v1beta1 for Gateway API implementations. The service block describes the generated ClusterIP service: port, targetPort, portName, optional match and rewrite rules, timeout, and portDiscovery to include the other container ports. The analysis block holds the timing and the key performance indicators: interval defaults to 60s, threshold is the maximum number of failed metric checks before rollback, maxWeight is the maximum percentage of traffic sent to the canary, and stepWeight is the increment per interval. Metrics are checked against thresholdRange. Two built-in Prometheus checks are named in the example, request-success-rate with a minimum percentage and request-duration with a maximum P99 in milliseconds, and custom checks are attached through templateRef. Webhooks run at defined points: a pre-rollout hook in the example calls a Helm test, and a rollout hook runs a load generator against the canary. Alerts route by severity to providers referenced by name.

The scheduling detail that is easy to miss: Flagger tracks the ConfigMaps and Secrets referenced by the Deployment and starts a canary analysis if any of them change. Code and configuration are synchronized on promotion, so a config-only change goes through the same gate as an image bump.

## Installing Flagger and running a first canary

The README points to the documentation for installation, specifically the page titled Flagger Install with Flux at docs.flagger.app. There is no install command in the README itself, so the honest starting point is that page plus the charts and kustomize directories visible in the repository root. What the README does give is a complete Canary manifest, which is the more useful artifact because it shows the shape of the CRD before you commit to anything.

The example below is the README's own manifest, trimmed to the parts you need for a first run against Istio. The provider is set to istio, the target is a Deployment named podinfo in namespace test, and the analysis promotes in 5 percent steps up to 50 percent, checking request success rate and request duration each minute.

```yaml
apiVersion: flagger.app/v1beta1
kind: Canary
metadata:
  name: podinfo
  namespace: test
spec:
  provider: istio
  targetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: podinfo
  progressDeadlineSeconds: 60
  service:
    port: 9898
    targetPort: 9898
    portName: http
    portDiscovery: true
  analysis:
    interval: 1m
    threshold: 10
    maxWeight: 50
    stepWeight: 5
    metrics:
    - name: request-success-rate
      thresholdRange:
        min: 99
      interval: 1m
```

Save that manifest and apply it in the namespace where the target Deployment already exists. The reader should expect Flagger to create the canary Deployment, the primary and canary ClusterIP services, and the Istio routes that carry the weights. Progress is observable through the status subresource of the Canary object, which is how you confirm the loop is actually running rather than waiting on a silent controller.

Two fields in that manifest are worth changing before production use. progressDeadlineSeconds is 60 in the example, which is aggressive; the README notes the default is 600s and describes it as the maximum time for the canary deployment to make progress before rollback. And skipAnalysis defaults to false, which is what you want, but setting it true promotes the canary without analysing it, a switch that exists for a reason and should be treated as one.

## Where Flagger is the wrong tool

Flagger does not implement traffic shifting itself. The provider field decides who does, and if you have neither a service mesh nor an ingress controller that supports weighted routing, the kubernetes provider is the fallback and the granularity of the rollout is necessarily coarser. The README's feature table is organized as a matrix of capabilities against Istio, Linkerd, Kuma, Knative and Kubernetes CNI, which is a clear signal that behaviour is not uniform across providers. A team that assumes an analysis tuned on Istio will behave identically on NGINX or Contour is making an assumption the documentation does not support.

The second boundary is metrics. The built-in checks are request-success-rate and request-duration, both described as Prometheus checks. If your workload does not emit request metrics that a Prometheus-compatible backend can query, the default analysis has nothing to evaluate, and you are left writing custom metric templates or running with skipAnalysis, which defeats the purpose. The example's third metric, named "database connections" with a templateRef, shows the escape hatch, but it is work you have to do.

Third, this is a controller with cluster-level responsibilities. The Dockerfile builds a single binary that runs as the flagger entrypoint under a non-root user, and the Makefile's version-set target rewrites the version string across pkg/version/version.go, artifacts/flagger/deployment.yaml, charts/flagger/values.yaml, the Helm chart metadata and kustomize/base/flagger/kustomization.yaml in one step. That is a maintenance surface: an upgrade touches the CRD, the deployment manifest, the chart and the kustomize base together, and the Makefile's crd target copies artifacts/flagger/crd.yaml into both charts/flagger/crds/crd.yaml and kustomize/base/flagger/crd.yaml, so the CRD exists in three places that must stay in sync.

## Argo Rollouts compared, and how the two differ

The obvious comparison for anyone evaluating progressive delivery on Kubernetes is Argo Rollouts, which also introduces a custom resource in place of a plain Deployment and also drives canary and blue/green promotion. The difference in approach is where the traffic control lives. Flagger keeps the Deployment as the source of truth and creates the surrounding services and routes itself, generating a primary and a canary workload from the one you already have. Argo Rollouts replaces the Deployment with a Rollout object, so the workload definition moves into the new kind. That distinction matters for GitOps: with Flagger your existing Deployment manifests stay as they are and the Canary is an additional object, while with Rollouts the manifest you already have is rewritten. Flagger's README also states that it keeps track of referenced ConfigMaps and Secrets and triggers analysis when they change, which ties configuration drift into the same promotion gate. Neither approach is free. Flagger's model means more generated objects in the namespace and a provider abstraction you must configure correctly; the Rollout model means a migration of every workload you want to manage.

## Maintenance, releases and licence

The repository is not archived and the last push was on 2026-09-21. Recent releases are v1.45.0 on 2026-09-01, v1.44.0 on 2026-07-21 and v1.43.0 on 2026-04-21, a cadence of roughly one minor release per quarter with a patch in between. Flagger is a CNCF graduated project and part of the Flux family of GitOps tools, which is relevant to the upgrade question because the documented install path is Flagger Install with Flux. If you already run Flux, the controller is managed the same way as everything else in the cluster; if you do not, you are adding a second GitOps tool to manage one operator, which is a real cost to weigh.

The licence is Apache-2.0, as stated in the repository metadata and present as the LICENSE file at the top level. That is a permissive licence with a patent grant, and it is the same licence as the wider Kubernetes ecosystem, so there is no licensing obstacle to embedding the operator in a commercial platform. This is not legal advice; check the LICENSE file and your own obligations.

Upgrade cost is dominated by the CRD and the generated objects rather than by the binary. The Makefile's verify-crd and test-codegen targets exist because the CRD and the generated client code are checked in, so a version bump is a coordinated change across pkg/version/version.go, the deployment manifest, the Helm chart and the kustomize base. Budget for that on every minor release, not just on majors.

## Conclusion

Adopt Flagger if you already run Istio, Linkerd, Kuma, Knative, NGINX, Contour, Gloo, Traefik or Skipper on Kubernetes and you want rollout promotion and rollback driven by metrics rather than by hand. Do not adopt it if you have no ingress or mesh provider to shift traffic with, or if you cannot expose Prometheus-style metrics for the built-in checks, because the analysis loop has nothing to read. Before committing, verify two things in your own cluster: that the provider value you intend to use is listed in the Canary spec, and that request-success-rate and request-duration return data for your workload, since the default analysis depends on them.

## FAQ

### What does Flagger do in Kubernetes?

It automates the release process for applications running on Kubernetes by gradually shifting traffic to a new version while measuring metrics and running conformance tests, and it rolls back when the checks fail. It implements canary releases, A/B testing and blue/green mirroring, and integrates with Kubernetes ingress controllers, service meshes and monitoring solutions.

### How do I install Flagger?

The README does not give install commands; it points to the documentation at docs.flagger.app, with a page titled Flagger Install with Flux. The repository also contains charts/ and kustomize/ directories that hold the deployment artifacts.

### Which providers can a Flagger Canary use?

The provider field can be kubernetes, istio, linkerd, kuma, knative, nginx, contour, gloo, traefik or skipper, and gatewayapi:v1 and gatewayapi:v1beta1 for Gateway API implementations. The README notes that the provider is optional.

### What do the analysis interval, threshold, maxWeight and stepWeight control?

interval is the schedule interval and defaults to 60s, threshold is the maximum number of failed metric checks before rollback, maxWeight is the maximum traffic percentage routed to the canary, and stepWeight is the increment step per interval. In the README example the step is 5 percent up to a maximum of 50 percent.

### Does Flagger react to ConfigMap and Secret changes?

Yes. The README states that Flagger keeps track of ConfigMaps and Secrets referenced by a Kubernetes Deployment and triggers a canary analysis if any of those objects change, so code and configuration are synchronized when a workload is promoted in production.

## Sources

- [fluxcd/flagger on GitHub](https://github.com/fluxcd/flagger)
- [License: Apache-2.0](https://github.com/fluxcd/flagger/blob/main/LICENSE)
- [Project website](https://docs.flagger.app/main)
- [README](https://github.com/fluxcd/flagger/blob/main/README.md)
- [Releases](https://github.com/fluxcd/flagger/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/fluxcd-flagger
