# k3s-ansible: An Ansible Collection for HA k3s with kube-vip and MetalLB

> timothystewart6/k3s-ansible builds a highly available k3s cluster with an etcd control plane, a kube-vip virtual IP and MetalLB service load balancing. It is aimed at operators who already run Ansible and want a reproducible cluster build rather than a hand-installed one.

**timothystewart6/k3s-ansible** — The easiest way to bootstrap a self-hosted High Availability Kubernetes cluster.  A fully automated HA k3s etcd install with kube-vip, MetalLB, and more.  Build. Destroy. Repeat.

- Repository: https://github.com/timothystewart6/k3s-ansible
- Website: https://technotim.com/posts/k3s-etcd-ansible/
- Stars: 3,017 · Forks: 1,175
- Language: Jinja
- License: Apache-2.0
- Published: 2026-09-24 · Updated: 2026-09-24 · Language: en
- Canonical page: https://hysenlabs.com/projects/timothystewart6-k3s-ansible

## What k3s-ansible builds and who it is for

The project is an Ansible collection that provisions a highly available Kubernetes cluster using k3s. It supports kube-vip for the control plane virtual IP, multiple CNI options, and either MetalLB or kube-vip for service load balancing. The README describes it as based on a fork of k3s-io/k3s-ansible, with kube-vip handling the control plane load balancer and MetalLB handling the service LoadBalancer.

The audience is narrow and specific. You need machines you control, passwordless SSH to them (or the --ask-pass --ask-become-pass flags), a control node with Ansible 2.11 or newer, and a network where you can dedicate a virtual IP to the API server. The README lists Debian 13, Ubuntu 26.04 LTS and Rocky 10 as tested, with x64, arm64 and armhf architectures. If you are renting a managed Kubernetes control plane, this project is not for you: it exists precisely to build the control plane yourself. If you have two or three spare boxes and want an etcd-backed cluster you can destroy and rebuild, it fits.

## How the playbooks, roles and inventory fit together

The repository is a standard Ansible layout plus a galaxy.yml, which is what makes it installable as a collection. site.yml is the entry point for creation, reset.yml tears the cluster down, and reboot.yml handles node reboots. The roles/ directory holds the actual work, templates/ holds rendered files, and inventory/ holds the sample inventory that you copy per cluster. The README notes that the inventory directory ignores custom content so credentials and environment details are not committed accidentally.

Data flow is conventional Ansible. You describe hosts in an INI inventory, group them as master and node under a k3s_cluster parent, and set variables in group_vars/all.yml. The playbook reads those variables, and the roles apply them. One behaviour is worth calling out because it is easy to miss: if the master group contains more than one host, the playbook automatically sets up k3s in HA mode with embedded etcd. There is no separate flag to enable HA, which means adding a second master later changes the datastore model. The README also states that site.yml asserts unique hostnames up front and fails fast on duplicates, because k3s registers nodes keyed by hostname and two machines with the same name cannot both join.

## Installing the collection and creating a first cluster

The control node needs the collections from the repository's requirements file before anything runs. The README gives this command, and netaddr must be importable by Ansible, which is already true if Ansible came from apt but needs an explicit install into the virtual environment if it came from pip.

```bash
ansible-galaxy collection install -r ./collections/requirements.yml
```

Next, copy the sample inventory and edit the host list. The README's example uses three masters and two agents, which is the shape that triggers etcd HA.

```bash
cp -R inventory/sample inventory/my-cluster
```

```ini
[master]
192.168.30.38
192.168.30.39
192.168.30.40

[node]
192.168.30.41
192.168.30.42

[k3s_cluster:children]
master
node
```

Copy ansible.example.cfg to ansible.cfg and update the inventory path; the local ansible.cfg is ignored by Git. Then run the site playbook against your inventory. The minimum k3s version is 1.19.1 and the version is selected with the k3s_version variable.

```bash
ansible-playbook site.yml -i inventory/my-cluster/hosts.ini
```

After deployment, the README states the control plane is reachable through the virtual IP defined by apiserver_endpoint. To use kubectl from your workstation, copy the kubeconfig from a master. The README warns to avoid world-writable permissions on it because it contains cluster credentials.

```bash
scp debian@master_ip:/etc/rancher/k3s/k3s.yaml ~/.kube/config
```

## Removing and rebooting a cluster without surprises

Two operational playbooks sit alongside site.yml. reset.yml removes k3s, and the README adds a warning that is easy to skip: reboot the nodes afterwards because the virtual IP may remain configured. That is a direct consequence of managing the VIP outside the cluster, and it means a teardown is not complete until the machines restart.

```bash
ansible-playbook reset.yml -i inventory/my-cluster/hosts.ini
```

reboot.yml restarts every node, or stages the restart. The README documents concurrent_reboots, which takes a node count or a percentage, and wait_seconds_after_reboot, which pauses after each batch so pods in the freshly rebooted batch can settle. The example passes both as extra vars.

```bash
ansible-playbook reboot.yml -i inventory/my-cluster/hosts.ini \
  --extra-vars 'concurrent_reboots=2 wait_seconds_after_reboot=30'
```

The staging option is the more interesting one. Rebooting all nodes at once is the default path, and with a single control plane VIP that is a full outage. The batch variables are the project's answer, but they are opt-in and the README does not say what values are safe for a given cluster size.

## Upgrades are the weak point, and the README says so

The upgrade section is unusually candid. The version variables select components for a fresh installation and are not a supported direct in-place upgrade path for an existing cluster. K3s, Calico and Cilium each require staged upgrades, and the playbook does not automate them, so backups and health checks at each step remain manual operational work.

The documented K3s path is strict: do not jump an embedded-etcd cluster straight to Kubernetes 1.36. Upgrade one minor version at a time. From the sample default of v1.30.2+k3s2, the sequence runs through the latest supported 1.30 patch, then 1.31, 1.32, a 1.33 patch containing etcd 3.5.26 such as v1.33.7+k3s3, then 1.34, 1.35 and finally 1.36, upgrading servers one at a time before agents. Cilium supports only consecutive minor upgrades, so 1.17 through 1.20 in order with preflight checks, and a direct jump from an old Cilium to 1.20 is explicitly ruled out. Calico changed v3 resource UID behaviour at 3.28, and OwnerReferences pointing at projectcalico.org/v3 resources must be removed and recreated around an in-place upgrade. MetalLB is pinned to application tag v0.16.0, and the README warns that a chart-only tag such as metallb-chart-0.16.1 is not an application or image release and must not be used as the controller or speaker image tag.

This is the honest trade-off of the project. Provisioning is automated; staying current is not. A team that installs once and forgets will fall far enough behind that the staged path becomes a multi-week project.

## Where k3s-ansible is the wrong tool

The clearest mismatch is scale and intent. If you want one k3s node for a homelab service, the HA machinery is dead weight: kube-vip, MetalLB and an embedded etcd quorum exist to survive node loss, and a single node cannot. The README's own HA trigger is a second host in the master group, so a one-master inventory builds a simpler cluster than the project's framing suggests.

The second mismatch is lifecycle. Because upgrades are manual and staged, this collection suits clusters that are rebuilt or deliberately maintained on a schedule rather than clusters that must track upstream Kubernetes continuously. A platform team with dozens of clusters and an upgrade policy enforced by tooling will find that the playbooks stop at installation.

The third is environment. The README lists Debian, Ubuntu and Rocky as tested. Other distributions are not covered by that list, and passwordless SSH or the ask-pass flags are a prerequisite rather than something the collection configures for you.

## How it differs from plain k3s and from k3s-io/k3s-ansible

Plain k3s installation is a single script or binary per node. It gives you a working cluster, but the control plane endpoint is whichever server you point at, and a service of type LoadBalancer stays pending until you install something to satisfy it. k3s-ansible adds two components on top: kube-vip provides a virtual IP for the API server so clients have one stable address, and MetalLB or kube-vip satisfies LoadBalancer services. It also wires the etcd datastore automatically once a second master appears.

The closer comparison is the upstream k3s-io/k3s-ansible, which this project states it is based on through an intermediate fork. The difference is in the added pieces: kube-vip for the control plane VIP, MetalLB for service load balancing, multiple CNI options, and the reset and reboot playbooks. If you only need k3s installed on a set of hosts and you already have your own load balancing and upgrade process, upstream is the smaller dependency. If you want the VIP and the service load balancer handled by the same playbook run, this project covers that ground. The trade-off is that you inherit the pinned component versions and the staged upgrade obligations that come with them.

## Maintenance, licensing and what to check before adopting

The repository is not archived, and the last push was on 2026-09-23, one day before this writing. Releases track k3s closely: v1.36.4+k3s1+tt1 on 2026-09-02, v1.36.3+k3s1+tt1 on 2026-08-15 and v1.36.2+k3s1+tt1 on 2026-08-02. The tt suffix marks the project's own packaging on top of the upstream k3s version, so a release here is not identical to the corresponding k3s release.

The licence is Apache-2.0. That is permissive and permits commercial use and modification, but it is worth noting that the collection installs other components, including k3s, MetalLB, Cilium and Calico, each under its own licence and version policy. Nothing in the repository changes those terms, and the pinned MetalLB tag is a reminder that component versions are chosen by this project rather than inherited from the cluster's Kubernetes version.

The upgrade cost is the real maintenance line item. The README's staged path from v1.30.2+k3s2 to 1.36 involves roughly six Kubernetes minor steps, each with a backup and a health check, plus separate consecutive-minor paths for Cilium and a manual OwnerReference fix for Calico. Budget for that before committing, and check the roles and group_vars in your inventory against the versions the current release actually installs.

## Conclusion

Adopt k3s-ansible if you already run Ansible, want an etcd-backed HA control plane behind a kube-vip virtual IP, and accept that version upgrades are manual work you perform one Kubernetes minor at a time. Do not adopt it if you want a single-node cluster, a managed control plane, or an automated in-place upgrade path, because the README states the version variables select components for a fresh installation only and the playbooks do not automate upgrades. Before running site.yml against production hardware, verify that every node has a unique hostname (site.yml asserts this and fails fast on duplicates), that apiserver_endpoint is free on your network, and that your k3s_version and CNI choice match what the roles actually install.

## FAQ

### Does k3s-ansible set up HA automatically, or do I have to enable it?

It is automatic. The README states that if multiple hosts are in the master group, the playbook will automatically set up k3s in HA mode with etcd. There is no separate HA flag to set.

### Can k3s-ansible upgrade an existing cluster in place?

No. The README states the version variables select components for a fresh installation and are not a supported direct in-place upgrade path, and that the playbook does not automate upgrades. K3s, Calico and Cilium each require staged upgrades that remain manual operational steps.

### Which operating systems and architectures does k3s-ansible support?

The README lists Debian (tested on version 13), Ubuntu (tested on version 26.04 LTS) and Rocky (tested on version 10), with x64, arm64 and armhf processor architectures supported.

### Why do the cluster nodes need unique hostnames with k3s-ansible?

The README explains that k3s registers each node keyed by its hostname, so two nodes with the same hostname cannot join the cluster. site.yml asserts this up front and fails fast if any duplicate is found.

### What do I need on the control node before running the k3s-ansible playbooks?

Ansible 2.11 or newer, the collections installed from collections/requirements.yml, and the netaddr package available to Ansible. The README notes netaddr is already handled if Ansible was installed via apt, but must be installed into the virtual environment if Ansible came from pip.

## Sources

- [License: Apache-2.0](https://github.com/timothystewart6/k3s-ansible/blob/master/LICENSE)
- [Project website](https://technotim.com/posts/k3s-etcd-ansible/)
- [README](https://github.com/timothystewart6/k3s-ansible/blob/master/README.md)
- [Releases](https://github.com/timothystewart6/k3s-ansible/releases)
- [timothystewart6/k3s-ansible on GitHub](https://github.com/timothystewart6/k3s-ansible)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/timothystewart6-k3s-ansible
