onedr0p/cluster-template: a Talos and Flux starting point for one Kubernetes cluster
A template for deploying a Talos Kubernetes cluster including Flux for GitOps
At a glance
- What is it?
- The template renders Talos machine configs and a Flux GitOps tree from a single TOML file. It suits engineers who already know containers, YAML and Git, and it assumes a domain and, for public exposure, a Cloudflare account.
- Who is it for?
- Adopt it if you are comfortable with containers, YAML and Git, you have a domain, and you want one cluster managed by Flux with sops secrets and Talos as the OS. Do not adopt it if you need a multi-cluster fleet, managed control planes, or a workflow that avoids a local Python and mise toolchain.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 4 days ago.
- What is it written in?
- Mainly YAML, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem: bootstrapping one cluster without hand-writing every manifest
Standing up a Kubernetes cluster at home or on a small rack is not hard because of any single step. It is hard because of the number of steps that must agree with each other: machine configs for every node, a CNI, a certificate issuer, a GitOps controller, secret handling, and DNS. Get one address or one extension wrong and the node boots into a state that is tedious to diagnose. onedr0p/cluster-template exists to collapse that set of decisions into a single file.
The README describes the target plainly: a template for deploying a single Kubernetes cluster, on bare metal or virtual machines, inspired by the author's own home-ops repository. The audience is someone who already knows containers, YAML and Git. A domain is listed as required, and a Cloudflare account is required only if you intend to expose applications to the public internet; internal-only clusters do not need it.
That scope statement matters more than it looks. This is a template for one cluster, not a fleet manager. If your problem is "I have twelve clusters and I want them to converge," this is the wrong shape of tool, and no amount of editing the template changes that.
How makejinja turns cluster.toml into machine configs and a Flux tree
The mechanism is template rendering, not a controller. makejinja reads cluster.toml, a pydantic model validates and defaults it, and the result is written out as the configuration files that Talos and Flux consume. The repository layout matches that description: cluster.sample.toml sits at the top level next to cluster.schema.json and makejinja.toml, with the source templates under template/.
So the data flow is one direction. You edit cluster.toml. You render. You commit the rendered output, and Flux reconciles the cluster from Git. The rendered artifacts are the interface, which is why the template can stay agnostic about which Git provider you use: the README lists GitHub, GitLab, Gitea, Forgejo, Codeberg or self-hosted as options.
The component list is opinionated rather than assembled from parts you choose. The README names Flux, Cilium, cert-manager, Spegel, Reloader, Envoy Gateway, external-dns and cloudflared as included. Secrets go through sops, and cloudflared is what reaches applications from outside your local network. If you have already standardized on a different ingress controller or a different secret store, you are not choosing between two options here. You are editing templates.
Two tooling choices support the workflow rather than the cluster: mise manages the development environment, and the README mentions flate for diffing Flux HelmRelease and Kustomization objects. Renovate and GitHub Actions handle dependency and workflow automation.
Installing it and rendering your first configuration
The README lays out six stages in order, and it is worth respecting that order because later stages assume earlier ones. Stage 1 is hardware selection, Stage 2 is machine preparation, and Stage 3 is the local workstation. The install work happens in Stage 3.
First, create your repository from the template. The README gives the GitHub CLI form, with the repository name in a variable so you can change it:
export REPONAME="home-ops"
gh repo create $REPONAME --template onedr0p/cluster-template --public --clone
cd $REPONAMEIf you are not on GitHub, the README says any Git provider works. Clone the template with `git clone --depth 1 https://github.com/onedr0p/cluster-template`, re-initialize it with `git init`, and push it to an empty repository on your provider.
Next, install the Mise CLI and activate it in your shell, following the guide the README links. Then install the pinned CLI tools:
mise trust
mise installThe README notes that if tool installation fails, unsetting the `GITHUB_TOKEN` environment variable and rerunning these commands is worth trying. Tool downloads are pinned per platform in `.mise/mise.lock`, under `lockfile_platforms` in the mise config.
The justfile exposes the template work as recipes. Initialization creates cluster.toml, an age key, a deploy key and a webhook token; configure renders and validates the configuration files:
just init
just configureAfter `just configure`, you should have rendered configuration files on disk and a validation pass against the schema. Read the rendered output before committing it. The README's Stage 2 also gives a network check you can run before any of this, to confirm your nodes are reachable on the Talos API port:
nmap -Pn -n -p 50000 192.168.1.0/24 -vv | grep 'Discovered'Replace the CIDR with your own network. This is a discovery check, not a test of the cluster.
Where this template fights you: hardware, Python and single-cluster scope
The most concrete limitation is stated in the README itself, and it is about disks. Bare metal is strongly recommended over virtualized platforms like Proxmox. Enterprise NVMe or SATA SSDs are preferred even when used; consumer drives carry risks the README names directly: latency spikes, corruption and fsync delays, particularly in multi-node setups. Proxmox with enterprise drives can work for testing or a carefully tuned production cluster, but it adds I/O contention. Replicated storage such as Rook-Ceph or Longhorn should use dedicated disks separate from control plane and etcd nodes.
That is not a soft preference. If your hardware is a handful of mini PCs with consumer SSDs and a hypervisor, you are outside the configuration the README recommends, and the failure modes it lists are the ones you should expect. The README also says the best way to know your hardware works is to test and benchmark it under realistic workloads, which is an admission that the guidance is a baseline rather than a guarantee.
The toolchain is a second constraint. Rendering runs through makejinja and pydantic, declared in pyproject.toml with `requires-python = ">=3.14"`. That is a recent Python floor, and it means the templating step happens on a workstation with a working mise and Python environment. There is no indication of a container-only or CI-only path for rendering, so a contributor without that environment is blocked at the render step even though the cluster itself is declarative.
Third, the single-cluster scope is a boundary, not a bug. The README describes one cluster. Workflows built around many clusters with per-cluster overlays are a different project's problem.
Finally, the bootstrap path depends on external services for the public-exposure case: the Talos Linux Image Factory for images and Cloudflare for cloudflared. The README instructs you to note the schematic ID from the factory, because you need it later. Lose that ID and you cannot reproduce the image you flashed.
How it differs from a Rancher or RKE2 cluster template
The obvious alternative class is a Rancher or RKE2 cluster template, which people search for alongside this project. The difference is where the operating system boundary sits and who owns the control plane.
RKE2 installs Kubernetes onto a general-purpose Linux distribution and manages the cluster through Rancher. The node is a Linux server that also runs Kubernetes. Talos, which this template uses, is an immutable OS whose API is the machine configuration itself; you do not SSH in and edit files, you apply config. That changes what "fixing a node" means, and it changes what a rendered machine config has to contain.
The second difference is the reconciliation model. A Rancher template typically hands cluster lifecycle to Rancher's own management plane. Here, Flux syncs from your Git provider, and the README's component list (Cilium, cert-manager, Envoy Gateway, external-dns, Spegel, Reloader, cloudflared) is the opinionated stack that Flux brings up. You are not selecting add-ons from a catalog; you are rendering a tree and letting Flux converge it.
That makes the two approaches answer different questions. If you want a UI that provisions clusters and shows you their state, Rancher's model fits. If you want the cluster's desired state to live entirely in a Git repository you control, with the OS configured from the same generated output, this template's model fits. The trade is that you give up the management UI and take on the templating toolchain.
Maintenance cost, release cadence and the MIT licence
The repository is not archived, and the last push was on 2026-09-19. Releases follow a monthly pattern: 2026.9.0 on 2026-09-01, 2026.8.0 on 2026-08-01, and 2026.7.0 on 2026-07-21. That cadence tells you what maintaining a cluster built from this template looks like in practice. You are not adopting a frozen scaffold. You are adopting a moving one, and Renovate is included precisely because dependencies need to move with it.
Because the output is committed to your repository, upgrades are merges rather than package installs. That is a real advantage: you can see the diff between the template's current state and yours before accepting it. It is also a real cost, because a rendered tree that you have edited will conflict with upstream changes, and resolving those conflicts is manual work.
There is no rollback procedure documented in the README. If a rendered change breaks reconciliation, the recovery path is whatever your Git history and Talos configuration allow, and the README does not describe one. Treat your commit history as the mechanism and verify that you can restore a previous rendered state before you need to.
The licence is MIT. That is permissive: it allows use, modification and redistribution with the licence and copyright notice retained. It says nothing about the licences of the components the template deploys, which are separate projects with their own terms, and it is not legal advice about your obligations when you run them.
Editorial conclusion
Adopt it if you are comfortable with containers, YAML and Git, you have a domain, and you want one cluster managed by Flux with sops secrets and Talos as the OS. Do not adopt it if you need a multi-cluster fleet, managed control planes, or a workflow that avoids a local Python and mise toolchain. Before you commit, verify three things: that the schematic ID from the Talos Image Factory matches the extensions you flashed, that your Git provider is among the ones the README lists for Flux sync, and that cluster.toml validates against cluster.schema.json with your node addresses and domain filled in.
Frequently asked questions
Do I need a Cloudflare account to use onedr0p/cluster-template?
Only if you expose applications to the public internet. The README states that internal-only clusters do not require a Cloudflare account, while public exposure does.
What does onedr0p/cluster-template use to manage secrets?
The README lists sops as the secret management component, and the initialization recipe creates an age key alongside cluster.toml, a deploy key and a webhook token.
Which Git providers can Flux sync from with onedr0p/cluster-template?
The README names GitHub, GitLab, Gitea, Forgejo, Codeberg or self-hosted as supported sources for Flux.
How many nodes does onedr0p/cluster-template need?
The README recommends making 3 of your nodes controller nodes if you have 3 or more, for a highly available control plane. It configures all nodes to run workloads, so worker nodes are optional. The minimum system requirements table lists 4 cores, 16GB memory and a 256GB SSD or NVMe per control or worker node.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/onedr0p-cluster-template)