Self-hosted service
otwld/ollama-helm avatar
otwld/ollama-helm

ollama-helm: the chart version tracks Ollama's own release

Helm chart for Ollama on Kubernetes

592 stars93 forksGo TemplateMIT

At a glance

What is it?
A community Helm chart that follows Ollama upstream version for version, with a CPU path that accepts clusters ten releases older than the GPU path, a Knative mode that swaps the Deployment for a Service and needs three cluster flags first, and a documented values rename between 0.x and 1.x. It also carries a registry address that changed.
Who is it for?
This chart's main design decision is that it does not lag Ollama, and the upgrade section makes the consequence explicit by telling you to read the upstream release notes before every upgrade.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 2 days ago.
What is it written in?
Mainly Go Template, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 8, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Chart versions are Ollama versions

The three newest release tags are `ollama-1.85.0`, `ollama-1.84.0` and `ollama-1.83.0`, published on 2026-10-02, 2026-09-26 and 2026-09-19. The chart's version number is the upstream application's version number, prefixed.

That coupling is stated as an instruction rather than left to be inferred. The upgrade section opens by telling you to read Ollama's own release notes first to make sure there are no backwards incompatible changes, and only then to adjust your values and run the upgrade. In other words, the page treats every chart bump as an application upgrade with its own migration notes.

The distribution channel follows from that. The declared homepage is the Artifact Hub package page rather than a documentation site, and the root carries `artifacthub-repo.yml` for the listing metadata plus a `banner.png` for the listing image. Installation itself goes through a Helm repository registry, not through Artifact Hub.

There is also an `update-cli.yaml` at the root whose contents the page never explains, in a repository that otherwise documents its own release process closely.

Two Kubernetes floors ten releases apart

The requirements section gives two different minimums for two different workloads. CPU-only deployment asks for Kubernetes `>= 1.16.0-0`, while stable GPU support asks for `>= 1.26.0-0` and names both NVIDIA and AMD.

So a CPU-only deployment runs on clusters that predate the GPU floor by ten minor releases, and the reason is that GPU support is the part that depends on cluster features that did not exist at 1.16. The zero patch suffix on both is Helm's own way of expressing a lower bound.

The GPU line then carries its own caveat in italics: not all GPUs are currently supported with Ollama, especially with AMD. And the examples section repeats the point from the other direction, calling it highly recommended to run an updated version of Kubernetes when deploying Ollama with a GPU.

Which means the two paths have different risk profiles in the same chart. The CPU path is a wide compatibility floor with no hardware caveat, and the GPU path is a narrow one with an explicit warning about one of the two vendors it nominally supports.

The registry moved and only the new address is shown

An important note sits above the install commands. The project is migrating its registry from a GitHub Pages URL to a Helm central registry, and it asks users to update their Helm registry accordingly. Every command on the page uses the new address:

console
helm repo add otwld https://helm.otwld.com/
helm repo update
helm install ollama otwld/ollama --namespace ollama --create-namespace

The old address appears only inside that migration note. So an installation that still points at the GitHub Pages repository is silently on the deprecated path, and nothing on the page maps the old index to the new one or says whether both serve the same chart versions. The upgrade and delete commands follow the same pattern, using `otwld/ollama` as the chart reference.

This is the one operational item on the page that will affect an automated pipeline before it affects a human, since a repository added once in CI keeps its old URL until someone edits it.

Knative mode needs three feature flags or the Service is rejected

Setting `knative.enabled=true` changes what gets deployed: a Knative `Service` instead of a standard Kubernetes `Deployment` and `Service`. That swap is the whole feature, and it carries a cluster prerequisite.

To mount a PVC in Knative mode, three feature flags must be enabled on the Knative Serving cluster: `kubernetes.podspec-persistent-volume-claim`, `kubernetes.podspec-persistent-volume-write` and `kubernetes.podspec-tolerations`. The stated consequence of leaving them off is precise: Knative rejects the rendered `Service` with webhook validation errors. So this is a render-time failure rather than a runtime one, which is a slightly kinder failure mode than a pod that will not start.

The page gives both the inspection and the fix:

console
kubectl get configmap/config-features -n knative-serving -o yaml

followed by a patch of that same ConfigMap setting all three data keys to enabled. And it names the case where that is the wrong tool: if the cluster is managed by the Knative Operator, the flags belong on the `KnativeServing` resource instead, rather than patched into the ConfigMap directly.

Model bootstrap is a Job that helm does not wait for

In Knative mode the chart cannot preload models in the serving container, because a Knative `Service` starts cold. So it renders a separate bootstrap `Job` when any of `ollama.models.pull`, `ollama.models.run` or `ollama.models.create` is configured.

Four behaviours are documented and they matter in different ways. The Job is non-blocking, so `helm install` and `helm upgrade` return without waiting for the preload to finish. The Job name includes a hash of `ollama.models.pull`, `ollama.models.run`, `ollama.models.create` and `ollama.models.clean`, so changing any of those on upgrade produces a new Job and reruns the bootstrap rather than leaving a stale one in place. Completed Jobs are cleaned up according to `knative.modelBootstrap.ttlSecondsAfterFinished`. And the whole mechanism requires `persistentVolume.enabled=true`, because the Job and the Service have to share the same model data.

That last requirement is the one to watch. Without a persistent volume the bootstrap pulls models into storage that the next Service instance does not see, which is exactly the failure the shared-claim example exists to prevent:

yaml
persistentVolume:
  enabled: true
  existingClaim: ollama-pvc

Three ways to prepare models, and a values rename at 1.x

Models can be pulled by name, built from a template, and then run. The template example is the most capable of the three, because it creates a derived model rather than pulling a published one:

yaml
ollama:
  models:
    create:
      - name: llama3.1-ctx32768
        template: |
          FROM llama3.1
          PARAMETER num_ctx 32768
    run:
      - llama3.1-ctx32768

A base model plus a parameter block becomes a named model, and the run list then references that derived name. So the chart can be configured to build a variant at startup rather than only fetching one.

The GPU example carries the other shape, with `gpu.enabled`, a `gpu.type` of `nvidia` or `amd` and a `gpu.number`, whose inline comment reads as a copy of the default value rather than a description of the range. The Ingress example enables `ingress.enabled` with a host and a prefix path, after which the API is reachable at the host you named.

The breaking change is the rename. Version 1.x added the ability to load models in memory at startup, which moved the list from `ollama.models` to `ollama.models.pull`, and the page shows the before and after so an upgrade does not silently drop every model.

A Kind cluster in CI, and a values table that starts with affinity

The repository carries its own test rig. There is a `ci/` directory and a `kind-config.yml`, which together mean chart changes are exercised against a local Kind cluster rather than only linted, and `.helmignore` sits beside them for excluding files from the packaged chart.

The published values table begins with a single row: `affinity`, of type object, defaulting to an empty object. That is an honest start for a chart, since scheduling is where a GPU workload is usually decided, and an empty default leaves the choice entirely to the operator. The table is cut off after that row on the page, with `values.yaml` named as the place to read the rest.

The rest of the root is small and conventional: `Chart.yaml`, `values.yaml`, `templates/`, `LICENSE`, `README.md` and `.github/` for the CI workflow the badge links to. The primary language is recorded as Go Template, which is the honest classification for a chart, since there is no compiled program here at all.

Editorial conclusion

This chart's main design decision is that it does not lag Ollama, and the upgrade section makes the consequence explicit by telling you to read the upstream release notes before every upgrade. It suits an operator who wants models pulled and run at startup rather than by hand, especially on Knative where a separate bootstrap Job does the loading, and it does not suit anyone who treats a chart bump as a routine dependency update, because a chart bump here is an application upgrade with its own breaking-change history. Before you upgrade, read the Ollama notes, check whether your Kubernetes version clears the GPU floor of 1.26, and if you use Knative, verify the three feature flags rather than patching them blindly.

Frequently asked questions

What Kubernetes version does the ollama-helm chart require?

Version 1.16.0 or newer for CPU-only deployments, and 1.26.0 or newer for GPU support on both NVIDIA and AMD. The page also notes that not all GPUs are currently supported with Ollama, especially with AMD, and recommends a current Kubernetes release when deploying with a GPU.

How do I install the ollama Helm chart?

Add the registry with helm repo add otwld https://helm.otwld.com/, run helm repo update, then helm install ollama otwld/ollama --namespace ollama --create-namespace. The registry has moved from a GitHub Pages URL to the Helm central registry at helm.otwld.com, so an existing repository entry may need updating.

What does enabling knative.enabled change in the ollama chart?

The chart deploys a Knative Service instead of a standard Deployment and Service. Mounting a PVC in that mode then requires three Knative Serving feature flags: kubernetes.podspec-persistent-volume-claim, kubernetes.podspec-persistent-volume-write and kubernetes.podspec-tolerations, or Knative rejects the rendered Service with webhook validation errors.

How does the ollama chart load models at startup?

Through ollama.models.pull, ollama.models.create and ollama.models.run, with create able to build a derived model from a base model and a parameter block. In Knative mode this becomes a separate bootstrap Job that helm does not wait for, whose name hashes those settings so changes rerun it, and which requires persistentVolume.enabled so the Job and the Service share model data.

Official sources

  1. License: MIT
  2. otwld/ollama-helm on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/otwld-ollama-helm.svg)](https://hysenlabs.com/projects/otwld-ollama-helm)