How to Upgrade the KSML Provisioner to Runtime Provisioner 0.8.0

This guide shows you how to fix the values that block the upgrade, install Runtime Provisioner 0.8.0 beside your running release, repoint Platform Manager and your monitoring at the new name, verify the result, and remove the old release.

Type

How-to guide

Goal

Move a running KSML Provisioner 0.7.0 installation to Runtime Provisioner 0.8.0 and repoint everything that referenced the old name.

Audience

Platform Operator with Helm access to the Provisioner’s namespace, working with a Tenant Admin who can edit the KSML Provisioner configuration in Self-Service.

When to use

Use this guide once per installation, when moving an existing KSML Provisioner to chart version 0.8.0 or higher.

The upgrade runs blue-green. The renamed chart comes up as a second release in the same namespace, the old release keeps serving throughout, and you remove it only once Step 10 has confirmed the new one works. The two releases do not interfere with each other: the chart names every resource after its release, both ServiceAccounts hold the same namespace permissions rather than competing for them, and the Provisioner acts only when Platform Manager calls it, at one address at a time. Step 3 clears the few values that would break that separation. Running KSML applications keep processing data throughout.

One step is not optional. Platform Manager stores the Provisioner’s address as free text and never checks it against the Kubernetes Service name, so the rename makes that address stale without producing an error. Until you complete Step 7: Change the Provisioner address in Platform Manager, Platform Manager cannot reach the Provisioner and every KSML action in Self-Service fails.

What the rename changes

The service keeps provisioning KSML applications and now also reads Kafka Connect (KC) worker logs, so it is called the Runtime Provisioner from 0.8.0 onwards. The chart, the image and every Kubernetes object it creates follow that name, which is what the rest of this guide repoints:

Thing Before After

Helm chart

ksml-provisioner

runtime-provisioner

Container image

registry.axual.io/axual/ksml-provisioner

registry.axual.io/axual/runtime-provisioner

Kubernetes resources

<RELEASE>-ksml-provisioner

<RELEASE>-runtime-provisioner

Identity label

app.kubernetes.io/name: ksml-provisioner

app.kubernetes.io/name: runtime-provisioner

Container name in the pod

ksml-provisioner

runtime-provisioner

Image user and home directory

ksml, /home/ksml

provisioner, /home/provisioner

Your monitoring carries the old name in several more places; Step 8 lists them. Nothing else changes: every endpoint path, every environment variable name, every chart values key, the KSML chart and your KSML applications all stay as they are.

The image user keeps User Identifier (UID) 1024, so no securityContext or Security Context Constraint (SCC) change is needed, and a volume mounted at an absolute path under /home/ksml still lands there. Repoint only a relative mount path in an extra init container or sidecar of your own.

Prerequisites

Confirm the following before you begin.

Access and permissions required

You need the following access and permissions:

  • Helm permission to install, upgrade and uninstall releases in the Provisioner’s namespace, and to create a ServiceAccount, Role and RoleBinding there.

  • Access to the Axual chart registry. See Working with Helm Charts for how to gain it.

  • A Tenant Admin account in Self-Service, or a colleague holding one, to change the Provisioner address in Step 7.

  • Edit access to your Prometheus rules, Grafana dashboards and NetworkPolicies, so you can repoint the ones that name the old resource.

Tools and versions required

You need the following tools:

  • helm >= 3.12.

  • kubectl >= 1.28.

Resources that must exist before starting

The following must already exist:

  • A running KSML Provisioner 0.7.0 release, installed with Helm.

  • Version 0.8.0 or higher of the runtime-provisioner chart, published in the Axual registry. The old ksml-provisioner chart stops at 0.7.x by design.

  • A mirror of registry.axual.io/axual/runtime-provisioner, if you mirror Axual images into your own registry rather than pulling them directly.

Replace every <VALUE> placeholder with your own value before running a command. <RELEASE> is the name of your existing Helm release and <NAMESPACE> is the namespace it runs in.

Step 1: Save your current values

Find your release and export the values it runs with, so you can reinstall with the same configuration.

helm list --namespace <NAMESPACE>

helm get values <RELEASE> --namespace <NAMESPACE> -o yaml > provisioner-values.yaml

Write the release name down. Every command below uses it as <RELEASE>. Open provisioner-values.yaml and confirm your real settings are in it.

A file containing only null means your values are supplied elsewhere, for example by an ArgoCD Application or a parent chart. Find that file before going further, because it is the one you have to edit in Step 2, Step 3 and Step 4, and a fresh install from an empty values file would drop your configuration.

Step 2: Fix the values that block the upgrade

Two settings stop the upgrade rather than degrading it, so correct both in your values file before you install.

First, replace any reference to the old chart name in a templated value. The chart renders five fields as Go templates: extraInitContainers, extraContainers, extraVolumes, extraVolumeMounts and prometheusRule.extraRules. In each of them, change ksml-provisioner.fullname to runtime-provisioner.fullname.

prometheusRule:
  extraRules: |
    - alert: MyOwnRule
      expr: up{service="{{ include "runtime-provisioner.fullname" . }}"} == 0
Check these five fields even if you did not write them yourself. The chart’s own commented extraRules example carried the old spelling, so a values file copied from it still contains ksml-provisioner.fullname. The upgrade then fails with no template "ksml-provisioner.fullname" and changes nothing.

Second, remove or raise a pinned image tag. The registry.axual.io/axual/runtime-provisioner repository starts at 0.8.0, because the older tags are not copied into it, so a values file pinning an earlier tag gives ImagePullBackOff.

image:
  tag: "0.8.0"   # or remove the tag and take the chart's appVersion

If you mirror Axual images into your own registry, mirror registry.axual.io/axual/runtime-provisioner:0.8.0 before you install.

Step 3: Clear the values that collide with the running release

The two releases stay apart because the chart names every resource after its release. Five values override that, and your saved values file carries whatever the old release set, so check each one before you install.

Values key What to do

fullnameOverride

Remove it. It pins one fixed name for the Deployment, Service and ConfigMap, so the second release tries to create objects the first one already owns and Helm refuses the install with invalid ownership metadata.

serviceAccount.name

Remove it, or give the new release a name of its own. A fixed name makes both releases claim one ServiceAccount, which fails the same way.

nameOverride

Remove it. The release name still keeps the resources apart, so nothing collides, but it keeps the old app name on the new release and defeats the rename.

ingress.hosts[].host

Give the new release its own hostname while both run. Two Ingresses claiming one host is settled by the ingress controller rather than by you, so traffic can keep reaching the old pod after you have repointed everything else.

route.host

Give the new release its own hostname, for the same reason. An OpenShift Route on a host another Route already holds is not admitted, so the new release publishes nothing.

Nothing else the chart creates needs attention. The Role and RoleBinding grant the same namespace permissions to both ServiceAccounts, which is additive; the ServiceMonitor and PrometheusRule are named per release and carry distinct job and alert names; and the PodDisruptionBudget applies to each release’s own pods.

Keep exactly one Provisioner address registered in Platform Manager. The Provisioner has no reconcile loop and acts only on a request, so the release nobody calls changes nothing. Both releases can read the same KSML applications through /status, because those applications are separate Helm releases keyed by tenant, instance, environment and application rather than by Provisioner. Two registered addresses would let both deploy and undeploy the same applications.

Step 4: Rename the values key if the Provisioner runs as a subchart

Skip this step if you install the Provisioner directly with its own values file. If your settings sit under a ksml-provisioner: key in a parent chart’s values.yaml or in an ArgoCD Application, rename that key to runtime-provisioner: and rename the matching dependency entry in the parent Chart.yaml with it.

runtime-provisioner:
  env:
    - name: NAMESPACE
      value: "<TARGET_NAMESPACE>"
Helm gives no warning for a values key that matches no subchart. The block goes unused, so every setting falls back to the chart default: NAMESPACE becomes default and KSML applications land in the wrong namespace, and connect.namespaces becomes empty so KC log reading finds no pods.

Step 5: Secure the Route or Ingress if you publish the Provisioner

Skip this step if route.enabled and ingress.enabled are both false. If either is true, settle Basic Authentication, the Route path and the Transport Layer Security (TLS) termination in your values file before you install. Step 3 covers the hostname on the same two settings.

Before 0.8.0 neither the Route nor the Ingress carried traffic. Both publish the service from 0.8.0 and BASIC_AUTH_ENABLED defaults to false, so an unauthenticated caller reaching the endpoint can deploy and undeploy KSML applications and read KC worker logs.

Set all three together:

env:
  - name: BASIC_AUTH_ENABLED
    value: "true"
  - name: BASIC_AUTH_USERNAME
    value: "<PROVISIONER_USERNAME>"
  - name: BASIC_AUTH_PASSWORD
    valueFrom:
      secretKeyRef:
        name: <PROVISIONER_AUTH_SECRET>
        key: password

route:
  enabled: true
  path: /                  (1)
  tls:
    termination: edge      (2)
1 / is the new default. A values file still pinning route.path: /api publishes a path that reaches no handler.
2 passthrough is still the default, and it needs SERVER_TLS_CERTIFICATE and SERVER_TLS_PRIVATE_KEY set on the pod. The chart sets neither, so the handshake fails with nothing logged on the Route. edge terminates TLS at the router instead.

Step 6: Install the renamed chart beside the running release

Install the renamed chart as a second release called runtime-provisioner, with the values file you edited. The chart does not repeat its own name when the release already carries it, so the resources come out as runtime-provisioner rather than runtime-provisioner-runtime-provisioner. The command is safe to re-run, because helm upgrade --install converges on the values you pass it.

helm registry login registry.axual.io --username <YOUR_USERNAME>

helm upgrade --install runtime-provisioner \
  oci://registry.axual.io/axual-charts/runtime-provisioner \
  --version 0.8.0 \
  --namespace <NAMESPACE> \
  -f provisioner-values.yaml

Check the new release answers before anything points at it.

kubectl port-forward --namespace <NAMESPACE> svc/runtime-provisioner 8000:80

curl -s localhost:8000/healthz
curl -s localhost:8000/readyz

Both endpoints answer {"status":"ok"} and need no credentials. Port 80 is the chart default for service.port; use your own value if you changed it.

Step 7: Change the Provisioner address in Platform Manager

In Self-Service, open the instance, open its KSML Provisioner configuration, and change the host in the URL to the new Service name. Keep the scheme and the port you use today.

http://<RELEASE>-ksml-provisioner:8000   becomes   http://runtime-provisioner:8000

The field keeps the name KSML Provisioner URL; only the address inside it changes. If you have registered a log viewer address on a KC cluster, change that one too. It is held per KC cluster and points at the same service.

If you move the address with a database migration against ksml_provisioner.url instead of through Self-Service, nothing orders that migration against the chart install. Run it only after the new release answers /readyz, or you reopen the gap the blue-green route exists to close.

Step 8: Repoint monitoring at the new resource name

Nothing in this step is needed for the Provisioner to run. The chart creates its own ServiceMonitor and PrometheusRule under the new name, so scraping and the built-in down alert work from the moment Step 6 finishes. This step is about the queries you wrote yourself: Grafana panels, custom alert rules and Alertmanager silences that name the old resource. Each one carries the resource name, so it stops matching until you point it at the new name. The old values assume a release called axual; the new ones follow the release name you installed in Step 6, shown here as runtime-provisioner:

Where it appears Old value New value

Alert name

<Resource>KsmlProvisionerDown

<Resource>RuntimeProvisionerDown

service label in the alert expression

axual-ksml-provisioner

runtime-provisioner

job label on the Provisioner’s series

axual-ksml-provisioner

runtime-provisioner

container label on pod and resource metrics

ksml-provisioner

runtime-provisioner

app.kubernetes.io/name in dashboard queries

ksml-provisioner

runtime-provisioner

app.kubernetes.io/instance in dashboard queries

axual

runtime-provisioner

service.name in trace queries

ksml-provisioner

runtime-provisioner

otel_scope_name in metric queries

axual.io/otel/ksml/provisioner

axual.io/otel/runtime-provisioner

<Resource> in the alert name is the release name the chart camel-cases into it, so a release called axual gives AxualKsmlProvisionerDown. Take the new name from the PrometheusRule the chart installed rather than deriving it by hand.

The container label also carries the container resource metrics, so a Central Processing Unit (CPU) or memory panel querying container_cpu_usage_seconds_total{container="ksml-provisioner"} goes blank, and kubectl logs -c ksml-provisioner stops working.

Existing Alertmanager silences for the old alert name never fire again. They do not expire with an error, they stop matching. Remove them, and create new ones against the new alert name if you still need them.

Nothing in your KC monitoring can break here. Log reading, its kafkaconnect_log_* metrics and its endpoints are all new in 0.8.0, so you have nothing of your own to repoint for them.

Step 9: Check your NetworkPolicies and admission policies

This one fails without a message, so run the check even if you expect no hits. A NetworkPolicy selecting app.kubernetes.io/name: ksml-provisioner stops selecting anything after the rename. That applies in the Provisioner’s own namespace and in egress policies elsewhere that allow traffic towards it. Traffic is then dropped, or allowed when it should not be, with nothing in any log to say why.

kubectl get networkpolicy --all-namespaces -o yaml | grep -n 'ksml-provisioner'

Repoint every match at runtime-provisioner, then run the same check over any admission policy you operate, for example Kyverno or Gatekeeper rules that name the workload.

Step 10: Verify the upgrade

Confirm the new resources exist and that the pod is ready.

kubectl get all,cm,sa,role,rolebinding --namespace <NAMESPACE> \
  -l app.kubernetes.io/name=runtime-provisioner

kubectl get pods --namespace <NAMESPACE> -l app.kubernetes.io/name=runtime-provisioner

The first command lists the Deployment, Service, ConfigMap, ServiceAccount, Role and RoleBinding under the new name. The second must show the pod Running and 1/1 ready.

/readyz answers 503 when the ServiceAccount cannot list pods in one of the namespaces named in connect.namespaces. Fix that permission before you close the maintenance window, or it surfaces later as an empty KC log window with no obvious cause.

Then finish through Platform Manager:

  1. Open a KSML application in Self-Service and check its status, its logs and its metrics.

  2. Deploy and undeploy one test KSML application.

  3. Confirm the Prometheus target is up with kubectl get servicemonitor --namespace <NAMESPACE>.

Monitoring lags behind by up to about 90 seconds, because the Prometheus Operator has to notice the renamed ServiceMonitor, write a new configuration and scrape once. A gap of that length in the graphs is expected and is not a fault.

Step 11: Remove the old release

Remove the old release only once Step 10 passes. Until then it is your rollback, and Prometheus scrapes both releases.

helm uninstall <RELEASE> --namespace <NAMESPACE>

kubectl get all --namespace <NAMESPACE> -l app.kubernetes.io/name=ksml-provisioner

The second command must report No resources found.

helm uninstall removes only the Provisioner’s own objects: its Deployment, Service, ServiceAccount, Role, RoleBinding and ConfigMap, plus its ServiceMonitor, PrometheusRule, Ingress, Route, HorizontalPodAutoscaler and PodDisruptionBudget where you enabled them. Your KSML applications are separate Helm releases with their own StatefulSets or Jobs, PersistentVolumeClaims and Secrets, so none of them is removed and no data is lost. The new Provisioner finds them again by their axual.io/tenant and axual.io/instance labels.

Roll back to KSML Provisioner 0.7.0

Before Step 11 your old release is still running, so put the old address back in Platform Manager and remove the release you added.

helm uninstall runtime-provisioner --namespace <NAMESPACE>

After Step 11 the old release is gone. Remove the new release and install 0.7.0 again under the old name, with the values file you saved in Step 1. Set image.tag back to 0.7.0 in that file first, because 0.7.0 is the only release published under the old chart name.

helm uninstall runtime-provisioner --namespace <NAMESPACE>

helm upgrade --install <RELEASE> \
  oci://registry.axual.io/axual-charts/ksml-provisioner \
  --version 0.7.0 \
  --namespace <NAMESPACE> \
  -f provisioner-values.yaml

Your KSML applications are not part of this release, so a rollback does not affect them.

Confirm the rollback in either case. The pod of the release that is left must show Running and 1/1 ready. A KSML application in Self-Service must report its status again.

kubectl get pods --namespace <NAMESPACE> -l app.kubernetes.io/name=ksml-provisioner
Going back to 0.7.0 removes KC log reading. That release has no /kafkaconnect/logs, no download endpoint and no Connect Role. The log viewer registration in Platform Manager stops working until you come forward again. KSML provisioning is unaffected.