Troubleshoot a Kafka Connect Deployment

This guide covers the failures a Kafka Connect deployment produces, symptom by symptom: a worker that crash-loops or cannot pull its image, a connector that reports running but moves nothing, a rejected registration, and a failing Vault probe.

Type

How-to guide

Goal

Identify and fix why a Kafka Connect cluster or one of its connectors is failing.

Audience

Platform Operator with read access to the Connect namespace and the Vault it reads.

When to use

Use this guide when a deployment step fails, or when a running connector stops working.

Each section names one observable failure state in a Kafka Connect (KC) deployment, its cause, and the fix. For the deployment procedures themselves, see How to Deploy a Kafka Connect Cluster.

Collect diagnostics before investigating

Gather the worker’s configuration, pod state and recent logs before looking at any specific symptom below. Every scenario reads from that same set.

Replace <…​> placeholders (angle brackets) with your values before running a command. <tenant>, <instance>, and <cluster-name> are placeholders for the cluster’s real short names (for example axual, dev, my-cluster). <connector-vault-path> is the Vault path holding this cluster’s connector credentials, shown as Vault path on its page in Self-Service. The commands assume fullnameOverride equals clusterName, as How to Deploy a Kafka Connect Cluster sets it, so the KafkaConnect resource, its pods and the -acl-bootstrap Job all carry <cluster-name>.

Run these commands:

# KafkaConnect Custom Resource (CR) condition
kubectl describe kafkaconnect/<cluster-name> --namespace <ns> | grep -A 10 Conditions

# Worker pod events and state
kubectl get pods -l strimzi.io/cluster=<cluster-name> --namespace <ns>
kubectl describe pod -l strimzi.io/cluster=<cluster-name> --namespace <ns>

# Worker pod logs (last 100 lines)
kubectl logs -l strimzi.io/cluster=<cluster-name> --namespace <ns> --tail=100

# aclBootstrap Job logs (if enabled)
kubectl logs job/<cluster-name>-acl-bootstrap --namespace <ns> --tail=60

The workers write one ECS JSON object per log entry, so kubectl logs shows raw JSON. The grep commands below work on it unchanged. To read the output by hand, pipe it through jq, which also passes through the plain-text lines Strimzi writes before log4j2 starts:

kubectl logs -l strimzi.io/cluster=<cluster-name> --namespace <ns> --tail=100 \
  | jq -R -r '. as $l | try (fromjson | "\(.["@timestamp"]) \(.["log.level"]) \(.message)") catch $l'

Worker pod CrashLoops with invalid role or secret ID

Cause: The worker AppRole credentials in the Kubernetes Secret named in vault.credentialsSecret are wrong or expired.

Fix:

  1. Verify the Kubernetes Secret content matches the AppRole credentials from Vault:

    kubectl get secret <cluster-name>-vault-creds \
      --namespace <ns> \
      -o jsonpath='{.data.VAULT_APPROLE_ROLE_ID}' | base64 -d

    Compare the output against the role_id from Vault:

    vault read auth/approle/role/<tenant>-<instance>-<cluster-name>-connect/role-id
  2. If the Secret is wrong, delete and recreate it, then follow the Vault credential steps in How to Deploy a Kafka Connect Cluster.

Worker pod stuck in ImagePullBackOff

Cause: A plugin OCI image is missing from the registry, or the Harbor pull Secret is missing or invalid.

Fix:

  1. Check which image is failing:

    kubectl describe pod -l strimzi.io/cluster=<cluster-name> --namespace <ns> \
      | grep -A 5 "Failed to pull"
  2. Confirm the plugin image exists in Harbor before adding it to plugins. Publish it via the CI pipeline first.

  3. For development images from the private internal Harbor project, confirm regcred-harbor exists in the namespace:

    kubectl get secret regcred-harbor --namespace <ns>

    If missing, create it by following the registry credential step in How to Deploy a Kafka Connect Cluster.

kubectl rollout status deploy/…​ returns NotFound

Cause: Strimzi does not create a Deployment for KC workers. It creates a StrimziPodSet instead.

Fix: Use the correct health check command:

kubectl wait kafkaconnect/<cluster-name> \
  --for=condition=Ready \
  --namespace <ns> \
  --timeout=120s

helm upgrade fails with this cluster’s …​ is …​ but these values give …​

Cause: The values give a different groupId, configStorageTopic, offsetStorageTopic or statusStorageTopic than the running KafkaConnect resource uses. The four names derive from tenant, instance and clusterName (for example _<tenant>-<instance>-<clusterName>-connect-configs), so changing any of the three moves them too. Applying the change would leave the worker on empty topics and lose its connector configurations, offsets and status, so the chart stops the upgrade during rendering and applies nothing.

Fix: Restore the original tenant, instance and clusterName, or pin the four names in the values to what the cluster runs with, then upgrade again. Read the current names off the live resource:

kubectl get kafkaconnect <cluster-name> --namespace <ns> \
  -o jsonpath='{.spec.groupId}{"\n"}{.spec.configStorageTopic}{"\n"}{.spec.offsetStorageTopic}{"\n"}{.spec.statusStorageTopic}{"\n"}'
The check does nothing when the identity running helm upgrade cannot read KafkaConnect resources, for example a scoped Argo CD service account. In that case the upgrade goes through and the worker starts on new, empty topics, so do not change the identity values of a running cluster.

KafkaConnect never reports Ready with the NetworkPolicy on

Cause: restApi.networkPolicy.enabled: true renders a NetworkPolicy named <cluster-name>-rest-api that admits port 8083 only from the cluster’s own workers, the Strimzi operator, the router pods and allowedFrom. The Strimzi operator calls the REST API on every reconcile, so the resource never turns Ready when restApi.networkPolicy.strimziOperator does not match the operator pod. On a CNI that does not let the kubelet health probe through by default, the workers also never turn Ready.

Fix:

  1. Check the operator’s namespace and labels against restApi.networkPolicy.strimziOperator (default label strimzi.io/kind: cluster-operator):

    kubectl get pod -A -l strimzi.io/kind=cluster-operator --show-labels
  2. If the pod readiness probe fails while the operator is allowed, add the node address range to restApi.networkPolicy.allowedFrom.

  3. Run helm upgrade --install with the corrected values.

A wrong restApi.networkPolicy.router selector fails the same way for the route: requests through the Ingress or HTTPRoute time out. Find the router pods with kubectl get pod -n <router-namespace> --show-labels and copy the real labels.

Unauthenticated call to the REST API route does not return 401

Cause: restApi.route.authMethod: basic only checks that something carries the authentication (restApi.basicAuth.annotations, restApi.route.filters or restApi.extraObjects), not that it works. Annotations for a controller you do not run, a filter naming a CRD that is not installed, or a misspelt Secret name all render and leave the route open.

Fix:

  1. Call the route with no credentials:

    curl -o /dev/null -s -w '%{http_code}\n' https://<connect-api-host>/connectors

    Anything other than 401 means the REST API is open.

  2. Match the wiring to the implementation that serves the route, using the example for it in The chart does not know your ingress implementation. The F5 NGINX Ingress Controller also needs restApi.basicAuth.secretType: nginx.org/htpasswd and secretKey: htpasswd.

  3. Confirm the direct path is closed from any other pod, which proves the NetworkPolicy is enforced:

    kubectl run probe --rm -it --image=curlimages/curl --restart=Never --namespace <ns> -- \
      curl -m 5 http://<cluster-name>-connect-api.<ns>.svc:8083/connectors

    The call must time out. A reply means the CNI does not enforce the policy.

Worker pods refused when runtimeClassName is set

Cause: runtimeClassName renders a namespaced Kyverno Policy that sets the runtime on the worker pods at admission, with failurePolicy: Fail. While Kyverno is unavailable, new worker pods are refused rather than started on the shared kernel. A worker that starts but stays Pending has no node offering the runtime, because affinity is not set or names the wrong node label.

Fix: Confirm Kyverno is running, and that affinity targets the node pool that provides the runtime. Existing workers keep their old runtime until replaced, so roll them once the policy is in place:

kubectl annotate --namespace <ns> strimzipodset <cluster-name>-connect strimzi.io/manual-rolling-update=true
kubectl get pod --namespace <ns> -l strimzi.io/cluster=<cluster-name> \
  -o jsonpath='{.items[*].spec.runtimeClassName}'

Source connector shows RUNNING but produces no data, UNKNOWN_TOPIC_OR_PARTITION in logs

Cause: The connector’s target topic does not exist. Axual brokers do not auto-create topics, and the worker principal has no Create authority to make one (see How to Prepare the Kafka Broker for Kafka Connect).

Fix: Provision the target topic through Self-Service, then confirm the connector’s application holds Write (source) or Read (sink) access to it. Platform Manager grants those Access Control Lists (ACLs) when the topic is authorised for the connector.

Connect’s topic.creation.* settings do not resolve this on a governed Axual Platform cluster. They require the worker to hold Create, which governance withholds by design so connectors cannot bypass Self-Service topic provisioning.

Creating a connector returns HTTP 500 with 404 in the response body

Cause: A ${vault:…​} placeholder in the connector configuration references a path or key that does not exist in the Connector Vault. The worker resolves the placeholder when you create the connector, so the error appears immediately.

Fix:

  1. Identify the failing Vault path from the worker log:

    kubectl logs -l strimzi.io/cluster=<cluster-name> --namespace <ns> \
      | grep -iE "vault|404|VaultConfigurationProvider"
  2. Verify the secret exists in Vault at the exact path referenced, where <connector-vault-path> is the cluster’s Vault path from its Self-Service registration (see Understanding the Vault KV v2 path structure):

    vault kv get connector-auth/<connector-vault-path>/<tenant>/<instance>/<cluster-name>/tls/<env>/<app>

    Read the cluster’s connectorVaultPath off its Self-Service registration form if you do not know it. Leaving it out is the most common reason this lookup reports a missing secret that is in fact present, because the secret sits under that prefix rather than directly under the mount. See Understanding the Vault KV v2 path structure for the full path layout, including the SASL and Schema Registry paths.

  3. If the secret is missing, Platform Manager has not yet written the connector credential. Complete the Self-Service connector registration flow first.

  4. If the path is wrong, fix the Vault path in the connector configuration and re-create the connector.

Vault probe connector fails with x509, connection, or 403

Cause: The worker reached the Connector Vault but the TLS trust chain, network path, or AppRole policy is wrong. This is distinct from a missing secret, which returns HTTP 500 with 404 in the body (see the previous scenario).

Fix: Read the worker log and match the error:

kubectl logs -l strimzi.io/cluster=<cluster-name> --namespace <ns> \
  | grep -iE "vault|x509|403|VaultConfigurationProvider"
  • x509 or PKIX: the worker does not trust the Vault server certificate. Confirm vault.truststoreSecret holds the Vault Certificate Authority (CA), and that vault.address uses a hostname in the certificate Subject Alternative Name (SAN), not the internal service name.

  • Connection refused or timeout: the worker cannot reach vault.address. Check the network policy and that the address is reachable from the worker pod.

  • 403: the AppRole policy does not cover the requested path. Recheck the per-cluster read policy in step 4 of How to Set Up the Connector Vault for Kafka Connect.

A probe connector that resolves correctly returns HTTP 201. A 500 with 404 in the body means trust and authentication work but the secret is missing.

Connector fails authentication while its credential exists in Vault

The connector reports RUNNING while its task is FAILED, and the task’s trace ends in a broker authentication error such as SaslAuthenticationException: Authentication failed during authentication due to invalid credentials with SASL mechanism SCRAM-SHA-512, or the mutual TLS (mTLS) equivalent. The credential is present in Vault and the worker itself is connected to the broker.

Cause: The worker has no Vault configuration provider registered, so the ${vault:…​} placeholder in the connector configuration was never resolved. A placeholder whose provider is not registered is not an error in Kafka Connect: the value is passed through exactly as written, so the connector authenticates with the literal string ${vault:…​} as its password. The broker sees a wrong credential and rejects it.

This failure is quiet, and the three marks below tell it apart from a genuinely wrong credential:

  • The worker itself is unaffected, because its own credentials come from a Kubernetes Secret rather than from Vault. Only connectors fail.

  • No Vault error appears anywhere in the worker log. Nothing calls Vault, so there is nothing to find.

  • Regenerating the credential in Self-Service changes nothing, however many times it is regenerated.

The usual root cause is that vault.enabled is not true in the values the cluster is running, most often because the values were updated in version control but the change never reached the cluster. The chart refuses to render when any vault.* value is set while vault.enabled is false, and prints a warning after every install without Vault, so a values file with no vault block at all is the case to look for.

Fix:

  1. Check whether the worker’s generated configuration registers a provider named vault:

    kubectl get configmap <cluster-name>-connect-config --namespace <ns> \
      -o jsonpath='{.data.kafka-connect\.properties}' | grep -E "^config\.providers"

    Expect a config.providers list containing vault, together with a config.providers.vault.class entry. A list holding only the Strimzi providers confirms this scenario: the placeholder cannot resolve, whatever is in Vault.

  2. Confirm the worker pod received the rest of the wiring:

    kubectl get pod <cluster-name>-connect-0 --namespace <ns> \
      -o jsonpath='{.spec.containers[0].env[*].name}{"\n"}{.spec.containers[0].volumeMounts[*].mountPath}{"\n"}'

    Expect VAULT_APPROLE_ROLE_ID and VAULT_APPROLE_SECRET_ID among the environment variables, and /mnt/vault-truststore among the mount paths when vault.truststoreSecret is set.

  3. Set vault.enabled: true and the credentials Secret in the cluster values, as in How to Deploy a Kafka Connect Cluster, then apply them.

  4. Confirm the change reached the cluster rather than only the values file:

    kubectl get kafkaconnect <cluster-name> --namespace <ns> \
      -o jsonpath='{.metadata.generation}{"\n"}'

    A generation still at 1, or one that has not moved since the values changed, means the cluster is running an older resource than the one in version control. Look at whatever applies the values, such as the continuous deployment application that owns the cluster, rather than at the values themselves.

  5. Once the worker has rolled with the provider registered, restart the failed task:

    curl -X POST -u "<connect-user>:<connect-password>" \
      "https://<connect-api-host>/connectors/<connector-name>/tasks/0/restart"

    The credential already in Vault is reused, so there is no need to regenerate it.

When two Kafka Connect clusters are deployed side by side and only one has working connectors, compare their two -connect-config ConfigMaps before anything else. The config.providers line is where they differ.

Simple Authentication and Security Layer (SASL) worker CrashLoops with TopicAuthorizationException: Not authorized to access topics

Cause: The SASL user authenticated successfully but holds no ACLs for the internal topics or consumer group.

Fix: Grant the required ACLs. See How to Prepare the Kafka Broker for Kafka Connect.

On Strimzi-managed brokers, add spec.authorization to the KafkaUser CR:

spec:
  authorization:
    type: simple
    acls:
      - resource:
          type: topic
          name: <configStorageTopic>
          patternType: literal
        operations: [Read, Write, Describe, DescribeConfigs]
      # Repeat for offsetStorageTopic and statusStorageTopic
      - resource:
          type: group
          name: <groupId>
          patternType: literal
        operations: [Read, Describe]

aclBootstrap Job fails with FAILED (exit …​)

Cause: The superuser credentials are wrong, the broker is unreachable, or aclBootstrap.principal does not match the worker identity exactly. A line FAILED: <value> is blank means a value override set a required setting to "". The Job also fails when it runs longer than aclBootstrap.activeDeadlineSeconds (default 240), which usually means the broker is unreachable.

Fix:

  1. Read the full Job logs:

    kubectl logs job/<cluster-name>-acl-bootstrap \
      --namespace <ns> \
      --tail=100

    helm install and helm upgrade wait for the Job but do not stream its logs. Follow them live from a second terminal with kubectl logs -f. The completed Job and its pod stay until the next run.

  2. Check the logged bootstrap address and superuser auth type match your broker.

  3. Confirm aclBootstrap.principal is the exact Distinguished Name (DN) (for mutual TLS (mTLS): User:CN=…​,O=…​,C=…​) or username (for SASL: User:<username>).

  4. The Job is idempotent. After fixing the values, re-run:

    helm upgrade <CLUSTER_NAME> \
      oci://registry.axual.io/axual-charts/kafka-connect \
      --version <CHART_VERSION> \
      --namespace <NAMESPACE> \
      -f connect-values.yaml

Self-Service cluster registration fails

Cause: One of three Platform Manager validation checks failed: the REST API is unreachable, Platform Manager cannot write to the Connector Vault (its AppRole policy does not cover the cluster’s paths, or on-premises writer credentials are wrong), or kafka-topic-name-transforms is not listed in the worker’s plugins. The Vault check runs as the kafka-connects:validate-vault pre-flight check and, when it fails, returns an error of the form: The supplied Vault AppRole cannot write to path '…​'. Grant 'create' or 'update' capability on this path (capabilities: [deny]). Three mistakes produce this error. All of them leave the AppRole without access to connector-auth/data/<connector-vault-path>/, which is where the check looks:

  • The policy grants the bare path with no trailing /*. Vault matches these patterns literally, so such a grant covers nothing.

  • The policy grants a path further down, naming the tenant, instance or cluster. The check looks just below <connector-vault-path> and never reaches those. Grant connector-auth/data/<connector-vault-path>/* instead.

  • The registered connectorVaultPath already includes the connector-auth mount, or the tenant, instance and cluster name. Platform Manager adds those itself, so the path ends up doubled and no policy matches it. See Understanding the Vault KV v2 path structure.

Fix:

  1. Confirm the worker REST API is reachable from the Platform Manager pod:

    kubectl exec -n <pm-namespace> <pm-pod> -- \
      curl -s -u "<connect-user>:<connect-password>" \
      "<connect-url>/connector-plugins?connectorsOnly=false" | jq '.[].class'

    Use the connectUrl registered in Self-Service as <connect-url>. With restApi.networkPolicy.enabled, it must be the route address: the policy blocks Platform Manager from the Service on port 8083, so a direct call there times out by design.

    Omit -u only for a cluster still running restApi.route.authMethod: none. A 401 here on a cluster meant to run Basic Auth means the credentials Platform Manager holds do not match restApi.basicAuth, not that the API is unreachable. Platform Manager reads the password from the Governance Vault key connect-api-password, so it has to match the htpasswd entry the route checks.
  2. Confirm kafka-topic-name-transforms appears in the output. If missing, add it to plugins and re-run helm upgrade --install.

  3. Confirm Platform Manager can write to the Connector Vault. On Axual Cloud, check that its AppRole policy grants connector-auth/data/<connector-vault-path>/, with the trailing / and nothing further down the path (step 4 of the Vault setup how-to). On strict on-premises, re-check the writer role_id and secret_id against Vault:

    vault read auth/approle/role/<tenant>-<instance>-<cluster-name>-connector-writer/role-id
    If the policy or credentials are wrong, follow step 4 of
    xref:components/run-time/kafka-connect/how-to-set-up-connector-vault.adoc[How to Set Up the Connector Vault for Kafka Connect]
    and, on-premises, provide the new credentials to the Tenant Admin.

Connector logs do not appear in Self-Service

Cause: The connector runs, but the Log Provisioner cannot return its logs. Either it is not reachable at the URL registered for the Kubernetes cluster, or it is reachable and finds no pods. It locates the pods by the axual.io/tenant, axual.io/instance and axual.io/connect-cluster labels, so it finds none when one of those values differs from what Self-Service holds for the cluster. The chart renders the labels from the tenant, instance and clusterName values, and the log viewer lowercases the Self-Service names before matching, so the values must be lowercase.

Fix:

  1. Confirm the Log Provisioner pod is running:

    kubectl -n <ns> get pods -l app.kubernetes.io/name=runtime-provisioner
  2. Read the labels the chart rendered onto the worker pods:

    kubectl -n <ns> get pods -l axual.io/connect-cluster=<cluster-name> \
      -o jsonpath='{range .items[*]}{.metadata.labels.axual\.io/tenant}{" "}{.metadata.labels.axual\.io/instance}{" "}{.metadata.labels.axual\.io/connect-cluster}{"\n"}{end}'

    An empty result means the labels do not carry the values you expect, which is the same failure the Log Provisioner hits.

  3. Compare the three values against the lowercased tenant short name, instance short name and cluster name in Self-Service. A cluster name longer than 63 characters cannot be a label value, so shorten it in Self-Service first. A mismatch in any one of them hides every log for the cluster. The labels come from the chart values, so correct them there and run helm upgrade --install rather than editing the pods.

  4. Confirm the Log Provisioner URL registered against the Kubernetes cluster in Self-Service is the one the pod serves. See How to Enable Kafka Connect Log Reading.

For what the Log Provisioner does and why it matches on labels, see How the Log Provisioner finds connector logs.