Troubleshoot a Kafka Connect Deployment
This guide covers the failures a Kafka Connect deployment produces, symptom by symptom: a worker that crash-loops or cannot pull its image, a connector that reports running but moves nothing, a rejected registration, and a failing Vault probe.
Type |
How-to guide |
Goal |
Identify and fix why a Kafka Connect cluster or one of its connectors is failing. |
Audience |
Platform Operator with read access to the Connect namespace and the Vault it reads. |
When to use |
Use this guide when a deployment step fails, or when a running connector stops working. |
Each section names one observable failure state in a Kafka Connect (KC) deployment, its cause, and the fix. For the deployment procedures themselves, see How to Deploy a Kafka Connect Cluster.
Contents
The sections below cover each task in this guide:
-
helm upgradefails withthis cluster’s … is … but these values give … -
Unauthenticated call to the REST API route does not return
401 -
Source connector shows
RUNNINGbut produces no data,UNKNOWN_TOPIC_OR_PARTITIONin logs -
Creating a connector returns HTTP 500 with
404in the response body -
Connector fails authentication while its credential exists in Vault
Collect diagnostics before investigating
Gather the worker’s configuration, pod state and recent logs before looking at any specific symptom below. Every scenario reads from that same set.
Replace <…> placeholders (angle brackets) with your values before running a command.
<tenant>, <instance>, and <cluster-name> are placeholders for the cluster’s real short names (for example axual, dev, my-cluster).
<connector-vault-path> is the Vault path holding this cluster’s connector credentials, shown as Vault path on its page in Self-Service.
The commands assume fullnameOverride equals clusterName, as How to Deploy a Kafka Connect Cluster sets it, so the KafkaConnect resource, its pods and the -acl-bootstrap Job all carry <cluster-name>.
|
Run these commands:
# KafkaConnect Custom Resource (CR) condition
kubectl describe kafkaconnect/<cluster-name> --namespace <ns> | grep -A 10 Conditions
# Worker pod events and state
kubectl get pods -l strimzi.io/cluster=<cluster-name> --namespace <ns>
kubectl describe pod -l strimzi.io/cluster=<cluster-name> --namespace <ns>
# Worker pod logs (last 100 lines)
kubectl logs -l strimzi.io/cluster=<cluster-name> --namespace <ns> --tail=100
# aclBootstrap Job logs (if enabled)
kubectl logs job/<cluster-name>-acl-bootstrap --namespace <ns> --tail=60
The workers write one ECS JSON object per log entry, so kubectl logs shows raw JSON.
The grep commands below work on it unchanged.
To read the output by hand, pipe it through jq, which also passes through the plain-text lines Strimzi writes before log4j2 starts:
kubectl logs -l strimzi.io/cluster=<cluster-name> --namespace <ns> --tail=100 \
| jq -R -r '. as $l | try (fromjson | "\(.["@timestamp"]) \(.["log.level"]) \(.message)") catch $l'
Worker pod CrashLoops with invalid role or secret ID
Cause: The worker AppRole credentials in the Kubernetes Secret named in vault.credentialsSecret are wrong or expired.
Fix:
-
Verify the Kubernetes Secret content matches the AppRole credentials from Vault:
kubectl get secret <cluster-name>-vault-creds \ --namespace <ns> \ -o jsonpath='{.data.VAULT_APPROLE_ROLE_ID}' | base64 -dCompare the output against the
role_idfrom Vault:vault read auth/approle/role/<tenant>-<instance>-<cluster-name>-connect/role-id -
If the Secret is wrong, delete and recreate it, then follow the Vault credential steps in How to Deploy a Kafka Connect Cluster.
Worker pod stuck in ImagePullBackOff
Cause: A plugin OCI image is missing from the registry, or the Harbor pull Secret is missing or invalid.
Fix:
-
Check which image is failing:
kubectl describe pod -l strimzi.io/cluster=<cluster-name> --namespace <ns> \ | grep -A 5 "Failed to pull" -
Confirm the plugin image exists in Harbor before adding it to
plugins. Publish it via the CI pipeline first. -
For development images from the private
internalHarbor project, confirmregcred-harborexists in the namespace:kubectl get secret regcred-harbor --namespace <ns>If missing, create it by following the registry credential step in How to Deploy a Kafka Connect Cluster.
kubectl rollout status deploy/… returns NotFound
Cause: Strimzi does not create a Deployment for KC workers.
It creates a StrimziPodSet instead.
Fix: Use the correct health check command:
kubectl wait kafkaconnect/<cluster-name> \
--for=condition=Ready \
--namespace <ns> \
--timeout=120s
helm upgrade fails with this cluster’s … is … but these values give …
Cause: The values give a different groupId, configStorageTopic, offsetStorageTopic or statusStorageTopic than the running KafkaConnect resource uses.
The four names derive from tenant, instance and clusterName (for example _<tenant>-<instance>-<clusterName>-connect-configs), so changing any of the three moves them too.
Applying the change would leave the worker on empty topics and lose its connector configurations, offsets and status, so the chart stops the upgrade during rendering and applies nothing.
Fix: Restore the original tenant, instance and clusterName, or pin the four names in the values to what the cluster runs with, then upgrade again.
Read the current names off the live resource:
kubectl get kafkaconnect <cluster-name> --namespace <ns> \
-o jsonpath='{.spec.groupId}{"\n"}{.spec.configStorageTopic}{"\n"}{.spec.offsetStorageTopic}{"\n"}{.spec.statusStorageTopic}{"\n"}'
The check does nothing when the identity running helm upgrade cannot read KafkaConnect resources, for example a scoped Argo CD service account.
In that case the upgrade goes through and the worker starts on new, empty topics, so do not change the identity values of a running cluster.
|
KafkaConnect never reports Ready with the NetworkPolicy on
Cause: restApi.networkPolicy.enabled: true renders a NetworkPolicy named <cluster-name>-rest-api that admits port 8083 only from the cluster’s own workers, the Strimzi operator, the router pods and allowedFrom.
The Strimzi operator calls the REST API on every reconcile, so the resource never turns Ready when restApi.networkPolicy.strimziOperator does not match the operator pod.
On a CNI that does not let the kubelet health probe through by default, the workers also never turn Ready.
Fix:
-
Check the operator’s namespace and labels against
restApi.networkPolicy.strimziOperator(default labelstrimzi.io/kind: cluster-operator):kubectl get pod -A -l strimzi.io/kind=cluster-operator --show-labels -
If the pod readiness probe fails while the operator is allowed, add the node address range to
restApi.networkPolicy.allowedFrom. -
Run
helm upgrade --installwith the corrected values.
A wrong restApi.networkPolicy.router selector fails the same way for the route: requests through the Ingress or HTTPRoute time out.
Find the router pods with kubectl get pod -n <router-namespace> --show-labels and copy the real labels.
Unauthenticated call to the REST API route does not return 401
Cause: restApi.route.authMethod: basic only checks that something carries the authentication (restApi.basicAuth.annotations, restApi.route.filters or restApi.extraObjects), not that it works.
Annotations for a controller you do not run, a filter naming a CRD that is not installed, or a misspelt Secret name all render and leave the route open.
Fix:
-
Call the route with no credentials:
curl -o /dev/null -s -w '%{http_code}\n' https://<connect-api-host>/connectorsAnything other than
401means the REST API is open. -
Match the wiring to the implementation that serves the route, using the example for it in The chart does not know your ingress implementation. The F5 NGINX Ingress Controller also needs
restApi.basicAuth.secretType: nginx.org/htpasswdandsecretKey: htpasswd. -
Confirm the direct path is closed from any other pod, which proves the
NetworkPolicyis enforced:kubectl run probe --rm -it --image=curlimages/curl --restart=Never --namespace <ns> -- \ curl -m 5 http://<cluster-name>-connect-api.<ns>.svc:8083/connectorsThe call must time out. A reply means the CNI does not enforce the policy.
Worker pods refused when runtimeClassName is set
Cause: runtimeClassName renders a namespaced Kyverno Policy that sets the runtime on the worker pods at admission, with failurePolicy: Fail.
While Kyverno is unavailable, new worker pods are refused rather than started on the shared kernel.
A worker that starts but stays Pending has no node offering the runtime, because affinity is not set or names the wrong node label.
Fix: Confirm Kyverno is running, and that affinity targets the node pool that provides the runtime.
Existing workers keep their old runtime until replaced, so roll them once the policy is in place:
kubectl annotate --namespace <ns> strimzipodset <cluster-name>-connect strimzi.io/manual-rolling-update=true
kubectl get pod --namespace <ns> -l strimzi.io/cluster=<cluster-name> \
-o jsonpath='{.items[*].spec.runtimeClassName}'
Source connector shows RUNNING but produces no data, UNKNOWN_TOPIC_OR_PARTITION in logs
Cause: The connector’s target topic does not exist.
Axual brokers do not auto-create topics, and the worker principal has no Create authority to make one (see How to Prepare the Kafka Broker for Kafka Connect).
Fix: Provision the target topic through Self-Service, then confirm the connector’s application holds Write (source) or Read (sink) access to it.
Platform Manager grants those Access Control Lists (ACLs) when the topic is authorised for the connector.
Connect’s topic.creation.* settings do not resolve this on a governed Axual Platform cluster.
They require the worker to hold Create, which governance withholds by design so connectors cannot bypass Self-Service topic provisioning.
|
Creating a connector returns HTTP 500 with 404 in the response body
Cause: A ${vault:…} placeholder in the connector configuration references a path or key that does not exist in the Connector Vault.
The worker resolves the placeholder when you create the connector, so the error appears immediately.
Fix:
-
Identify the failing Vault path from the worker log:
kubectl logs -l strimzi.io/cluster=<cluster-name> --namespace <ns> \ | grep -iE "vault|404|VaultConfigurationProvider" -
Verify the secret exists in Vault at the exact path referenced, where
<connector-vault-path>is the cluster’s Vault path from its Self-Service registration (see Understanding the Vault KV v2 path structure):vault kv get connector-auth/<connector-vault-path>/<tenant>/<instance>/<cluster-name>/tls/<env>/<app>Read the cluster’s
connectorVaultPathoff its Self-Service registration form if you do not know it. Leaving it out is the most common reason this lookup reports a missing secret that is in fact present, because the secret sits under that prefix rather than directly under the mount. See Understanding the Vault KV v2 path structure for the full path layout, including the SASL and Schema Registry paths. -
If the secret is missing, Platform Manager has not yet written the connector credential. Complete the Self-Service connector registration flow first.
-
If the path is wrong, fix the Vault path in the connector configuration and re-create the connector.
Vault probe connector fails with x509, connection, or 403
Cause: The worker reached the Connector Vault but the TLS trust chain, network path, or AppRole policy is wrong.
This is distinct from a missing secret, which returns HTTP 500 with 404 in the body (see the previous scenario).
Fix: Read the worker log and match the error:
kubectl logs -l strimzi.io/cluster=<cluster-name> --namespace <ns> \
| grep -iE "vault|x509|403|VaultConfigurationProvider"
-
x509orPKIX: the worker does not trust the Vault server certificate. Confirmvault.truststoreSecretholds the Vault Certificate Authority (CA), and thatvault.addressuses a hostname in the certificate Subject Alternative Name (SAN), not the internal service name. -
Connection refused or timeout: the worker cannot reach
vault.address. Check the network policy and that the address is reachable from the worker pod. -
403: the AppRole policy does not cover the requested path. Recheck the per-cluster read policy in step 4 of How to Set Up the Connector Vault for Kafka Connect.
A probe connector that resolves correctly returns HTTP 201.
A 500 with 404 in the body means trust and authentication work but the secret is missing.
|
Connector fails authentication while its credential exists in Vault
The connector reports RUNNING while its task is FAILED, and the task’s trace ends in a broker authentication error such as
SaslAuthenticationException: Authentication failed during authentication due to invalid credentials with SASL mechanism SCRAM-SHA-512,
or the mutual TLS (mTLS) equivalent. The credential is present in Vault and the worker itself is connected to the broker.
Cause: The worker has no Vault configuration provider registered, so the ${vault:…} placeholder in the connector configuration was never resolved.
A placeholder whose provider is not registered is not an error in Kafka Connect: the value is passed through exactly as written, so the connector authenticates with the literal string ${vault:…} as its password.
The broker sees a wrong credential and rejects it.
This failure is quiet, and the three marks below tell it apart from a genuinely wrong credential:
-
The worker itself is unaffected, because its own credentials come from a Kubernetes Secret rather than from Vault. Only connectors fail.
-
No Vault error appears anywhere in the worker log. Nothing calls Vault, so there is nothing to find.
-
Regenerating the credential in Self-Service changes nothing, however many times it is regenerated.
The usual root cause is that vault.enabled is not true in the values the cluster is running, most often because the values were updated in version control but the change never reached the cluster.
The chart refuses to render when any vault.* value is set while vault.enabled is false, and prints a warning after every install without Vault, so a values file with no vault block at all is the case to look for.
Fix:
-
Check whether the worker’s generated configuration registers a provider named
vault:kubectl get configmap <cluster-name>-connect-config --namespace <ns> \ -o jsonpath='{.data.kafka-connect\.properties}' | grep -E "^config\.providers"Expect a
config.providerslist containingvault, together with aconfig.providers.vault.classentry. A list holding only the Strimzi providers confirms this scenario: the placeholder cannot resolve, whatever is in Vault. -
Confirm the worker pod received the rest of the wiring:
kubectl get pod <cluster-name>-connect-0 --namespace <ns> \ -o jsonpath='{.spec.containers[0].env[*].name}{"\n"}{.spec.containers[0].volumeMounts[*].mountPath}{"\n"}'Expect
VAULT_APPROLE_ROLE_IDandVAULT_APPROLE_SECRET_IDamong the environment variables, and/mnt/vault-truststoreamong the mount paths whenvault.truststoreSecretis set. -
Set
vault.enabled: trueand the credentials Secret in the cluster values, as in How to Deploy a Kafka Connect Cluster, then apply them. -
Confirm the change reached the cluster rather than only the values file:
kubectl get kafkaconnect <cluster-name> --namespace <ns> \ -o jsonpath='{.metadata.generation}{"\n"}'A generation still at
1, or one that has not moved since the values changed, means the cluster is running an older resource than the one in version control. Look at whatever applies the values, such as the continuous deployment application that owns the cluster, rather than at the values themselves. -
Once the worker has rolled with the provider registered, restart the failed task:
curl -X POST -u "<connect-user>:<connect-password>" \ "https://<connect-api-host>/connectors/<connector-name>/tasks/0/restart"The credential already in Vault is reused, so there is no need to regenerate it.
When two Kafka Connect clusters are deployed side by side and only one has working connectors, compare their two -connect-config ConfigMaps before anything else. The config.providers line is where they differ.
|
Simple Authentication and Security Layer (SASL) worker CrashLoops with TopicAuthorizationException: Not authorized to access topics
Cause: The SASL user authenticated successfully but holds no ACLs for the internal topics or consumer group.
Fix: Grant the required ACLs. See How to Prepare the Kafka Broker for Kafka Connect.
On Strimzi-managed brokers, add spec.authorization to the KafkaUser CR:
spec:
authorization:
type: simple
acls:
- resource:
type: topic
name: <configStorageTopic>
patternType: literal
operations: [Read, Write, Describe, DescribeConfigs]
# Repeat for offsetStorageTopic and statusStorageTopic
- resource:
type: group
name: <groupId>
patternType: literal
operations: [Read, Describe]
aclBootstrap Job fails with FAILED (exit …)
Cause: The superuser credentials are wrong, the broker is unreachable, or aclBootstrap.principal does not match the worker identity exactly.
A line FAILED: <value> is blank means a value override set a required setting to "".
The Job also fails when it runs longer than aclBootstrap.activeDeadlineSeconds (default 240), which usually means the broker is unreachable.
Fix:
-
Read the full Job logs:
kubectl logs job/<cluster-name>-acl-bootstrap \ --namespace <ns> \ --tail=100helm installandhelm upgradewait for the Job but do not stream its logs. Follow them live from a second terminal withkubectl logs -f. The completed Job and its pod stay until the next run. -
Check the logged bootstrap address and superuser auth type match your broker.
-
Confirm
aclBootstrap.principalis the exact Distinguished Name (DN) (for mutual TLS (mTLS):User:CN=…,O=…,C=…) or username (for SASL:User:<username>). -
The Job is idempotent. After fixing the values, re-run:
helm upgrade <CLUSTER_NAME> \ oci://registry.axual.io/axual-charts/kafka-connect \ --version <CHART_VERSION> \ --namespace <NAMESPACE> \ -f connect-values.yaml
Self-Service cluster registration fails
Cause: One of three Platform Manager validation checks failed:
the REST API is unreachable, Platform Manager cannot write to the Connector Vault (its AppRole policy does not cover the cluster’s paths, or on-premises writer credentials are wrong),
or kafka-topic-name-transforms is not listed in the worker’s plugins.
The Vault check runs as the kafka-connects:validate-vault pre-flight check and, when it fails, returns an error of the form:
The supplied Vault AppRole cannot write to path '…'. Grant 'create' or 'update' capability on this path (capabilities: [deny]).
Three mistakes produce this error. All of them leave the AppRole without access to connector-auth/data/<connector-vault-path>/, which is where the check looks:
-
The policy grants the bare path with no trailing
/*. Vault matches these patterns literally, so such a grant covers nothing. -
The policy grants a path further down, naming the tenant, instance or cluster. The check looks just below
<connector-vault-path>and never reaches those. Grantconnector-auth/data/<connector-vault-path>/*instead. -
The registered
connectorVaultPathalready includes theconnector-authmount, or the tenant, instance and cluster name. Platform Manager adds those itself, so the path ends up doubled and no policy matches it. See Understanding the Vault KV v2 path structure.
Fix:
-
Confirm the worker REST API is reachable from the Platform Manager pod:
kubectl exec -n <pm-namespace> <pm-pod> -- \ curl -s -u "<connect-user>:<connect-password>" \ "<connect-url>/connector-plugins?connectorsOnly=false" | jq '.[].class'Use the
connectUrlregistered in Self-Service as<connect-url>. WithrestApi.networkPolicy.enabled, it must be the route address: the policy blocks Platform Manager from the Service on port 8083, so a direct call there times out by design.Omit -uonly for a cluster still runningrestApi.route.authMethod: none. A401here on a cluster meant to run Basic Auth means the credentials Platform Manager holds do not matchrestApi.basicAuth, not that the API is unreachable. Platform Manager reads the password from the Governance Vault keyconnect-api-password, so it has to match the htpasswd entry the route checks. -
Confirm
kafka-topic-name-transformsappears in the output. If missing, add it topluginsand re-runhelm upgrade --install. -
Confirm Platform Manager can write to the Connector Vault. On Axual Cloud, check that its AppRole policy grants
connector-auth/data/<connector-vault-path>/, with the trailing/and nothing further down the path (step 4 of the Vault setup how-to). On strict on-premises, re-check the writerrole_idandsecret_idagainst Vault:vault read auth/approle/role/<tenant>-<instance>-<cluster-name>-connector-writer/role-idIf the policy or credentials are wrong, follow step 4 of xref:components/run-time/kafka-connect/how-to-set-up-connector-vault.adoc[How to Set Up the Connector Vault for Kafka Connect] and, on-premises, provide the new credentials to the Tenant Admin.
Connector logs do not appear in Self-Service
Cause: The connector runs, but the Log Provisioner cannot return its logs. Either it is not reachable at the URL registered for the Kubernetes cluster, or it is reachable and finds no pods. It locates the pods by the axual.io/tenant, axual.io/instance and axual.io/connect-cluster labels, so it finds none when one of those values differs from what Self-Service holds for the cluster.
The chart renders the labels from the tenant, instance and clusterName values, and the log viewer lowercases the Self-Service names before matching, so the values must be lowercase.
Fix:
-
Confirm the Log Provisioner pod is running:
kubectl -n <ns> get pods -l app.kubernetes.io/name=runtime-provisioner -
Read the labels the chart rendered onto the worker pods:
kubectl -n <ns> get pods -l axual.io/connect-cluster=<cluster-name> \ -o jsonpath='{range .items[*]}{.metadata.labels.axual\.io/tenant}{" "}{.metadata.labels.axual\.io/instance}{" "}{.metadata.labels.axual\.io/connect-cluster}{"\n"}{end}'An empty result means the labels do not carry the values you expect, which is the same failure the Log Provisioner hits.
-
Compare the three values against the lowercased tenant short name, instance short name and cluster name in Self-Service. A cluster name longer than 63 characters cannot be a label value, so shorten it in Self-Service first. A mismatch in any one of them hides every log for the cluster. The labels come from the chart values, so correct them there and run
helm upgrade --installrather than editing the pods. -
Confirm the Log Provisioner URL registered against the Kubernetes cluster in Self-Service is the one the pod serves. See How to Enable Kafka Connect Log Reading.
For what the Log Provisioner does and why it matches on labels, see How the Log Provisioner finds connector logs.