Troubleshooting
This guide introduces where to look when an installation or a running platform misbehaves, from a cluster-wide symptom down to a single component. It is not a stage of The installation order but a destination from any stage in it, and from a platform that is already live.
Type |
Reference |
Goal |
Find the page that covers the symptom you are seeing. |
Audience |
Platform Operators with |
When to use |
When a stage of the installation fails, or when a production alert has fired. |
The pages in this section move from general diagnosis to specific components:
-
How to Inspect Kubernetes Workloads covers the first moves on any problem: pod status, the events behind a pod that will not start, and reading its logs.
-
How to Restart a Service covers restarting a component that is running but unresponsive, and why an Apache Kafka broker is rolled through Strimzi instead.
-
How to Debug Authorisation Errors in Platform Manager covers making Platform Manager log the access decision behind an unexpected refusal.
-
Deployment Troubleshooting covers the symptoms specific to installing, such as an ingress that cannot be reached and a blank Self-Service page.
-
Acting on Alerts is the standby runbook: each Kafka and Strimzi alert, what it means, and what to do about it.
-
Troubleshoot a Kafka Connect Deployment covers the per-tenant Connect clusters, from a worker that will not start to a cluster registration that Platform Manager rejects.
-
Troubleshoot an Axual Connect Deployment covers the deprecated shared Connect cluster.
Reading logs is the first move in most of these, and How to Configure Logging is the one page that covers setting the levels.