Deployment Troubleshooting
This guide covers the failures specific to installing the platform: an ingress that cannot be reached from outside or from inside the cluster, and a Self-Service page that loads blank.
Type |
How-to guide |
Goal |
Get past an installation that completed but is not reachable. |
Audience |
Platform Operator with read access to the namespaces the platform runs in. |
When to use |
Use this guide when the charts installed cleanly but the platform cannot be used. |
Each section below names an observable symptom, what causes it, and how to fix it.
Replace every <VALUE> placeholder with your own value before running a command.
|
Prerequisites
Confirm the following before you begin.
Access and permissions required
You need the following access and permissions:
-
Read access to the namespaces the platform runs in, and permission to run
kubectl exec, because several of the checks below run from inside a pod. -
A shell on the machine you reach the platform from, to read its
/etc/hostsand runcurlagainst the ingress.
Tools and versions required
You need the following tools:
-
kubectl>= 1.28, configured for the target cluster. -
curl, on the machine you test from. The checks that run inside a pod assume the pod’s image carries it too.
Resources that must exist before starting
The following must already exist:
-
An installation whose charts finished and whose pods are running. A pod that never starts is a different failure: see How to Inspect Kubernetes Workloads.
-
The host name the platform is published on, and the ingress that serves it. Without both there is nothing to resolve and nothing to reach.
Ingress cannot be reached
Cause: The host name resolves to the wrong address, the ingress controller’s Service has no external address, or the Service is of a type the cluster cannot give one to. Which of the three it is shows up as the layer the request stops at.
Fix: Work outwards from the application. The first check that fails names the layer to repair.
-
Confirm the application answers inside its own pod:
kubectl exec -n <NAMESPACE> <API_GATEWAY_POD> -- curl -vvvk http://localhostA response means the application is healthy and the fault sits in front of it. No response makes the pod itself the problem, which How to Inspect Kubernetes Workloads covers.
-
Confirm the Service resolves and forwards, from a second pod in the cluster:
kubectl exec -n <NAMESPACE> <ANOTHER_POD> -- curl -vvvk http://<API_GATEWAY_SERVICE_NAME>A failure here after a successful first check puts the fault on the Service rather than on the application.
-
Confirm the ingress publishes the host name, from your own machine:
curl -vvvk https://axual.<DOMAIN>The
-voutput names the addresscurlresolved the host name to. When that is not the address the ingress is published on, correct the DNS record, or the entry in this machine’s/etc/hosts. Where one name has two entries, the first match wins, which may not be the one you added. -
Confirm the ingress controller’s Service holds an external address:
kubectl get service -n ingress-nginxEXTERNAL-IPstays<pending>while nothing can assign one. On a local cluster whose tool needs a tunnel or proxy process to exposeLoadBalancerServices, that is expected until that process runs, so start it and read the Service again. On a local macOS installation, check the loopback alias as well: see How to Set Up a Local Loopback Alias.
A Service of the wrong type never receives an external address. LoadBalancer is the type the chart in Install an Ingress Controller creates, and the type a local tunnel process serves. Where a cloud load balancer needs annotations on the Service, confirm those are present too.
|
Ingress cannot be reached from a pod
Cause: The ingress host name does not resolve inside the pod. A pod resolves a name in a fixed order: its own /etc/hosts first, then the nameservers listed in its /etc/resolv.conf, which point at the cluster DNS service (CoreDNS), and last whatever CoreDNS forwards to upstream. A name that resolves on your workstation and fails in the pod is answered by something only the workstation holds. Most often that is an entry in the workstation’s own /etc/hosts, which no pod ever reads.
Fix: Find where the chain stops, then give the cluster the same answer the workstation has.
-
Read the resolver configuration the pod starts from:
kubectl exec -n <NAMESPACE> <POD_NAME> -- cat /etc/resolv.conf kubectl exec -n <NAMESPACE> <POD_NAME> -- cat /etc/hostsBetween them these say which nameserver the pod asks and which names it answers itself, which is where a name that never leaves the pod is decided.
-
Check the workstation’s
/etc/hostsfor an entry for the same name. An entry there, and nowhere else, is exactly the case that resolves locally and fails in every pod. -
Add the name to the cluster’s DNS configuration rather than to individual pods. A pod’s
/etc/hostsis rebuilt every time the pod restarts, so an entry written into a running pod does not survive.
Blank page in the UI
Cause: The browser is serving a Self-Service session it cached earlier, and the session or authentication details it holds are stale.
Fix: Rule the browser cache out first, then refresh the session.
-
Open Self-Service in a private or incognito window. A page that renders there confirms the cache is the cause, and clearing the site’s stored data in your normal window fixes it there too.
-
If the page is still blank, open the identity provider endpoint directly, then reload Self-Service. That refreshes the session and authentication details the UI reads.
On the local quick-setup installation the identity provider is reached at https://apicurio-keycloak.<DOMAIN>, one of the host names that installation’s/etc/hostsentry defines. See Deploying Axual Kafka and Axual Governance on Kubernetes.