Deployment Troubleshooting

This guide covers the failures specific to installing the platform: an ingress that cannot be reached from outside or from inside the cluster, and a Self-Service page that loads blank.

Type

How-to guide

Goal

Get past an installation that completed but is not reachable.

Audience

Platform Operator with read access to the namespaces the platform runs in.

When to use

Use this guide when the charts installed cleanly but the platform cannot be used.

Each section below names an observable symptom, what causes it, and how to fix it.

Replace every <VALUE> placeholder with your own value before running a command.

Contents

Prerequisites

Confirm the following before you begin.

Access and permissions required

You need the following access and permissions:

  • Read access to the namespaces the platform runs in, and permission to run kubectl exec, because several of the checks below run from inside a pod.

  • A shell on the machine you reach the platform from, to read its /etc/hosts and run curl against the ingress.

Tools and versions required

You need the following tools:

  • kubectl >= 1.28, configured for the target cluster.

  • curl, on the machine you test from. The checks that run inside a pod assume the pod’s image carries it too.

Resources that must exist before starting

The following must already exist:

  • An installation whose charts finished and whose pods are running. A pod that never starts is a different failure: see How to Inspect Kubernetes Workloads.

  • The host name the platform is published on, and the ingress that serves it. Without both there is nothing to resolve and nothing to reach.

Ingress cannot be reached

Cause: The host name resolves to the wrong address, the ingress controller’s Service has no external address, or the Service is of a type the cluster cannot give one to. Which of the three it is shows up as the layer the request stops at.

Fix: Work outwards from the application. The first check that fails names the layer to repair.

  1. Confirm the application answers inside its own pod:

    kubectl exec -n <NAMESPACE> <API_GATEWAY_POD> -- curl -vvvk http://localhost

    A response means the application is healthy and the fault sits in front of it. No response makes the pod itself the problem, which How to Inspect Kubernetes Workloads covers.

  2. Confirm the Service resolves and forwards, from a second pod in the cluster:

    kubectl exec -n <NAMESPACE> <ANOTHER_POD> -- curl -vvvk http://<API_GATEWAY_SERVICE_NAME>

    A failure here after a successful first check puts the fault on the Service rather than on the application.

  3. Confirm the ingress publishes the host name, from your own machine:

    curl -vvvk https://axual.<DOMAIN>

    The -v output names the address curl resolved the host name to. When that is not the address the ingress is published on, correct the DNS record, or the entry in this machine’s /etc/hosts. Where one name has two entries, the first match wins, which may not be the one you added.

  4. Confirm the ingress controller’s Service holds an external address:

    kubectl get service -n ingress-nginx

    EXTERNAL-IP stays <pending> while nothing can assign one. On a local cluster whose tool needs a tunnel or proxy process to expose LoadBalancer Services, that is expected until that process runs, so start it and read the Service again. On a local macOS installation, check the loopback alias as well: see How to Set Up a Local Loopback Alias.

A Service of the wrong type never receives an external address. LoadBalancer is the type the chart in Install an Ingress Controller creates, and the type a local tunnel process serves. Where a cloud load balancer needs annotations on the Service, confirm those are present too.

Ingress cannot be reached from a pod

Cause: The ingress host name does not resolve inside the pod. A pod resolves a name in a fixed order: its own /etc/hosts first, then the nameservers listed in its /etc/resolv.conf, which point at the cluster DNS service (CoreDNS), and last whatever CoreDNS forwards to upstream. A name that resolves on your workstation and fails in the pod is answered by something only the workstation holds. Most often that is an entry in the workstation’s own /etc/hosts, which no pod ever reads.

Fix: Find where the chain stops, then give the cluster the same answer the workstation has.

  1. Read the resolver configuration the pod starts from:

    kubectl exec -n <NAMESPACE> <POD_NAME> -- cat /etc/resolv.conf
    kubectl exec -n <NAMESPACE> <POD_NAME> -- cat /etc/hosts

    Between them these say which nameserver the pod asks and which names it answers itself, which is where a name that never leaves the pod is decided.

  2. Check the workstation’s /etc/hosts for an entry for the same name. An entry there, and nowhere else, is exactly the case that resolves locally and fails in every pod.

  3. Add the name to the cluster’s DNS configuration rather than to individual pods. A pod’s /etc/hosts is rebuilt every time the pod restarts, so an entry written into a running pod does not survive.

Blank page in the UI

Cause: The browser is serving a Self-Service session it cached earlier, and the session or authentication details it holds are stale.

Fix: Rule the browser cache out first, then refresh the session.

  1. Open Self-Service in a private or incognito window. A page that renders there confirms the cache is the cause, and clearing the site’s stored data in your normal window fixes it there too.

  2. If the page is still blank, open the identity provider endpoint directly, then reload Self-Service. That refreshes the session and authentication details the UI reads.

    On the local quick-setup installation the identity provider is reached at https://apicurio-keycloak.<DOMAIN>;, one of the host names that installation’s /etc/hosts entry defines. See Deploying Axual Kafka and Axual Governance on Kubernetes.