How to Improve Rest Proxy Availability

This guide shows you how to put several Rest Proxy instances behind a load balancer without breaking client authentication: choosing a load balancer that supports the right session behaviour, deciding which names go on the server certificate, and configuring the balancer’s forwarding and health checks.

Type

How-to guide

Goal

Run several Rest Proxy instances behind one endpoint, so a single instance failing does not take the REST interface down.

Audience

Platform Operator working with whoever administers the load balancer and issues server certificates.

When to use

Use this guide when one Rest Proxy instance is not enough, after the streaming layer is installed.

Rest Proxy is stateful and authenticates clients with TLS, and those two facts decide the whole design. A load balancer that terminates TLS destroys the client identity the proxy authorises on, and one that spreads a client’s requests across instances breaks the session. Both are configuration choices, and the common defaults get both of them wrong.

Prerequisites

Confirm the following before you begin.

Access and permissions required

You need the following access and permissions:

  • Permission to configure the load balancer, or an administrator who can.

  • The ability to request a server certificate from the certificate authority the platform’s clients trust.

Resources that must exist before starting

The following must already exist:

  • Several Rest Proxy instances, reachable on the network. See Install Axual Streaming.

  • A load balancer that can forward TCP and support sticky sessions. Confirm both before designing around it, because a balancer that cannot do either forces a different topology.

Choose a load balancer and public endpoint

The load balancer must forward TCP rather than terminate it. It must also support sticky sessions, which send every request from one source IP address to the same backend. Rest Proxy holds per-client state, so a request landing on a different instance mid-session fails.

Decide the public host name or IP address clients will use. It goes on the Rest Proxy server certificate, so settle it before requesting the certificate.

Choose the backends and ports

Collect the host name, IP address and port numbers of each Rest Proxy instance. Note both the HTTPS port and the management port: the first is what the balancer forwards to, and the second is what it health-checks.

This list becomes the balancer’s backend configuration. Add the backend host names to the certificate as well if clients will ever reach an instance directly rather than through the balancer.

Request the server certificate

Request a new server certificate for Rest Proxy with the public host name in its Subject Alternative Names, plus any backend host names from the previous step.

Deploy the instances with that certificate before configuring the balancer. A client reaching the public endpoint validates the name it dialled against the certificate the instance presents, so the certificate has to be in place first.

If the platform’s Prometheus is the Axual managed one, restart it with the new configuration so it scrapes every instance rather than the original one. Skip this otherwise.

Configure the load balancer

Replace every <VALUE> placeholder with your own value, then apply the following four settings. The last two are the ones most often missed:

  • Forward TCP requests to the HTTPS port of each backend.

  • Use http://<BACKEND_IP>:<MANAGEMENT_PORT>/actuator/health as the health probe.

  • Disable SSL offloading. Offloading terminates TLS at the balancer, which strips the client certificate Rest Proxy authenticates on.

  • Enable sticky sessions based on client port in the load balancing rules.

Verify the balanced endpoint

Check the backends and the public endpoint separately, because they fail for different reasons. A backend the probe cannot reach is a port or health-check problem, while a reachable backend that refuses an authenticated call is TLS being terminated at the balancer.

Run the health probe against each backend, using the same URL you gave the load balancer:

curl -s http://<BACKEND_IP>:<MANAGEMENT_PORT>/actuator/health

Every instance returns {"status":"UP"}, and every node shows as active in the load balancer’s own status view. A node marked unhealthy while the instance is running usually means the probe points at the HTTPS port instead of the management port.

Then produce a record through the public endpoint with a client certificate, using the request in Example: Producing data with Rest Proxy with its REST_PROXY_URL set to the public endpoint rather than to a single instance. Send the same request several times, keeping one axual-producer-uuid across all of them.

  • Every call returns 200. A TLS or authorisation failure here, on a request that succeeds against an instance directly, means the balancer is terminating TLS and stripping the client certificate, so re-check that SSL offloading is disabled.

  • Every call in the repeated sequence succeeds. One that fails partway through means the balancer sent it to a different instance from the one holding the session, so re-check the sticky session rule.