How to Check Distribution Health

This guide shows you how to read the consumer group lag of a running distribution deployment from inside a broker pod, and how to tell from the target cluster’s offsets that the Offset Committer is applying what the Offset Distributor sends it.

Type

How-to guide

Goal

Confirm that the connectors of a running distribution deployment are keeping up with the load.

Audience

Platform Operator with a shell into a broker pod on both clusters.

When to use

Use this guide as a routine check on a deployment that is already distributing, and when a consumer reports that it reprocesses records after moving to another cluster.

Connector state alone can look healthy while nothing progresses, because a connector reports itself running whether or not its consumer group gains ground. Lag is the figure that separates the two. In a bilateral setup, up to four connector tasks run per cluster: message distributor, offset distributor, offset committer, and optionally the schema distributor. For why offsets travel as timestamps rather than as numbers, see Offset Distribution.

Contents

Prerequisites

Confirm the following before you begin.

Access and permissions required

You need the following access and permissions:

  • A shell into a broker pod on the source cluster and on the target cluster, or another way to run kafka-consumer-groups.sh against both.

  • A Kafka principal allowed to list and describe consumer groups on both clusters.

Tools and versions required

You need the following tools:

  • kubectl >= 1.28, to open the shell into the broker pod.

  • kafka-consumer-groups.sh, which ships in the broker image at /opt/kafka/bin/.

Resources that must exist before starting

The following must already exist:

  • A distribution deployment whose connectors are running on both clusters. See How to Deploy the Axual Distributor.

  • The authentication material the brokers expect: the truststore and keystore on a Transport Layer Security (TLS) cluster, or the Simple Authentication and Security Layer (SASL) credentials on a SASL cluster. The commands below read it from a client properties file you create in the pod.

Create the client properties file

Every command in this guide runs inside a broker pod and authenticates to the broker with a client properties file. Create that file first, in the same pod, using the same authentication the Axual Distributor uses.

Replace every <VALUE> placeholder with your own value before running a command.
cat > /tmp/client.properties <<'EOF'
security.protocol=SSL
ssl.truststore.location=<TRUSTSTORE_PATH>
ssl.truststore.password=<TRUSTSTORE_PASSWORD>
ssl.keystore.location=<KEYSTORE_PATH>
ssl.keystore.password=<KEYSTORE_PASSWORD>
EOF

On a SASL cluster, use the sasl. properties the brokers expect instead of the ssl.keystore. pair.

Check that both connectors are keeping up

Offsets only arrive when the Offset Distributor is running on the source cluster and the Offset Committer is running on the target, and neither side reports the other as missing. Read the source side from the Offset Distributor’s lag, then read the target side by comparing consumer group offsets between the two clusters.

Read a distributor’s consumer group lag

List the groups on the source cluster first, because the group name carries the cluster names and the level it distributes at:

/opt/kafka/bin/kafka-consumer-groups.sh --command-config /tmp/client.properties \
  --bootstrap-server <KAFKA_HOST> --list

The Offset Distributor’s group is named _<CLUSTER_NAME>-offset-distributor-level-<LEVEL>-to-<TARGET_CLUSTER_NAME>. All three variable parts come from the chart:

  • <CLUSTER_NAME> is distribution.sourceCluster.name.

  • <LEVEL> is the level’s key under distribution.levels, a number from 1 to 32.

  • <TARGET_CLUSTER_NAME> is that level’s targetClusterName.

Take the example in levels: message and offset distributors, where cluster01 distributes to cluster02 at level 1. Its group is _cluster01-offset-distributor-level-1-to-cluster02.

Describe the group the listing returned:

/opt/kafka/bin/kafka-consumer-groups.sh --command-config /tmp/client.properties \
  --bootstrap-server <KAFKA_HOST> --describe \
  --group _<CLUSTER_NAME>-offset-distributor-level-<LEVEL>-to-<TARGET_CLUSTER_NAME>

Expected output gives current-offset and log-end-offset across the partitions of __consumer_offsets. Watch LAG: a figure that keeps falling means the Offset Distributor is working, and one that only grows means it has stalled.

The Message Distributor has its own consumer group per level, which the same --list returns and the same --describe reads, so message distribution lag is checked the same way.

Compare consumer group offsets between the clusters

The offsets themselves confirm the receiving side. Run the following against the source cluster, then against the target cluster, and compare the two:

/opt/kafka/bin/kafka-consumer-groups.sh --command-config /tmp/client.properties \
  --bootstrap-server <KAFKA_HOST> --describe --offsets --all-groups

Working offset distribution shows the target cluster’s offsets updating quickly for consumer groups with no active consumer there. Nothing but the Offset Committer moves those offsets.

There is normally under a minute of delay between a commit on the source cluster and the matching commit on the target, so a small amount of reprocessing after a migration is expected rather than a fault.