Acting on Alerts

This reference lists the alerts the Axual charts provide, what each one means, and what to do when it fires, for Apache Kafka and for the Strimzi operator.

Type

Reference

Goal

Look up a firing alert and find the response for it.

Audience

Platform Operators on standby, and whoever writes the standby runbook.

When to use

While an alert is firing, and when setting up the standby rota.

Contents

The sections below cover each area in this reference:

Process

With Monitoring & Metrics in place, PrometheusRules define the alerts that trigger AlertManager, which calls out the standby engineers who resolve the issue.

A common sequence of events runs as follows:

  • A production cluster PrometheusRule alert triggers.

  • AlertManager evaluates the active alerts and routes the notifications its configuration defines.

  • The standby team receives the notification, by telephone or otherwise.

  • The responsible standby engineer investigates the issue, reports it and resolves the underlying problem.

  • Additional colleagues join for incident management, when the incident needs them.

  • Prometheus resolves the alerts once the condition clears, and AlertManager sends the green alert notifications where it is configured to.

The sections below help the responsible standby engineer investigate and resolve the problem.

Alert details

Each alert below gives its severity, the threshold that triggers it, what it means and what to do about it. The Axual Streaming chart ships the PrometheusRule that produces them, and Kafka Chart Values Reference covers the values that switch the rule on.

Example Strimzi PrometheusRules

The alerts derive from the Strimzi PrometheusRules example. Open the block below for the full rule definitions.

Click to open strimzi_prometheus_rules.yaml
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
  labels:
    role: alert-rules
    app: strimzi
  name: prometheus-k8s-rules
spec:
  groups:
    - name: kafka
      rules:
        - alert: KafkaRunningOutOfSpace
          expr: kubelet_volume_stats_available_bytes{persistentvolumeclaim=~"data(-[0-9]+)?-(.+)-kafka-[0-9]+"} * 100 / kubelet_volume_stats_capacity_bytes{persistentvolumeclaim=~"data(-[0-9]+)?-(.+)-kafka-[0-9]+"} < 15
          for: 10s
          labels:
            severity: warning
          annotations:
            summary: 'Kafka is running out of free disk space'
            description: 'There are only {{ $value }} percent available at {{ $labels.persistentvolumeclaim }} PVC'
        - alert: UnderReplicatedPartitions
          expr: kafka_server_replicamanager_underreplicatedpartitions > 0
          for: 10s
          labels:
            severity: warning
          annotations:
            summary: 'Kafka under replicated partitions'
            description: 'There are {{ $value }} under replicated partitions on {{ $labels.kubernetes_pod_name }}'
        - alert: AbnormalControllerState
          expr: sum(kafka_controller_kafkacontroller_activecontrollercount) by (strimzi_io_name) != 1
          for: 10s
          labels:
            severity: warning
          annotations:
            summary: 'Kafka abnormal controller state'
            description: 'There are {{ $value }} active controllers in the cluster'
        - alert: OfflinePartitions
          expr: sum(kafka_controller_kafkacontroller_offlinepartitionscount) > 0
          for: 10s
          labels:
            severity: warning
          annotations:
            summary: 'Kafka offline partitions'
            description: 'One or more partitions have no leader'
        - alert: UnderMinIsrPartitionCount
          expr: kafka_server_replicamanager_underminisrpartitioncount > 0
          for: 10s
          labels:
            severity: warning
          annotations:
            summary: 'Kafka under min ISR partitions'
            description: 'There are {{ $value }} partitions under the min ISR on {{ $labels.kubernetes_pod_name }}'
        - alert: OfflineLogDirectoryCount
          expr: kafka_log_logmanager_offlinelogdirectorycount > 0
          for: 10s
          labels:
            severity: warning
          annotations:
            summary: 'Kafka offline log directories'
            description: 'There are {{ $value }} offline log directories on {{ $labels.kubernetes_pod_name }}'
        - alert: ScrapeProblem
          expr: up{kubernetes_namespace!~"openshift-.+",kubernetes_pod_name=~".+-kafka-[0-9]+"} == 0
          for: 3m
          labels:
            severity: major
          annotations:
            summary: 'Prometheus unable to scrape metrics from {{ $labels.kubernetes_pod_name }}/{{ $labels.instance }}'
            description: 'Prometheus was unable to scrape metrics from {{ $labels.kubernetes_pod_name }}/{{ $labels.instance }} for more than 3 minutes'
        - alert: ClusterOperatorContainerDown
          expr: count((container_last_seen{container="strimzi-cluster-operator"} > (time() - 90))) < 1 or absent(container_last_seen{container="strimzi-cluster-operator"})
          for: 1m
          labels:
            severity: major
          annotations:
            summary: 'Cluster Operator down'
            description: 'The Cluster Operator has been down for longer than 90 seconds'
        - alert: KafkaBrokerContainersDown
          expr: absent(container_last_seen{container="kafka",pod=~".+-kafka-[0-9]+"})
          for: 3m
          labels:
            severity: major
          annotations:
            summary: 'All `kafka` containers down or in CrashLookBackOff status'
            description: 'All `kafka` containers have been down or in CrashLookBackOff status for 3 minutes'
        - alert: KafkaContainerRestartedInTheLast5Minutes
          expr: count(count_over_time(container_last_seen{container="kafka"}[5m])) > 2 * count(container_last_seen{container="kafka",pod=~".+-kafka-[0-9]+"})
          for: 5m
          labels:
            severity: warning
          annotations:
            summary: 'One or more Kafka containers restarted too often'
            description: 'One or more Kafka containers were restarted too often within the last 5 minutes'
    - name: entityOperator
      rules:
        - alert: TopicOperatorContainerDown
          expr: absent(container_last_seen{container="topic-operator",pod=~".+-entity-operator-.+"})
          for: 3m
          labels:
            severity: major
          annotations:
            summary: 'Container topic-operator in Entity Operator pod down or in CrashLookBackOff status'
            description: 'Container topic-operator in Entity Operator pod has been or in CrashLookBackOff status for 3 minutes'
        - alert: UserOperatorContainerDown
          expr: absent(container_last_seen{container="user-operator",pod=~".+-entity-operator-.+"})
          for: 3m
          labels:
            severity: major
          annotations:
            summary: 'Container user-operator in Entity Operator pod down or in CrashLookBackOff status'
            description: 'Container user-operator in Entity Operator pod have been down or in CrashLookBackOff status for 3 minutes'
    - name: connect
      rules:
        - alert: ConnectContainersDown
          expr: absent(container_last_seen{container=~".+-connect",pod=~".+-connect-.+"})
          for: 3m
          labels:
            severity: major
          annotations:
            summary: 'All Kafka Connect containers down or in CrashLookBackOff status'
            description: 'All Kafka Connect containers have been down or in CrashLookBackOff status for 3 minutes'
    - name: bridge
      rules:
        - alert: BridgeContainersDown
          expr: absent(container_last_seen{container=~".+-bridge",pod=~".+-bridge-.+"})
          for: 3m
          labels:
            severity: major
          annotations:
            summary: 'All Kafka Bridge containers down or in CrashLookBackOff status'
            description: 'All Kafka Bridge containers have been down or in CrashLookBackOff status for 3 minutes'
        - alert: AvgProducerLatency
          expr: strimzi_bridge_kafka_producer_request_latency_avg > 10
          for: 10s
          labels:
            severity: warning
          annotations:
            summary: 'Kafka Bridge average consumer fetch latency'
            description: 'The average fetch latency is {{ $value }} on {{ $labels.clientId }}'
        - alert: AvgConsumerFetchLatency
          expr: strimzi_bridge_kafka_consumer_fetch_latency_avg > 500
          for: 10s
          labels:
            severity: warning
          annotations:
            summary: 'Kafka Bridge consumer average fetch latency'
            description: 'The average consumer commit latency is {{ $value }} on {{ $labels.clientId }}'
        - alert: AvgConsumerCommitLatency
          expr: strimzi_bridge_kafka_consumer_commit_latency_avg > 200
          for: 10s
          labels:
            severity: warning
          annotations:
            summary: 'Kafka Bridge consumer average commit latency'
            description: 'The average consumer commit latency is {{ $value }} on {{ $labels.clientId }}'
        - alert: Http4xxErrorRate
          expr: strimzi_bridge_http_server_requestCount_total{code=~"^4..$", container=~"^.+-bridge", path !="/favicon.ico"} > 10
          for: 1m
          labels:
            severity: warning
          annotations:
            summary: 'Kafka Bridge returns code 4xx too often'
            description: 'Kafka Bridge returns code 4xx too much ({{ $value }}) for the path {{ $labels.path }}'
        - alert: Http5xxErrorRate
          expr: strimzi_bridge_http_server_requestCount_total{code=~"^5..$", container=~"^.+-bridge"} > 10
          for: 1m
          labels:
            severity: warning
          annotations:
            summary: 'Kafka Bridge returns code 5xx too often'
            description: 'Kafka Bridge returns code 5xx too much ({{ $value }}) for the path {{ $labels.path }}'
    - name: mirrorMaker
      rules:
        - alert: MirrorMakerContainerDown
          expr: absent(container_last_seen{container=~".+-mirror-maker",pod=~".+-mirror-maker-.+"})
          for: 3m
          labels:
            severity: major
          annotations:
            summary: 'All Kafka Mirror Maker containers down or in CrashLookBackOff status'
            description: 'All Kafka Mirror Maker containers have been down or in CrashLookBackOff status for 3 minutes'
    - name: kafkaExporter
      rules:
        - alert: UnderReplicatedPartition
          expr: kafka_topic_partition_under_replicated_partition > 0
          for: 10s
          labels:
            severity: warning
          annotations:
            summary: 'Topic has under-replicated partitions'
            description: 'Topic  {{ $labels.topic }} has {{ $value }} under-replicated partition {{ $labels.partition }}'
        - alert: TooLargeConsumerGroupLag
          expr: kafka_consumergroup_lag > 1000
          for: 10s
          labels:
            severity: warning
          annotations:
            summary: 'Consumer group lag is too big'
            description: 'Consumer group {{ $labels.consumergroup}} lag is too big ({{ $value }}) on topic {{ $labels.topic }}/partition {{ $labels.partition }}'
        - alert: NoMessageForTooLong
          expr: changes(kafka_topic_partition_current_offset[10m]) == 0
          for: 10s
          labels:
            severity: warning
          annotations:
            summary: 'No message for 10 minutes'
            description: 'There is no messages in topic {{ $labels.topic}}/partition {{ $labels.partition }} for 10 minutes'
    - name: certificates
      interval: 1m0s
      rules:
        - alert: CertificateExpiration
          expr: |
            strimzi_certificate_expiration_timestamp_ms/1000 - time() < 30 * 24 * 60 * 60
          for: 5m
          labels:
            severity: warning
          annotations:
            summary: 'Certificate will expire in less than 30 days'
            description: 'Certificate of type {{ $labels.type }} in cluster {{ $labels.cluster }} in namespace {{ $labels.resource_namespace }} will expire in less than 30 days'

Kafka

These alerts come from the Apache Kafka brokers and their storage.

KafkaRunningOutOfSpace

Kafka is running out of free disk space.

  • Severity: major

  • Default Threshold: 85% full

  • Description: 'There are only {{ $value }} percent available at {{ $labels.persistentvolumeclaim }} PVC'

  • Comment: This is a potential cluster breaking issue. Brokers keep writing data until the disks are completely full, after which they cannot function properly. They then have trouble starting and cannot run the housekeeping that removes data past its retention time.
    Once the disks are full, only resizing them dynamically saves the brokers.

  • Action:

    • Identify which topics are the largest. Where these are not compaction topics, consider reducing the retention time.

    • Then identify which producers add the most data. Reach the responsible teams and decide whether to increase the disks, lower the retention or change the producer behaviour.

    • In an emergency, where a producer floods the cluster, revoke that producer’s Access Control List (ACL) entries on the topic. This stops new data arriving without deleting the data already there.

UnderReplicatedPartitions

Kafka under replicated partitions.

  • Severity: warning

  • Default Threshold: More than 0

  • Description: 'There are {{ $value }} under replicated partitions on {{ $labels.kubernetes_pod_name }}'

  • Comment:

    • Some partitions are online but have fewer In Sync Replicas (ISRs) than expected.

    • This is common and happens when a transient network issue makes a broker unavailable.

    • It also occurs when a broker is offline. The Kafka cluster is still fully functional then, but in a strained state.

    • It can be an early warning for a larger problem when it does not resolve automatically.

  • Action: Wait for the issue to resolve itself. Investigate further if it lasts longer than 5 minutes.

AbnormalControllerState

Kafka abnormal controller state.

  • Severity: warning

  • Default Threshold: Other than 1

  • Description: 'There are {{ $value }} active controllers in the cluster'

  • Comment: A quick rolling restart of the brokers combined with slow metrics scraping temporarily registers two controllers, which reports as multiple active controllers.

  • Action: Wait for the issue to resolve itself. Investigate further if it lasts longer than 5 minutes.

OfflinePartitions

Kafka offline partitions.

  • Severity: warning

  • Default Threshold: More than 0

  • Description: 'One or more partitions have no leader'

  • Comment: This should never occur when partitions have enough replicas and no more than one broker is unavailable.

  • Action:

    • If multiple brokers are offline, this is a major incident and probably a loss of service. Quickly escalate and investigate why brokers are not available.

    • If one broker is offline and this alert triggers, identify which topics are affected and ensure the correct number of replicas.

UnderMinIsrPartitionCount

Kafka under min ISR partitions.

  • Severity: warning

  • Default Threshold: More than 0

  • Description: 'There are {{ $value }} partitions under the min ISR on {{ $labels.kubernetes_pod_name }}'

  • Comment: When below minimum In Sync Replicas, most producers stop writing records into Kafka to ensure no data loss on broker failure.

  • Action: This can only occur when at least two brokers are down, so this is part of a cluster outage. Quickly escalate and investigate why multiple brokers are unavailable.

OfflineLogDirectoryCount

Kafka offline log directories.

  • Severity: warning

  • Default Threshold: More than 0

  • Description: 'There are {{ $value }} offline log directories on {{ $labels.kubernetes_pod_name }}'

  • Comment: This signals an issue with the underlying storage. Some log directories are missing or unreachable.

  • Action: Involve the infrastructure team and investigate the root cause of the failure.

ScrapeProblem

Prometheus unable to scrape metrics from {{ $labels.kubernetes_pod_name }}/{{ $labels.instance }}.

  • Severity: major

  • Default Threshold: More than 0

  • Description: 'Prometheus was unable to scrape metrics from {{ $labels.kubernetes_pod_name }}/{{ $labels.instance }} for more than 3 minutes'

  • Comment: None of the alerts work while metrics are not scraped, so this is a critical issue.

  • Action:

    • Check the cluster health and the broker Pod status straight away. Without Prometheus the cluster is unmonitored, so watch it directly until scraping returns.

    • Investigate where the problem originates. It can sit in the upstream Prometheus when the platform runs in a federated setup.

KafkaBrokerContainersDown

All kafka containers down or in CrashLookBackOff status.

  • Severity: major

  • Default Threshold: All absent

  • Description: 'All kafka containers have been down or in CrashLookBackOff status for 3 minutes'

  • Comment: The whole Kafka cluster is broken. This is a critical incident.

  • Action:

    • Escalate, the Kafka service is unavailable.

    • Look into Kubernetes Events for scheduling issues or any missing volumeMounts like Secrets.

    • UnderReplicatedPartitions and OfflinePartitions fire before this one, so it should never be the first alert you see.

KafkaContainerRestartedInTheLast5Minutes

One or more Kafka containers restarted too often.

  • Severity: warning

  • Default Threshold: > 2 restarts

  • Description: 'One or more Kafka containers were restarted too often within the last 5 minutes'

  • Comment: This also occurs during broker updates, where it is not a problem.

  • Action: If this occurs outside of maintenance, check for memory usage (OOMKilled Pods) or crash looping Pods. Check Kubernetes Events.

Strimzi

These alerts come from the Strimzi operator rather than from Kafka itself.

ClusterOperatorContainerDown

Cluster Operator down.

  • Severity: major

  • Default Threshold: Less than 1

  • Description: 'The Cluster Operator has been down for longer than 90 seconds'

  • Comment: On a stable cluster the impact is low, because Kafka keeps functioning without an active Strimzi Cluster Operator.
    It becomes a problem as soon as the cluster needs a change applied.

  • Action: The Strimzi Operator Pod should be in a Ready state. Investigate whether the Pod is crash looping or whether Kubernetes cannot schedule it.