Acting on Alerts
This reference lists the alerts the Axual charts provide, what each one means, and what to do when it fires, for Apache Kafka and for the Strimzi operator.
Type |
Reference |
Goal |
Look up a firing alert and find the response for it. |
Audience |
Platform Operators on standby, and whoever writes the standby runbook. |
When to use |
While an alert is firing, and when setting up the standby rota. |
Process
With Monitoring & Metrics in place, PrometheusRules define the alerts that trigger AlertManager, which calls out the standby engineers who resolve the issue.
A common sequence of events runs as follows:
-
A production cluster PrometheusRule alert triggers.
-
AlertManager evaluates the active alerts and routes the notifications its configuration defines.
-
The standby team receives the notification, by telephone or otherwise.
-
The responsible standby engineer investigates the issue, reports it and resolves the underlying problem.
-
Additional colleagues join for incident management, when the incident needs them.
-
Prometheus resolves the alerts once the condition clears, and AlertManager sends the green alert notifications where it is configured to.
The sections below help the responsible standby engineer investigate and resolve the problem.
Alert details
Each alert below gives its severity, the threshold that triggers it, what it means and what to do about it. The Axual Streaming chart ships the PrometheusRule that produces them, and Kafka Chart Values Reference covers the values that switch the rule on.
Example Strimzi PrometheusRules
The alerts derive from the Strimzi PrometheusRules example. Open the block below for the full rule definitions.
Click to open strimzi_prometheus_rules.yaml
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
labels:
role: alert-rules
app: strimzi
name: prometheus-k8s-rules
spec:
groups:
- name: kafka
rules:
- alert: KafkaRunningOutOfSpace
expr: kubelet_volume_stats_available_bytes{persistentvolumeclaim=~"data(-[0-9]+)?-(.+)-kafka-[0-9]+"} * 100 / kubelet_volume_stats_capacity_bytes{persistentvolumeclaim=~"data(-[0-9]+)?-(.+)-kafka-[0-9]+"} < 15
for: 10s
labels:
severity: warning
annotations:
summary: 'Kafka is running out of free disk space'
description: 'There are only {{ $value }} percent available at {{ $labels.persistentvolumeclaim }} PVC'
- alert: UnderReplicatedPartitions
expr: kafka_server_replicamanager_underreplicatedpartitions > 0
for: 10s
labels:
severity: warning
annotations:
summary: 'Kafka under replicated partitions'
description: 'There are {{ $value }} under replicated partitions on {{ $labels.kubernetes_pod_name }}'
- alert: AbnormalControllerState
expr: sum(kafka_controller_kafkacontroller_activecontrollercount) by (strimzi_io_name) != 1
for: 10s
labels:
severity: warning
annotations:
summary: 'Kafka abnormal controller state'
description: 'There are {{ $value }} active controllers in the cluster'
- alert: OfflinePartitions
expr: sum(kafka_controller_kafkacontroller_offlinepartitionscount) > 0
for: 10s
labels:
severity: warning
annotations:
summary: 'Kafka offline partitions'
description: 'One or more partitions have no leader'
- alert: UnderMinIsrPartitionCount
expr: kafka_server_replicamanager_underminisrpartitioncount > 0
for: 10s
labels:
severity: warning
annotations:
summary: 'Kafka under min ISR partitions'
description: 'There are {{ $value }} partitions under the min ISR on {{ $labels.kubernetes_pod_name }}'
- alert: OfflineLogDirectoryCount
expr: kafka_log_logmanager_offlinelogdirectorycount > 0
for: 10s
labels:
severity: warning
annotations:
summary: 'Kafka offline log directories'
description: 'There are {{ $value }} offline log directories on {{ $labels.kubernetes_pod_name }}'
- alert: ScrapeProblem
expr: up{kubernetes_namespace!~"openshift-.+",kubernetes_pod_name=~".+-kafka-[0-9]+"} == 0
for: 3m
labels:
severity: major
annotations:
summary: 'Prometheus unable to scrape metrics from {{ $labels.kubernetes_pod_name }}/{{ $labels.instance }}'
description: 'Prometheus was unable to scrape metrics from {{ $labels.kubernetes_pod_name }}/{{ $labels.instance }} for more than 3 minutes'
- alert: ClusterOperatorContainerDown
expr: count((container_last_seen{container="strimzi-cluster-operator"} > (time() - 90))) < 1 or absent(container_last_seen{container="strimzi-cluster-operator"})
for: 1m
labels:
severity: major
annotations:
summary: 'Cluster Operator down'
description: 'The Cluster Operator has been down for longer than 90 seconds'
- alert: KafkaBrokerContainersDown
expr: absent(container_last_seen{container="kafka",pod=~".+-kafka-[0-9]+"})
for: 3m
labels:
severity: major
annotations:
summary: 'All `kafka` containers down or in CrashLookBackOff status'
description: 'All `kafka` containers have been down or in CrashLookBackOff status for 3 minutes'
- alert: KafkaContainerRestartedInTheLast5Minutes
expr: count(count_over_time(container_last_seen{container="kafka"}[5m])) > 2 * count(container_last_seen{container="kafka",pod=~".+-kafka-[0-9]+"})
for: 5m
labels:
severity: warning
annotations:
summary: 'One or more Kafka containers restarted too often'
description: 'One or more Kafka containers were restarted too often within the last 5 minutes'
- name: entityOperator
rules:
- alert: TopicOperatorContainerDown
expr: absent(container_last_seen{container="topic-operator",pod=~".+-entity-operator-.+"})
for: 3m
labels:
severity: major
annotations:
summary: 'Container topic-operator in Entity Operator pod down or in CrashLookBackOff status'
description: 'Container topic-operator in Entity Operator pod has been or in CrashLookBackOff status for 3 minutes'
- alert: UserOperatorContainerDown
expr: absent(container_last_seen{container="user-operator",pod=~".+-entity-operator-.+"})
for: 3m
labels:
severity: major
annotations:
summary: 'Container user-operator in Entity Operator pod down or in CrashLookBackOff status'
description: 'Container user-operator in Entity Operator pod have been down or in CrashLookBackOff status for 3 minutes'
- name: connect
rules:
- alert: ConnectContainersDown
expr: absent(container_last_seen{container=~".+-connect",pod=~".+-connect-.+"})
for: 3m
labels:
severity: major
annotations:
summary: 'All Kafka Connect containers down or in CrashLookBackOff status'
description: 'All Kafka Connect containers have been down or in CrashLookBackOff status for 3 minutes'
- name: bridge
rules:
- alert: BridgeContainersDown
expr: absent(container_last_seen{container=~".+-bridge",pod=~".+-bridge-.+"})
for: 3m
labels:
severity: major
annotations:
summary: 'All Kafka Bridge containers down or in CrashLookBackOff status'
description: 'All Kafka Bridge containers have been down or in CrashLookBackOff status for 3 minutes'
- alert: AvgProducerLatency
expr: strimzi_bridge_kafka_producer_request_latency_avg > 10
for: 10s
labels:
severity: warning
annotations:
summary: 'Kafka Bridge average consumer fetch latency'
description: 'The average fetch latency is {{ $value }} on {{ $labels.clientId }}'
- alert: AvgConsumerFetchLatency
expr: strimzi_bridge_kafka_consumer_fetch_latency_avg > 500
for: 10s
labels:
severity: warning
annotations:
summary: 'Kafka Bridge consumer average fetch latency'
description: 'The average consumer commit latency is {{ $value }} on {{ $labels.clientId }}'
- alert: AvgConsumerCommitLatency
expr: strimzi_bridge_kafka_consumer_commit_latency_avg > 200
for: 10s
labels:
severity: warning
annotations:
summary: 'Kafka Bridge consumer average commit latency'
description: 'The average consumer commit latency is {{ $value }} on {{ $labels.clientId }}'
- alert: Http4xxErrorRate
expr: strimzi_bridge_http_server_requestCount_total{code=~"^4..$", container=~"^.+-bridge", path !="/favicon.ico"} > 10
for: 1m
labels:
severity: warning
annotations:
summary: 'Kafka Bridge returns code 4xx too often'
description: 'Kafka Bridge returns code 4xx too much ({{ $value }}) for the path {{ $labels.path }}'
- alert: Http5xxErrorRate
expr: strimzi_bridge_http_server_requestCount_total{code=~"^5..$", container=~"^.+-bridge"} > 10
for: 1m
labels:
severity: warning
annotations:
summary: 'Kafka Bridge returns code 5xx too often'
description: 'Kafka Bridge returns code 5xx too much ({{ $value }}) for the path {{ $labels.path }}'
- name: mirrorMaker
rules:
- alert: MirrorMakerContainerDown
expr: absent(container_last_seen{container=~".+-mirror-maker",pod=~".+-mirror-maker-.+"})
for: 3m
labels:
severity: major
annotations:
summary: 'All Kafka Mirror Maker containers down or in CrashLookBackOff status'
description: 'All Kafka Mirror Maker containers have been down or in CrashLookBackOff status for 3 minutes'
- name: kafkaExporter
rules:
- alert: UnderReplicatedPartition
expr: kafka_topic_partition_under_replicated_partition > 0
for: 10s
labels:
severity: warning
annotations:
summary: 'Topic has under-replicated partitions'
description: 'Topic {{ $labels.topic }} has {{ $value }} under-replicated partition {{ $labels.partition }}'
- alert: TooLargeConsumerGroupLag
expr: kafka_consumergroup_lag > 1000
for: 10s
labels:
severity: warning
annotations:
summary: 'Consumer group lag is too big'
description: 'Consumer group {{ $labels.consumergroup}} lag is too big ({{ $value }}) on topic {{ $labels.topic }}/partition {{ $labels.partition }}'
- alert: NoMessageForTooLong
expr: changes(kafka_topic_partition_current_offset[10m]) == 0
for: 10s
labels:
severity: warning
annotations:
summary: 'No message for 10 minutes'
description: 'There is no messages in topic {{ $labels.topic}}/partition {{ $labels.partition }} for 10 minutes'
- name: certificates
interval: 1m0s
rules:
- alert: CertificateExpiration
expr: |
strimzi_certificate_expiration_timestamp_ms/1000 - time() < 30 * 24 * 60 * 60
for: 5m
labels:
severity: warning
annotations:
summary: 'Certificate will expire in less than 30 days'
description: 'Certificate of type {{ $labels.type }} in cluster {{ $labels.cluster }} in namespace {{ $labels.resource_namespace }} will expire in less than 30 days'
Kafka
These alerts come from the Apache Kafka brokers and their storage.
KafkaRunningOutOfSpace
Kafka is running out of free disk space.
-
Severity: major
-
Default Threshold: 85% full
-
Description: 'There are only {{ $value }} percent available at {{ $labels.persistentvolumeclaim }} PVC'
-
Comment: This is a potential cluster breaking issue. Brokers keep writing data until the disks are completely full, after which they cannot function properly. They then have trouble starting and cannot run the housekeeping that removes data past its retention time.
Once the disks are full, only resizing them dynamically saves the brokers. -
Action:
-
Identify which topics are the largest. Where these are not compaction topics, consider reducing the retention time.
-
Then identify which producers add the most data. Reach the responsible teams and decide whether to increase the disks, lower the retention or change the producer behaviour.
-
In an emergency, where a producer floods the cluster, revoke that producer’s Access Control List (ACL) entries on the topic. This stops new data arriving without deleting the data already there.
-
UnderReplicatedPartitions
Kafka under replicated partitions.
-
Severity: warning
-
Default Threshold: More than 0
-
Description: 'There are {{ $value }} under replicated partitions on {{ $labels.kubernetes_pod_name }}'
-
Comment:
-
Some partitions are online but have fewer
In Sync Replicas(ISRs) than expected. -
This is common and happens when a transient network issue makes a broker unavailable.
-
It also occurs when a broker is offline. The Kafka cluster is still fully functional then, but in a strained state.
-
It can be an early warning for a larger problem when it does not resolve automatically.
-
-
Action: Wait for the issue to resolve itself. Investigate further if it lasts longer than 5 minutes.
AbnormalControllerState
Kafka abnormal controller state.
-
Severity: warning
-
Default Threshold: Other than 1
-
Description: 'There are {{ $value }} active controllers in the cluster'
-
Comment: A quick rolling restart of the brokers combined with slow metrics scraping temporarily registers two controllers, which reports as multiple active controllers.
-
Action: Wait for the issue to resolve itself. Investigate further if it lasts longer than 5 minutes.
OfflinePartitions
Kafka offline partitions.
-
Severity: warning
-
Default Threshold: More than 0
-
Description: 'One or more partitions have no leader'
-
Comment: This should never occur when partitions have enough replicas and no more than one broker is unavailable.
-
Action:
-
If multiple brokers are offline, this is a major incident and probably a loss of service. Quickly escalate and investigate why brokers are not available.
-
If one broker is offline and this alert triggers, identify which topics are affected and ensure the correct number of replicas.
-
UnderMinIsrPartitionCount
Kafka under min ISR partitions.
-
Severity: warning
-
Default Threshold: More than 0
-
Description: 'There are {{ $value }} partitions under the min ISR on {{ $labels.kubernetes_pod_name }}'
-
Comment: When below minimum
In Sync Replicas, most producers stop writing records into Kafka to ensure no data loss on broker failure. -
Action: This can only occur when at least two brokers are down, so this is part of a cluster outage. Quickly escalate and investigate why multiple brokers are unavailable.
OfflineLogDirectoryCount
Kafka offline log directories.
-
Severity: warning
-
Default Threshold: More than 0
-
Description: 'There are {{ $value }} offline log directories on {{ $labels.kubernetes_pod_name }}'
-
Comment: This signals an issue with the underlying storage. Some log directories are missing or unreachable.
-
Action: Involve the infrastructure team and investigate the root cause of the failure.
ScrapeProblem
Prometheus unable to scrape metrics from {{ $labels.kubernetes_pod_name }}/{{ $labels.instance }}.
-
Severity: major
-
Default Threshold: More than 0
-
Description: 'Prometheus was unable to scrape metrics from {{ $labels.kubernetes_pod_name }}/{{ $labels.instance }} for more than 3 minutes'
-
Comment: None of the alerts work while metrics are not scraped, so this is a critical issue.
-
Action:
-
Check the cluster health and the broker Pod status straight away. Without Prometheus the cluster is unmonitored, so watch it directly until scraping returns.
-
Investigate where the problem originates. It can sit in the upstream Prometheus when the platform runs in a federated setup.
-
KafkaBrokerContainersDown
All kafka containers down or in CrashLookBackOff status.
-
Severity: major
-
Default Threshold: All absent
-
Description: 'All
kafkacontainers have been down or in CrashLookBackOff status for 3 minutes' -
Comment: The whole Kafka cluster is broken. This is a critical incident.
-
Action:
-
Escalate, the Kafka service is unavailable.
-
Look into Kubernetes Events for scheduling issues or any missing volumeMounts like Secrets.
-
UnderReplicatedPartitionsandOfflinePartitionsfire before this one, so it should never be the first alert you see.
-
KafkaContainerRestartedInTheLast5Minutes
One or more Kafka containers restarted too often.
-
Severity: warning
-
Default Threshold: > 2 restarts
-
Description: 'One or more Kafka containers were restarted too often within the last 5 minutes'
-
Comment: This also occurs during broker updates, where it is not a problem.
-
Action: If this occurs outside of maintenance, check for memory usage (OOMKilled Pods) or crash looping Pods. Check Kubernetes Events.
Strimzi
These alerts come from the Strimzi operator rather than from Kafka itself.
ClusterOperatorContainerDown
Cluster Operator down.
-
Severity: major
-
Default Threshold: Less than 1
-
Description: 'The Cluster Operator has been down for longer than 90 seconds'
-
Comment: On a stable cluster the impact is low, because Kafka keeps functioning without an active Strimzi Cluster Operator.
It becomes a problem as soon as the cluster needs a change applied. -
Action: The Strimzi Operator Pod should be in a Ready state. Investigate whether the Pod is crash looping or whether Kubernetes cannot schedule it.