Monitoring & Metrics
This guide explains what the Axual Platform emits as metrics, how each component exposes them, what the Kubernetes resources that collect them are for, and why Prometheus, Grafana and AlertManager are the tools Axual builds around.
Type |
Explanation |
Goal |
Understand what the platform measures and which tool is responsible for which part of it. |
Audience |
Platform Operators and anyone designing monitoring or alerting for the platform. |
When to use |
Read before setting monitoring up, and when deciding what to alert on. |
None of the tools discussed here ships with the Axual Platform. The platform’s side of monitoring is emitting metrics and declaring how to collect them; running the stack that collects them belongs to the infrastructure. For the procedure, see Installing the Monitoring Stack.
What the platform emits
Metrics are numerical data an application emits while running: how long it has been up, how many requests a second it handles, how much disk or memory is left. On their own they are numbers; their value is that graphs, dashboards and alerts are all built from them, so what is not emitted cannot be watched.
Two mechanisms cover the platform.
- Spring Boot Actuator
-
Most Axual components are Spring Boot applications and expose metrics through Actuator. Axual enables only a few of its endpoints, deliberately, because the rest expose more than a monitoring system needs. Actuator also backs the Kubernetes liveness and readiness probes, on
/actuator/health/liveness, so the same endpoint that reports metrics is what decides whether Kubernetes considers a pod healthy. - Kafka Exporter
-
Apache Kafka is not a Spring Boot application, so its metrics come from the Strimzi Kafka Exporter, which Axual enables by default on the clusters it installs.
For the metrics an individual component emits, see that component’s page under Axual Architecture & Components, or read them directly from its /actuator/prometheus endpoint after port-forwarding the Service.
How the metrics are collected
Emitting a metric and collecting it are separate problems, and Kubernetes solves the second with custom resources rather than configuration files. The Axual charts ship three of them, so a Prometheus Operator in the cluster discovers what to scrape without anyone editing a Prometheus configuration:
-
PodMonitors, which target pods directly. Apache Kafka uses these.
-
ServiceMonitors, which target a Service. Most components use these.
-
PrometheusRules, which declare alerting rules alongside the thing they alert on.
That design has one consequence worth knowing before you debug it: the resources are declarative and the operator decides whether to act on them, usually by matching a label. A monitor that exists and is never scraped is the normal failure, and it looks like missing metrics rather than a configuration mistake. Make the Axual monitors discoverable covers the label side of it.
Why Prometheus
Prometheus is what Axual uses and recommends, because the operator pattern above is native to it and because it is what most Kubernetes estates already run. Another monitoring, visualisation and alerting stack works on the same metrics, and this documentation does not cover configuring one.
Prometheus pulls rather than receives. It scrapes its targets on an interval, typically 20 to 30 seconds, which is why a dashboard refreshed faster than that shows the same numbers.
| Prometheus, Grafana and AlertManager are not part of the Axual Platform. An infrastructure team is expected to provide them. |
Most deployments end up with a central Prometheus that federates several Kubernetes clusters, so dashboards and alerts have one data source rather than one per cluster. That is worth knowing early, because it decides whether the platform’s monitors are scraped locally and federated up, or scraped by the central instance directly.
What Grafana is for
Grafana turns the metrics Prometheus holds into dashboards. Axual publishes a set for the platform, and importing them is a procedure: see How to Import the Axual Grafana Dashboards.
The reason to install them early, rather than when something breaks, is that they answer questions which only have answers while the system is healthy:
-
What is the normal system load?
-
Can the brokers handle it?
-
At what level do CPU, network or disk usage start to correlate with poor performance?
-
Are any scheduled jobs slowing the system at particular times?
-
Do the Connect or Rest Proxy instances need scaling for extra load?
-
Are the brokers loaded evenly, or is something like Cruise Control needed to balance them?
-
What peak load can the system handle comfortably?
| Record baseline statistics while performance is good. Grafana is most often opened during an incident, and at that point a graph without a baseline shows a number with nothing to compare it to. |
Why alerting is a collaboration
Axual suggests AlertManager, because it integrates with Prometheus directly. In practice most production estates already have an alerting system that the on-call rota watches, and routing into that is worth more than introducing another one. Alerting on log data instead is also viable: see Create the threshold alert.
Alerts fire from PrometheusRules. Some ship with the Axual charts, and others belong in the central location the on-call operators watch, which is why installing them is a deliberate step rather than a side effect of installing the platform.
| Agree the rules with the team that holds the standby rota. A gap in alerting surfaces as a Kafka outage, typically a filled disk or an unavailable cluster, which costs far more than writing the rule that would have caught it. |
For each alert the charts provide and what to do when it fires, see Acting on Alerts.