How to Import the Axual Grafana Dashboards
This guide shows you which dashboards Axual publishes for the platform, the two ways to get them into Grafana, and how to add the alerting rules that belong beside them.
Type |
How-to guide |
Goal |
Get the Axual dashboards and alerting rules into the Grafana and Prometheus the on-call rota watches. |
Audience |
Platform Operator with permission to import Grafana dashboards, or to commit to the repository Grafana reads them from. |
When to use |
Once the monitoring stack is collecting the platform’s metrics, and before go-live. |
Install these in the central Grafana rather than a per-cluster one, for the same reason the metrics are usually federated. The people who look at a dashboard during an incident open one Grafana, not one per cluster.
Prerequisites
Confirm the following before you begin.
Access and permissions required
You need the following access and permissions:
-
Permission to import dashboards in the target Grafana, or write access to the repository its dashboards are loaded from.
-
Permission to create
PrometheusRuleresources, for the alerting section.
Resources that must exist before starting
The following must already exist:
-
A Grafana instance with the platform’s Prometheus as a data source. A dashboard imported against a Prometheus that is not scraping the platform renders empty panels rather than an error, so confirm collection first: see Verify the platform is being collected.
Choose how to import
Two routes exist, and the difference is whether the dashboards become managed state or stay a one-off action:
-
Import the JSON file through the Grafana user interface. Quickest, and appropriate for trying a dashboard out.
-
Create a ConfigMap holding the JSON, so Grafana picks it up on start. Slower to set up, and the right choice for anything permanent, because the dashboard then lives in git like the rest of the deployment.
Prefer the ConfigMap route for a production estate. A dashboard imported by hand is invisible to review and lost when the instance is rebuilt.
The dashboards
Six dashboards cover the platform. The first two are Axual’s, and the rest come from Strimzi and cover Apache Kafka and the operator.
| Dashboard | Contents |
|---|---|
Cluster status at a high level: health, data rates and request rates. Start here. |
|
Rest Proxy specifics: schema status, produce and consume status, and the common JVM metrics. |
|
The Kafka clusters, their brokers and associated components. |
|
Per consumer group and per topic: offsets, lag and partition metrics. |
|
The Strimzi operator itself: resource usage, reconciliation status and other operational metrics. |
|
Kafka Connect connectors: throughput and worker performance. |
Once they render, use them to record a baseline while the system is healthy. What Grafana is for lists the questions a baseline answers and an incident cannot.
Add the alerting rules
Alerts fire from PrometheusRule resources. The Axual charts ship some, enabled per component with prometheusRule.enabled, and those cover the platform’s own components. Apache Kafka’s need adding.
Start from the Kafka PrometheusRules Strimzi publishes as an example and tailor the thresholds to your estate. Install them wherever the on-call operators' Prometheus reads its rules, which for a federated setup is the central instance rather than the cluster.
| Set the thresholds with the team that holds the standby rota. A rule with the wrong threshold is worse than no rule: it either fires constantly and gets muted, or never fires and is trusted. |
Confirm the rules loaded in Prometheus under Status > Rules. For what each Axual alert means when it fires, see Acting on Alerts.