Installing the Monitoring Stack

This guide covers standing up the stack that collects what the Axual Platform emits, or pointing an existing one at it. Prometheus, Grafana and AlertManager are not part of the Axual Platform and are expected from the infrastructure it runs on, so this guide is about connecting the platform to that stack rather than about operating the stack itself.

Type

How-to guide

Goal

Collect the Axual Platform’s metrics, logs and traces in a monitoring stack, and alert on them.

Audience

Platform Operators, working with whoever runs the organisation’s central monitoring stack.

When to use

Stage 7 of The installation order, alongside or after the layers are installed.

What each signal contains, and how to work with it once collected, is covered by the pages listed under Monitoring.

Contents

The sections below cover each task in this guide:

Prerequisites

Confirm the following before you begin.

Access and permissions required

You need the following access and permissions:

  • Permission to install Helm releases in the cluster, and to register custom resource definitions, which are cluster-scoped. Where a central team runs the monitoring stack, that team does this part and you supply the labels the Axual monitors need.

  • Read access to the namespaces the platform runs in, for the kubectl checks below.

  • Access to the Prometheus and Grafana interfaces, because two of the checks are made there.

Tools and versions required

You need the following tools:

  • helm >= 3.12, to install the monitoring chart.

  • kubectl >= 1.28, configured for the target cluster.

Resources that must exist before starting

The following must already exist:

  • An Axual Platform installation, since its charts are what render the monitors. This is stage 7 of The installation order.

  • The Prometheus Operator’s custom resource definitions, registered before any Axual chart is applied with a monitor enabled. A chart applied ahead of them leaves the monitor out, which is one of the two causes of a missing monitor in Verify the platform is being collected.

Install or connect the stack

Two routes exist, and the choice depends on whether the cluster already runs monitoring. The Prometheus stack is one of the components Axual relies on and does not ship, listed in External Dependencies.

With no monitoring in place, install the kube-prometheus-stack chart, which brings Prometheus, the Prometheus Operator, AlertManager and Grafana together. No install command for it appears here, because Axual neither ships nor versions that chart. The chart’s own documentation holds the values and the helm install line, and that is the copy to follow. Axual publishes an example configuration of the chart in its professional services local-config repository.

A finished install leaves four things running:

  • Prometheus.

  • The Prometheus Operator, with its custom resource definitions registered.

  • AlertManager.

  • Grafana, to import the Axual dashboards into.

That operator is the one capability the platform needs from whichever stack you run. The Axual charts create PodMonitor, ServiceMonitor and PrometheusRule resources and nothing else. All three are custom resource types the Prometheus Operator defines, so a stack without the operator collects no Axual metrics however well it is configured. See Monitoring & Metrics for what each component exposes.

Once Prometheus is running, components that read from it are given its in-cluster address. Metrics Exposer takes a map of them, so a platform can point at more than one Prometheus:

prometheus-urls:
  default: http://kube-prometheus-stack-prometheus:9090
  axual: http://axual-prometheus-stack-prometheus:9090

With a central Prometheus already in place, federate from this cluster into it rather than running a second stack. The operator’s custom resource definitions still have to be registered in the cluster the platform runs in, since that is where the Axual monitors are created. The scrape configuration that performs the federation is not given here.

Make the Axual monitors discoverable

The Axual charts create the monitors, but they do not label them. Every chart exposes the same three blocks, and labels defaults to {}:

Replace every <VALUE> placeholder with your own value before applying the configuration.
<COMPONENT>:            (1)
  serviceMonitor:
    enabled: true
    interval: 30s
    scrapeTimeout: 10s
    labels: {}          (2)
  prometheusRule:
    enabled: true
1 The component’s key in values.yaml, for example platform-manager or api-gateway.
2 Empty by default, which is what makes this the step people miss.
A Prometheus Operator is usually configured to discover only the monitors carrying a specific label, and the Axual charts add none. Until labels matches what your operator selects on, the monitors exist and are never scraped, which looks like a metrics problem rather than a configuration one. Either set labels to the value your stack selects on, or widen the operator’s own selector to match everything.

Apache Kafka is the exception: its monitors are podMonitor rather than serviceMonitor, and they are keyed per workload, so each one is enabled and labelled separately.

podMonitor:
  kafka:
    enabled: true
    labels: {}
    interval: "30s"
    scrapeTimeout: "20s"
  entityOperator:
    enabled: true
    labels: {}

Each chart lists its own monitor keys. Kafka’s podMonitor keys are in Kafka Chart Values Reference and the Axual Distributor’s in Distributor Chart Values Reference. Every other component documents its serviceMonitor and prometheusRule keys in its generated Helm readme, for example Metrics Exposer 1.7.0 Helm Readme and Platform Manager 16.0.0 Helm Readme.

Verify the platform is being collected

Check each signal separately. A stack that collects two of the three looks healthy from the dashboard that happens to work.

Start with the monitors themselves, because a monitor the operator never discovered is the most common failure and the least visible:

kubectl get servicemonitor,podmonitor,prometheusrule --all-namespaces

Every enabled component should appear. A component whose monitor is missing has serviceMonitor.enabled false, or the chart was applied before the Prometheus Operator’s custom resource definitions existed.

Then confirm Prometheus is scraping them, rather than only knowing about them. Open Prometheus and check Status > Targets: each Axual target should read UP. A monitor that exists while no target appears is the label problem described above. Monitoring & Metrics covers which components expose what, and which Grafana dashboards to import once the targets are up.

For logs, confirm lines from an Axual pod arrive in the central store and are indexed as fields rather than as unparsed text. How to Search Logs in Kibana covers querying them, and How to Ship Logs to a Central Stack covers the JSON format that makes the fields queryable.

For traces, confirm spans from a request through the API Gateway reach the tracing backend. How to Enable Distributed Tracing covers enabling OpenTelemetry and pointing it at a collector.

Finally, confirm alerting works end to end by checking that the rules the charts ship are loaded, in Prometheus under Status > Rules. Acting on Alerts lists what each one means when it fires.